Millimeter wave radar virtual channel prediction method and system based on self-supervised learning, and medium

By generating virtual channels through self-supervised learning, the problem of insufficient angular resolution of millimeter-wave radar is solved, and high angular resolution radar data generation is achieved, which improves target detection accuracy and perception performance and is applicable to various hardware and annotation methods.

CN121958779APending Publication Date: 2026-05-01XIDIAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XIDIAN UNIV
Filing Date
2026-01-15
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing millimeter-wave radars suffer from insufficient angular resolution due to hardware limitations, making it difficult to effectively distinguish targets in complex environments. Current technologies rely on high-resolution data training or costly hardware upgrades, and the bounding boxes are not accurately covered, affecting the target detection effect.

Method used

A self-supervised learning method is adopted to predict virtual channels through deep convolutional networks. By utilizing the radar's own channel structure and inherent correlations, high-angle resolution radar data is generated. This includes radar channel signal preprocessing, multi-scale feature extraction, feature pyramid fusion, and self-supervised training to generate high-resolution maps.

Benefits of technology

Without increasing hardware costs, it significantly improves angular resolution, enhances target detection accuracy and perception performance, is applicable to different hardware and annotation methods, and strengthens target discrimination and detection accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121958779A_ABST
    Figure CN121958779A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of millimeter wave radar and deep learning, in particular to a millimeter wave radar virtual channel prediction method and system based on self-supervised learning and a medium. According to the method, after radar channel signal preprocessing and recombination, multi-scale residual feature extraction, feature pyramid fusion and angle dimension up-sampling, virtual channel signal regression prediction and self-supervised training and loss constraint are carried out in sequence, radar data with high angle resolution are obtained; according to the method, on the premise of not depending on high-resolution radar data as external supervision, prediction and extension of a virtual channel are realized by using a channel structure and an internal association relationship of low-resolution radar data; self-supervised virtual channel prediction and high-angle-resolution atlas generation can be stably realized under multiple data sets, different hardware and different labeling modes, and a reliable, efficient and migratable technical scheme is provided for target detection and environmental perception of a millimeter wave radar with only a small number of antennas or low-angle-resolution low-hardware configuration.
Need to check novelty before this filing date? Find Prior Art

Description

A method, system, and medium for predicting virtual channels in millimeter-wave radar based on self-supervised learning. Technical Field

[0001] This invention relates to the fields of millimeter-wave radar and deep learning technology, and in particular to a method, system and medium for predicting virtual channels in millimeter-wave radar based on self-supervised learning. Background Technology

[0002] Millimeter-wave radar is widely used in fields such as autonomous driving and intelligent transportation due to its insensitivity to ambient light, strong penetration capability, and high speed resolution. Radar detects, locates, and tracks objects in the environment by emitting high-frequency electromagnetic waves and receiving reflected signals. Typically, millimeter-wave radar employs a MIMO (Multiple-Input Multiple-Output) antenna array, generating two-dimensional or three-dimensional maps of the target in range-Doppler (RD) and range-angle (RA) spaces by receiving multi-channel signals, thus providing spatial information about the target.

[0003] However, due to limitations in hardware such as size, cost, and power consumption, the number of transmitting and receiving antennas for vehicle-mounted millimeter-wave radar is usually limited. This results in a limited equivalent aperture of the radar array in the angular dimension and insufficient angular resolution. Insufficient angular resolution makes it difficult for the radar to effectively distinguish multiple targets that are close in spatial location in the angular direction, leading to problems such as target energy diffusion and sidelobe widening in the angular dimension. These problems are particularly prominent in application scenarios with dense targets or complex traffic environments, which can easily cause mutual interference between adjacent targets, resulting in confusion, missed detections, or decreased positioning accuracy, seriously affecting the reliability of the perception system.

[0004] To address the aforementioned issues, existing technologies typically improve angular resolution by increasing the number of radar antennas or enlarging the array aperture. However, this method relies on hardware upgrades, is costly, and has significant limitations in practical engineering applications. To address this, some research has attempted to introduce neural network methods. These methods utilize high-angular-resolution radar data as training samples, degrading it to low-resolution data for model training. This allows the system to learn the mapping relationships between multiple channels, predicting virtual channels or generating enhanced radar maps, thereby improving angular resolution. However, these methods usually rely on high-resolution, multi-channel radar data as supervisory signals. Their training process requires pre-acquiring high-quality radar data under full array conditions, making them unsuitable for applications with only a few antennas or low-angular-resolution radar.

[0005] Furthermore, existing radar-generated RA or RAD maps have limited spatial detail, making it difficult to meet the feature resolution requirements of high-precision target detection models. Simultaneously, radar data annotation typically relies on clustering reflection intensity on low-angular-resolution maps to generate bounding boxes surrounding the target. Because the annotation process is based on the original map with low angular resolution, the target energy distribution exhibits significant sidelobes and diffusion, resulting in insufficient precision in the spatial range covered by the bounding boxes. With improved radar map resolution, these original bounding boxes cannot accurately reflect the true spatial distribution of the target; direct use may introduce supervision bias, affecting the training performance of the target detection model.

[0006] How to improve the angular resolution of millimeter-wave radar without increasing hardware costs or interrupting the existing radar structure, and further construct reliable supervisory labels suitable for high-resolution maps to ensure the training stability and detection accuracy of the target detection model, is a key problem that urgently needs to be solved in this field.

[0007] In the prior art, patent publication number CN120687944A, entitled "A Self-Supervised Contrast Mask Reconstruction Method and Apparatus for Radar Signal Modulation Recognition," discloses a radar signal modulation recognition technology based on self-supervised learning. This invention acquires radar signals and constructs a radar modulation signal dataset containing labeled and unlabeled data. The radar signals are then converted into time-frequency images, and a self-supervised contrast mask image reconstruction model, including online and target branches, is built. In this technical solution, data augmentation and random masking operations are applied to the time-frequency images to generate dual views, which are then input into the online and target branches respectively. The model is pre-trained using unlabeled data to jointly optimize the reconstruction loss and contrast loss. Subsequently, the pre-trained weights are transferred to the downstream radar signal modulation recognition network and fine-tuned using a small amount of labeled data, thereby achieving radar signal modulation recognition in complex electromagnetic environments and improving modulation recognition accuracy.

[0008] The aforementioned technical solution, by introducing self-supervised contrastive learning and mask reconstruction mechanisms, improves radar signal modulation recognition performance while reducing reliance on large amounts of labeled data, demonstrating certain technical effectiveness. However, the self-supervised learning process of this invention primarily targets the modulation type recognition task of radar signals, aiming to extract abstract feature representations beneficial for classification and recognition. It does not involve modeling the mapping relationships between radar array channel data, nor does it address the issues of virtual channel prediction or angular resolution improvement under low-channel, low-angle-resolution radar conditions. Furthermore, the mask reconstruction and contrastive learning process operates at the time-frequency image level of radar signals, failing to model the multi-channel spatial information or array structural characteristics of the radar. Consequently, it cannot solve the problem of insufficient angular resolution caused by the limited number of radar hardware channels, nor can it generate equivalent high-angle-resolution radar data using only low-angle-resolution radar data. Therefore, it still falls short of meeting the angular resolution requirements of target detection and spatial perception tasks. Summary of the Invention

[0009] To overcome the shortcomings of the prior art, the present invention aims to propose a method, system, and medium for predicting virtual channels of millimeter-wave radar based on self-supervised learning. This method, without relying on high-resolution radar data as external supervision, utilizes the channel structure and inherent correlation of low-resolution radar data itself to predict and expand virtual channels, thereby obtaining radar data with higher angular resolution to improve target detection and perception performance.

[0010] To achieve the above objectives, the technical solution adopted by the present invention is as follows: Firstly, a method for predicting virtual channels of millimeter-wave radar based on self-supervised learning, comprising the following steps: Step S1: Radar channel signal preprocessing and reconstruction: The first channel signal from the multi-channel signal of the millimeter-wave radar is input into a deep convolutional network model, and the first channel signal is reconstructed using the multiple-input multiple-output (MIMO) precoding module in the deep convolutional network model to obtain reconstructed features; Step S2: Multi-scale residual feature extraction (backbone network): The reconstructed features obtained in Step S1 are input into the backbone network of the deep convolutional network model for feature extraction to obtain multi-scale features; Step S3: Feature pyramid fusion and angular dimension upsampling: The multi-scale features obtained in Step S2 are input into the feature pyramid decoding network (FPN) of the deep convolutional network model. Upsampling is performed in the Decoder, and the upsampled high-level features are concatenated with the corresponding low-level features at the channel dimension. The concatenated multi-scale features are then fused using convolutional blocks, which are used to reshape and enhance the multi-scale features. Through the above top-down feature fusion method, a radar feature map with high spatial resolution is obtained. Step S4: Virtual channel signal regression prediction: The radar feature map obtained in step S3 is input into the prediction head of the deep convolutional network model. After convolutional operation of the fused features and angular dimension prediction, the predicted virtual second channel signal is output. The prediction head includes multiple two-dimensional convolutional layers, batch normalization, and activation functions, which are used to predict and reconstruct angular dimension features. This prediction head realizes the modeling of virtual angular information of unobserved channels by convolutional operation of the fused features and angular dimension prediction, thereby improving the angular resolution of the radar spectrum without relying on an increase in the number of physical antennas. Step S5: Self-supervised training and loss constraint: The second channel signal in the millimeter-wave radar multi-channel signal obtained in step S1 is used as the supervision target, and the training is performed on the signal obtained in step S4. The virtual second channel signal is self-supervised to obtain a virtual channel prediction module. The second channel signal is input into the virtual channel prediction module, and the virtual third channel signal is output through the prediction head of the virtual channel prediction module.

[0011] Furthermore, in step S1, the first channel signal is a two-dimensional or three-dimensional tensor, which includes at least distance dimension and Doppler dimension information.

[0012] Furthermore, the multiple-input multiple-output precoding module in step S1 adopts a two-dimensional convolutional structure with a kernel size of "1, N_Tx" and an expansion rate of "1, N_Rx" to fuse channel information in the dimensions of the transmit antenna and the receive antenna.

[0013] Furthermore, the backbone network in step S2 includes an initial convolutional layer and a multi-level residual feature extraction module. The initial convolutional layer uses a 3×3 convolutional kernel with a stride of 1 and padding of 1 for two-dimensional convolution, combined with batch normalization and the ReLU activation function. The multi-level residual feature extraction module is a bottleneck structure, which includes the following layers in sequence: a first convolutional layer with a 1×1 convolutional kernel for channel dimensionality reduction; a second convolutional layer with a 3×3 convolutional kernel with a stride of 1 or 2 for spatial feature extraction; and a third convolutional layer with a 1×1 convolutional kernel for channel expansion with an expansion ratio of 4. In the residual connection path, when the number of input channels is inconsistent with the number of output channels, a downsampling operation with a 1×1 convolution and a stride of 2 is introduced. The multi-level residual feature extraction module is cascaded to form four feature levels, with the feature resolution decreasing and the number of channels increasing at each level to obtain multi-scale radar spatial features.

[0014] Furthermore, the feature pyramid decoding network described in step S3 includes multiple deconvolution (transposed convolution) layers. The kernel size of the deconvolution layer is 3×3, the stride is 2, the padding is 1, and the output padding is 1, which is used to progressively restore the resolution of the feature map in the distance and angle dimensions.

[0015] Furthermore, in the prediction head described in step S4, the first convolutional layer uses a 3×3 convolutional kernel, with 256 input channels and 144 output channels. The stride in the angular dimension can be adaptively set to 1 or 2 according to the input size, and the padding is 1. Batch normalization and ReLU activation are applied after convolution. The second convolutional layer uses a 3×3 convolutional kernel, with 144 input channels and 96 output channels, a stride of 1, and padding of 1. Batch normalization and ReLU activation are applied after convolution. Subsequently, two 3×3 convolutional layers are used to maintain the number of channels at 96, further refining the angular dimension features. Finally, a single 3×3 convolutional layer outputs the predicted virtual channel signal, with the number of channels equal to the number of virtual channels to be predicted, thereby obtaining a distance-angle or distance-angle-Doppler spectrum with higher angular resolution.

[0016] Furthermore, in the self-supervised training described in step S5, L1 loss, L2 loss, and / or structural similarity loss (SSIM) are introduced to constrain the consistency of the virtual second channel signal with the second channel signal in terms of amplitude distribution and spatial structure, thereby obtaining the trained virtual channel prediction module.

[0017] Secondly, a high-angle-resolution radar map generation method based on the aforementioned virtual channel prediction method includes the following steps: Step T1: Channel stitching: The virtual third channel signal described in step S5 is stitched together with other channel signals of the millimeter-wave radar according to the channel dimension to form an extended radar array signal matrix, thereby constructing a denser virtual radar array. The prediction process of the virtual third channel signal described in step S5 does not rely on real labels and is used for virtual channel generation during the inference stage. Step T2: Frequency domain transformation: The extended radar array signal matrix described in step T1 is transformed by Fast Fourier Transform (FFT) in the channel dimension to generate a range-angle (RA) map or a range-angle-Doppler (RAD) map. Step T3: Amplitude normalization and map enhancement: The range-angle (RA) map or range-angle-Doppler (RAD) map described in step T2 is subjected to amplitude normalization and energy consistency correction processing to maintain the continuity and consistency of the target energy distribution, ensuring clear angular dimension information and obtaining a high-angle-resolution radar map.

[0018] Based on the original target bounding box, the target's center position, target category, and distance relative to the radar are obtained. Then, using specific parameters and algorithms, a corresponding confidence map is generated.

[0019] Thirdly, a system based on the aforementioned virtual channel prediction method is characterized by comprising the following modules: a radar channel signal preprocessing and reconstruction module: inputting the first channel signal from the millimeter-wave radar multi-channel signal into a deep convolutional network model, and using the multi-input multi-output precoding module in the deep convolutional network model to reconstruct the first channel signal to obtain reconstructed features; a multi-scale residual feature extraction module: inputting the reconstructed features into the backbone network of the deep convolutional network model for feature extraction to obtain multi-scale features; and a feature pyramid fusion and angular dimension upsampling module: inputting the multi-scale features into the feature pyramid decoding network of the deep convolutional network model for upsampling, and fusing the upsampled high-level features with the corresponding low-level features in the channel. The radar features are stitched together along the channel dimension, and the stitched multi-scale features are fused through convolutional blocks to obtain a radar feature map with high spatial resolution. A virtual channel signal regression prediction module inputs the radar feature map into the prediction head of a deep convolutional network model. After convolutional operations on the fused features and angular dimension prediction, it outputs the predicted virtual second channel signal. The prediction head includes multiple two-dimensional convolutional layers, batch normalization, and activation functions. A self-supervised training and loss constraint module uses the second channel signal in the millimeter-wave radar multi-channel signal as the supervision target, performs self-supervised training on the virtual second channel signal, and obtains a virtual channel prediction module. The second channel signal is input into the virtual channel prediction module, and the prediction head of the virtual channel prediction module outputs a virtual third channel signal.

[0020] Fourthly, a computer-readable storage medium stores a computer program that, when executed by a processor, implements the virtual channel prediction method; the computer-readable storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard disk, magnetic disk, or optical disk.

[0021] Compared with existing technologies, the beneficial effects of this invention are as follows: 1. Cross-dataset prediction capability: On the RADDet dataset, this invention uses the transmitting array signal of millimeter-wave radar as input, generates virtual receiving array channels through a self-supervised training virtual channel prediction module, and stitches them together to obtain a high-angle resolution RA / RAD map; simultaneously, on the RADinal dataset, using the receiving array signal as input, it generates virtual transmitting array channels through a self-supervised training virtual channel prediction module, and similarly stitches together a high-angle resolution map; thus, it is demonstrated that this invention can flexibly select the input channel type for different datasets, realize virtual channel prediction of transmitting and receiving arrays, and the prediction process is entirely based on self-supervision and does not rely on target detection annotation.

[0022] 2. Cross-hardware applicability: The method of this invention has good adaptability to radar hardware configuration. Regardless of the differences in the number of transmitting antennas, the number of receiving antennas, or the array arrangement, the virtual channel prediction module can effectively generate virtual channels. Therefore, this invention can support universal prediction for different radar models or array designs, enhancing the feasibility and versatility of actual deployment.

[0023] 3. Feasibility Verification Across Annotation Methods: The virtual channel prediction stage employs self-supervised training, independent of object detection annotations. To verify the amplified effect of the virtual channel prediction, annotation methods from different datasets can be used: The RADDet dataset, due to its 2 transmit and 4 receive antennas, has poor angular resolution, resulting in errors in the edge accuracy of bounding box annotations. To reduce the impact of edge supervision caused by incorrect bounding box annotations, the object detection annotations use Gaussian spot confidence maps transformed from bounding box annotations. The RADIAL dataset, with its 12 transmit and 16 receive antennas, has high angular resolution and extremely high annotation accuracy. Annotations such as laser point clouds from its original dataset are used for object detection training and testing. This demonstrates that the virtual channel model prediction results of this invention can be effectively used for subsequent object detection performance verification under different annotation methods.

[0024] 4. Improved angular resolution and target discrimination capability: By generating virtual channels and performing channel stitching and FFT processing, this invention can significantly improve the angular resolution of RA / RAD maps, making adjacent targets clearer in the angular dimension and reducing energy diffusion and sidelobe interference; thereby improving the accuracy of target detection in complex traffic scenarios, especially for dense targets or small targets.

[0025] In summary, this invention utilizes the channel structure and inherent correlations of low-resolution radar data to achieve virtual channel prediction and expansion, thereby obtaining radar data with higher angular resolution. It can stably achieve self-supervised virtual channel prediction and high angular resolution map generation under multiple datasets, different hardware, and different annotation methods. It provides a reliable, efficient, and transferable technical solution for target detection and environmental perception of millimeter-wave radars with only a few antennas or low angular resolution and low hardware configuration. Meanwhile, the annotation is only used to verify the target detection effect after prediction and expansion. Attached Figure Description

[0026] Figure 1 is a schematic diagram of the prediction structure of the virtual channel prediction module of the present invention.

[0027] Figure 2 is a schematic diagram of the virtual channel prediction and high-angle resolution radar map generation process of the present invention.

[0028] Figure 3 is a schematic diagram of RADDet target labeling, channel-enhanced radar representation, and a supervision mechanism based on confidence map in an embodiment of the present invention.

[0029] Figure 4 is a comparative example of the RAD representation of the top 2T4R and bottom 3T4R in the RADDet dataset in an embodiment of the present invention. Detailed Implementation

[0030] The invention will be further described in detail below with reference to Figures 1 to 4: The millimeter-wave radar virtual channel prediction method and high-angle resolution radar map generation method of the present invention actually include three main parts: a virtual channel prediction module, a channel expansion and map generation module, and a confidence map generation module. On this basis, the target detection and evaluation module is used to train and evaluate the target detection of the channel-expanded high-resolution radar RA / RAD map. The specific descriptions of each module are as follows: 1. Virtual Channel Prediction Module: The virtual channel prediction module of the present invention is based on a deep convolutional network architecture. It learns the mapping relationship between millimeter-wave radar channels through self-supervised learning to predict and generate virtual channels that have not yet been acquired, thereby improving the radar angular resolution.

[0031] As shown in Figure 1, the virtual channel prediction module is not limited to the RADDet dataset in the example. In the experiment, it also works on another dataset with different specifications, the RADIAL dataset. The virtual channel prediction module used for these different datasets only requires changing some parameters.

[0032] 1.1 Structural Description The module mainly includes four sub-modules: (1) Radar channel signal preprocessing and reconstruction sub-module The first channel signal of the millimeter-wave radar multi-channel signal is input into the deep convolutional network model. The multi-input multi-output precoding module in the deep convolutional network model is used to reconstruct the first channel signal to obtain the reconstructed channel features; (2) Multi-scale residual feature extraction sub-module Constructs a feature pyramid through four residual blocks (Bottleneck or BasicBlock) to generate multi-scale features x1~x4 and retain the bottom feature x0. The residual blocks are stacked through convolution, batch normalization and ReLU activation function to ensure the expressive power of deep features and the stability of gradient propagation. The multi-scale residual feature extraction submodule constructed in this invention is used to extract multi-scale spatial features from the original radar channel signals. The specific steps are as follows: On the preprocessed output channel, features are extracted sequentially through four residual blocks (Bottleneck or BasicBlock) to form a feature pyramid structure. The four residual blocks are denoted as block1, block2, block3, and block4, respectively. Block1: extracts lower-level features to capture local details and basic channel relationships; the output feature is denoted as x1. Block2: further extracts mid-level features to enhance spatial context information and outputs feature x2. Block3: extracts high-level semantic features to capture target scale changes and more global angle information and outputs feature x3. Block4: extracts the deepest features to enhance deep semantic expression and inter-channel correlation and outputs feature x4. Each residual block consists of convolutional layers (1×1, 3×3, 1×1), batch normalization layers (BatchNorm), and ReLU activation functions stacked together. The residual connections ensure the expressive power of deep features and the stability of gradient propagation.

[0033] Low-level feature preservation: The feature preserved after initial convolution processing before passing through the residual block is denoted as x0. This low-level feature contains the low-level spatial structure information and antenna channel distribution information of the original signal, which can be used for high-resolution recovery in the decoding stage.

[0034] Multi-scale feature output: Finally, the backbone network outputs five multi-scale feature sets x0, x1, x2, x3, and x4, which are used by the subsequent feature pyramid fusion and angular dimension upsampling submodules to achieve high angular resolution radar map generation.

[0035] (3) The feature pyramid fusion and angular dimension upsampling submodule uses top-down deconvolution (ConvTranspose2d) to progressively upsample features and restore angular resolution. After upsampling at each layer, the corresponding backbone network features are fused through channel concatenation to enhance the expressive power of spatial information. After each fusion, high-level features are further extracted through convolutional residual blocks (BasicBlock) to obtain the final predicted feature map.

[0036] (4) The virtual channel signal regression prediction submodule maps the number of channels in the final predicted feature map to the number of target virtual channels, generating the final virtual channel radar signal. It includes multi-layer convolution and batch normalization, and can adaptively adjust the number of channels and convolution stride according to the input angle size.

[0037] 1.2 Principle Description As shown in Figure 2, the original multi-channel radar signal is used as input, and a self-supervised mechanism is used to predict neighboring virtual channels: Input channel x → Radar channel signal preprocessing and reassembly submodule → Channel reassembly features → Multi-scale residual feature extraction submodule → Multi-scale features → Feature pyramid fusion and angular dimension upsampling submodule → Final prediction feature map → Virtual channel signal regression prediction submodule → Generation of final virtual channel radar data Self-supervised training objective: Using the original channel data as input, neighboring or subsequent channels are used as prediction targets, and the network learns the implicit mapping relationship between channels. The loss function can be L1, L2, and / or structural similarity loss (SSIM), and spectral fidelity loss, phase consistency loss, energy preservation loss, or adversarial loss are also used as prediction optimization objectives. Different loss functions can constrain the virtual channel to maintain the real physical structure, thereby obtaining the same angular resolution improvement effect and ensuring that the predicted channel is consistent with the real channel in amplitude and spatial structure.

[0038] In this way, virtual channel signals with equivalent aperture expansion can be generated without external sensors or additional annotations, thereby improving angular resolution.

[0039] 1.3 Action Relationship Description: Data Preprocessing: Extract the backbone input part of the radar multi-channel signal into the multi-channel feature extraction of the deep convolutional network. After the radar channel signal preprocessing and reconstruction submodule, the channels are reconstructed, expanded and convolved to generate preliminary features.

[0040] High-level feature extraction: The preliminary features are processed layer by layer by residual blocks to extract high-dimensional spatial information to obtain high-level features, while retaining feature maps x0~x4 at different scales.

[0041] Virtual channel decoding: The high-level features are upsampled through deconvolution and sequentially concatenated and fused with the corresponding backbone layer features. Each fusion layer further extracts high-level features through BasicBlock convolutional blocks, gradually restoring the radar angular resolution.

[0042] Virtual channel generation: The final decoder outputs the final predicted feature map and inputs it into the virtual channel signal regression prediction submodule. The features of the final predicted feature map are mapped to the target number of virtual channels to obtain virtual channel signals that can be directly spliced.

[0043] 1.4 Module Features and Advantages No additional hardware or external sensors required: Utilizing the multi-receiver channel signals of existing radar, virtual channels are generated through self-supervised learning prediction to achieve equivalent aperture expansion without adding antennas or introducing multi-modal sensors.

[0044] Improve angular resolution: By predicting virtual channels and stitching them with the original channels, the angular resolution of the RA / RAD map is significantly improved, making the target energy more concentrated and the sidelobe compressed, thereby enhancing the target's distinguishability in the angular dimension.

[0045] Data-driven, self-supervised learning: The module learns the implicit spatial mapping relationship between channels through the CNN structure and is trained in a self-supervised manner, eliminating the need for manual annotation, reducing labor costs and improving training efficiency.

[0046] It is compatible with multiple subsequent detection models: the generated high-resolution maps can be directly used as input to object detection models (such as CNN, FPN, U-Net, etc.), improving detection accuracy and model generalization ability.

[0047] Highly scalable: The module design is flexible and can adjust the number of input channels and output virtual channels according to the number of radar antennas and target detection task requirements, making it suitable for different radar equipment and application scenarios.

[0048] 2. Channel Expansion and Map Generation Module: The channel expansion and map generation module of this invention is responsible for stitching the predicted virtual channels with the original radar channels and generating an enhanced high-resolution RA / RAD map, providing more refined spatial features for subsequent target detection.

[0049] 2.1 The structural description module mainly includes the following parts: (1) The channel splicing unit splices the predicted virtual channel signal with the original radar channel signal in the channel dimension to form an expanded multi-channel signal. The splicing method is flexible, and furthermore, the channel dimension length can be adjusted according to the radar array aperture and the number of predicted channels.

[0050] (2) The FFT and spectrum generation unit performs Fast Fourier Transform (FFT) on the extended multi-channel signal to generate a distance-angle (RA) or distance-angle-Doppler (RAD) spectrum; furthermore, a window function (such as Hanning window) can be used to process the spectrum to improve the smoothness of the spectrum and the sidelobe suppression effect.

[0051] (3) The standardization and post-processing unit normalizes the amplitude of the generated map to ensure the scale consistency of data between different channels. Furthermore, amplitude thresholding or confidence weighting can be performed to provide reliable input for subsequent target detection.

[0052] 2.2 Principle Explanation: The virtual channel prediction module generates virtual channel signals that simulate uncollected radar receiving channels. After stitching, this is equivalent to increasing the array aperture. The added channels can significantly improve angular resolution, making the energy distribution of the target more concentrated in the angular dimension, compressing sidelobes, and reducing the overlap of angular energy distributions of adjacent targets.

[0053] The principle of spectrum generation: By performing FFT on the expanded multi-channel signal, the time-domain or sampled-domain signal can be mapped to the frequency domain space to obtain the RA or RAD spectrum. The RA / RAD spectrum can intuitively reflect the spatial distribution characteristics of the target in the dimensions of distance, angle, and velocity.

[0054] The expanded map has higher angular resolution than the original map, providing finer-grained target features for downstream detection models.

[0055] 2.3 Description of Action Relationships: Virtual Channel Stitching: The virtual channels output by the virtual channel prediction module are stitched together with the original radar channels according to the channel dimension to form an extended signal matrix.

[0056] FFT Transform: Performs an FFT transform on the extended signal in the angular and distance dimensions to generate an RA or RAD spectrum. Windowing can be applied to the signal before the FFT to increase the main lobe concentration and reduce the sidelobes.

[0057] Atlas normalization and enhancement: Amplitude normalization is performed on the FFT output to maintain consistency in amplitude across different channels. Confidence-weighted processing can be applied to the atlas to generate training or inference inputs, depending on the target detection requirements.

[0058] Output extended map: Obtain a high-angular-resolution RA / RAD map for subsequent target detection or other radar sensing tasks.

[0059] 2.4 Module Features and Advantages: This module achieves equivalent aperture expansion and high-resolution map generation without requiring additional hardware or external sensors. The generated maps are fully compatible with the original signals and can be directly used for training and inference of existing target detection models. It supports adaptive channel expansion, which can be flexibly adjusted according to different radar configurations.

[0060] 3. Confidence Map Generation Module: As shown in Figure 3, the confidence map generation module of the present invention is used to map millimeter-wave radar target annotation information into a confidence map on a high-resolution RA / RAD map, thereby providing a more accurate supervision signal for the target detection model.

[0061] 3.1 The structural description module mainly includes the following parts: (1) Target center extraction unit: Input the original annotation information (such as target category, target center coordinates, distance information). Extract the spatial center position of each target on the map as the center point of the Gaussian response distribution.

[0062] (2) The Gaussian heatmap generation unit determines the scale and intensity of the Gaussian kernel based on the target category and distance information to generate a two-dimensional Gaussian distribution. The kernel size can be adaptively adjusted for different categories (such as vehicles, people, and motorcycles) or targets at near and far distances to reflect the real spatial distribution.

[0063] (3) The confidence map fusion unit overlays the Gaussian heatmaps of all targets to generate a complete confidence map. The overlapping areas can be normalized or weighted to ensure that the contribution of each target to the map is reasonable.

[0064] 3.2 Principle Explanation Core Idea: In high-angle resolution maps, target energy is concentrated, side lobes are compressed, and the original bounding box annotations are no longer spatially accurate. Converting target annotations into Gaussian heatmaps can simulate the probability distribution of targets in high-resolution maps, making the monitoring signal smoother and more spatially continuous.

[0065] Category and distance adaptation: Different target categories have different sizes and reflectivity characteristics, and the generated Gaussian kernel can adjust its standard deviation according to the category. The farther the target is, the smaller the pixel range it occupies on the map, so the Gaussian kernel will be reduced accordingly. This adaptive method ensures that the confidence map can more realistically reflect the spatial distribution of the target.

[0066] 3.3 Action Relationship Description Input Target Information: Read the original annotation box, target category, and distance information.

[0067] Calculate Gaussian kernel parameters: Determine the standard deviation and amplitude of the Gaussian kernel based on the target category and distance information.

[0068] Generate a single-target heatmap: Generate a two-dimensional Gaussian response map, i.e., a single-target heatmap, at the coordinates of the target's center.

[0069] Overlay all target heatmaps: The heatmaps of individual targets are overlaid onto the entire spectrum to form a complete confidence spectrum. The overlapping areas are normalized to ensure a reasonable overall probability distribution, and the confidence spectrum is output.

[0070] Output confidence map: The generated confidence map is used as a supervision label for training the object detection model, and is used as input for training the high-resolution map after channel expansion.

[0071] 3.4 Module Features and Advantages: Provides a smooth spatial probability distribution, avoiding supervision bias caused by low-resolution bounding boxes.

[0072] It supports category and distance adaptation, which can effectively reflect the characteristics of different targets.

[0073] It is fully compatible with high-resolution maps generated by virtual channel expansion, improving the stability and accuracy of detection model training.

[0074] The generated confidence map can be directly used in CNN, FPN or other deep learning object detection architectures.

[0075] 4. Target Detection and Evaluation Module: This module is designed to train and evaluate target detection on the high-resolution radar RA / RAD map after channel expansion. By using the object detection methods from the papers corresponding to the original datasets (RADDet's corresponding paper is Zhang, A.; Nowruzi, FE; Laganière, R. RADDet: Range–azimuth–doppler based radar objectdetection for dynamic road users. In Proceedings of the IEEE / CVF Conferenceon Computer Vision and Pattern Recognition (CVPR), Nashville, TN, USA, 2021; RADIAL's corresponding paper is Rebut, J.; Ouaknine, A.; Malik, W.; Pérez, P. Raw high-definition radar for multi-task learning. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 2022; pp. 17021–17030.), we can verify that the angular resolution improvement brought about by virtual channel prediction enhances the object detection performance. Experiments show that, under different hardware and corresponding datasets, the extended channel data has significant effects on small-angle interval target detection, sparse target recognition, and background noise suppression.

[0076] This invention possesses significant advantages for engineering application and industrial promotion. Its core feature is that it achieves equivalent array aperture expansion and angular resolution improvement without relying on hardware modifications or increasing the number of sensors, solely through self-supervised learning of existing radar multi-channel signals. Therefore, this invention has broad application prospects and commercial value in multiple fields.

[0077] Firstly, in the fields of autonomous driving and intelligent transportation, millimeter-wave radar has become a core component of environmental perception systems. However, due to limitations in hardware quantity and cost, it generally suffers from insufficient angular resolution. This invention can significantly improve detection performance without replacing the radar or adding new sensors, enabling vehicle-mounted radar to perform better in key tasks such as long-range vehicle recognition, pedestrian lateral motion detection, and multi-target separation, thereby enhancing the safety and stability of autonomous driving systems. This method can be directly deployed on existing pre-installed automotive-grade radar, helping to promote the widespread adoption of high-performance perception in low- and medium-cost vehicle models.

[0078] Secondly, in industrial security and robotics perception scenarios, this invention can enhance the spatial resolution capability of radar in complex environments, enabling radar equipment to have higher target recognition accuracy and scene adaptability. For example, in applications such as indoor robot positioning, security monitoring, and personnel trajectory recognition, higher angular resolution can effectively improve the recognition capability of low-speed or occluded targets, significantly improving system reliability.

[0079] Furthermore, this invention is also highly adaptable to fields such as drones, smart infrastructure, and traffic monitoring. Lightweight radars for drones are often limited by size, preventing an increase in the number of antennas. This invention, however, can achieve near-high-end array radar sensing performance through software. In road infrastructure, radar is used for tasks such as traffic flow statistics and anomaly detection. The improved resolution of this invention can enhance detection accuracy in multi-vehicle parallel scenarios, providing higher-quality data support for intelligent transportation systems.

[0080] Furthermore, this invention uses a self-supervised approach to construct the training system, relying entirely on the radar's own data without requiring bounding box annotation or external sensors, and also without relying on synchronous calibration with visual data. This makes the invention highly versatile and platform compatible, allowing for rapid adaptation to different radar brands and antenna configurations, providing advantages for large-scale deployment, and seamless integration with subsequent deep learning detection frameworks to further amplify the model performance gains brought about by improved angular resolution.

[0081] In summary, this invention not only overcomes the limitations of existing millimeter-wave radar hardware, achieving significant performance improvements at extremely low cost, but also possesses broad cross-industry adaptability and strong engineering feasibility. With the accelerated development of intelligent driving, smart cities, and intelligent robotics industries, this invention will have sustained market demand and extremely high application value, demonstrating promising industrialization prospects.

[0082] In this embodiment, the proposed virtual channel prediction method was evaluated using two publicly available radar datasets, RADDet and RADIAL, which have significantly different characteristics. RADDet represents low-angular-resolution scenes and uses virtual transmit channel prediction, while RADIAL corresponds to high-angular-resolution radar systems and uses virtual receive channel prediction. Although the two datasets differ significantly in angular resolution, radar configuration, prediction channel type, and annotation strategy, the proposed method consistently improves radar-based target detection performance on both datasets, demonstrating strong robustness and generalization ability under different radar configurations.

[0083] As shown in Figure 1, the specific structure of the virtual channel prediction module in this embodiment is as follows: (1) Radar channel signal preprocessing and reassembly submodule Input: The first channel signal in the millimeter-wave radar multi-channel signal is input into the deep convolutional network model. The multi-input multi-output precoding module in the deep convolutional network model is used to reassemble the first channel signal to obtain the channel reassembly features. (2) Multi-scale residual feature extraction submodule Input: Channel reassembly features Feature extraction: Use four residual blocks (Bottleneck or BasicBlock) to construct a feature pyramid: block1 → x1: Low-level features, capturing local amplitude changes and basic channel correlations block2 → x2: Mid-level features, enhancing spatial context block3 → x3: High-level features, capturing target scale and angle patterns block4 → x4: Deepest features, enhancing inter-channel connections and semantic information while retaining x0: Initial low-level features, containing the most original low-level spatial information Output: features = {x0, x1, x2, x3, x4}, providing multi-scale features for the decoder.

[0084] (3) Feature pyramid fusion and angle dimension upsampling submodule Input: Multi-scale features output by the backbone Process: Use upsampling convolution (ConvTranspose2d) and channel splicing (skip connections) to fuse features of different scales layer by layer. Further integrate features through several convolutional blocks Output: Final predicted feature map, the size is usually consistent with or slightly larger than the original angle resolution Function: Map the multi-scale features extracted by the backbone back to the angle-distance space, preparing for the generation of virtual channels. It retains the spatial structure information of the original signal, which is convenient for prediction to be consistent with the real channel. (4) Virtual channel signal regression prediction submodule Input: Final predicted feature map Network structure: Consists of several layers of 3×3 convolution + BN + ReLU stacked to form the last layer of convolution Output the predicted virtual channel, which is consistent with the dimension of the real channel Output: Predicted virtual radar channel Loss function: Self-supervised training: Compare with the corresponding real channel. L1 / L2 / SSIM loss can be used to ensure the consistency of amplitude and spatial structure As shown in Figure 3, the original RADDet annotation, channel-enhanced radar representation and the supervision method based on confidence map are intuitively compared. The top-left subplot shows the original bounding box annotations overlaid on the RAD map obtained from the native 2T4R radar configuration, while the top-right subplot shows color legends corresponding to the six target categories: pedestrians, bicycles, motorcycles, cars, buses, and trucks. The bottom-left subplot shows the original 8-channel RAD representation extracted from the 2T4R measurement data. After virtual channel prediction, the channel-enhanced 12-channel RAD representation corresponding to the virtual 3T4R configuration, as shown in the bottom-middle subplot, exhibits a denser and more continuous angular response. The bottom-right subplot visualizes a Gaussian confidence map constructed according to the CRUW dataset protocol, which provides spatial smoothing supervision aligned with the underlying radar response and is more robust to angular localization uncertainties.

[0085] Figure 4 compares the RAD representations of the original 2T4R and predicted 3T4R data. The top row shows the results obtained with the original 2T4R configuration, while the bottom row corresponds to the 3T4R representation after channel enhancement. From left to right, each row shows the RD plot, RA plot, and RA plot in Cartesian coordinates, respectively. Compared to the original 2T4R representation, the 3T4R results exhibit a significantly narrower main lobe and significantly suppressed sidelobes in the angular dimension, while maintaining good alignment of the target center position. This observation indicates that the predicted virtual channel enhances angular resolution without introducing spatial bias, resulting in a more focused and sharper angular response.

[0086] Table 1. Virtual Channel Prediction Accuracy on RADDet and RADIAL Datasets | Dataset | Channel Prediction | L1 | PSNR (dB) | |---|---|---|---| | RADDet | 1 | T4R (Group 1) → 1 | T4R (Group 2) | 1.38745 | 2.4433 | | RADIAL | 1 | 2 | T1R (Group 1) → 1 | 2 | T1R (Group 2) | 4.76355 | 8.1492 | Table 1 lists the prediction performance on the RADDet and RADIAL datasets under their respective channel prediction settings. For the RADDet dataset, the prediction task is defined as using measurement data from one transmit antenna to predict the measurement data of four receive channels associated with another transmit antenna, i.e., 1T4R (Group 1) → 1T4R (Group 2). Under this setting, the proposed model achieves an L1 error of 1.3874 and a PSNR of 52.4433 dB, demonstrating that the model can accurately predict the amplitude and phase-dependent structures of the signal in the RD domain.

[0087] For the RADIAL dataset, due to channel structure decoupling, virtual channel prediction is performed along the receive antenna dimension. The model predicts twelve transmit channel measurements for one receive antenna based on measurements from another receive antenna, i.e., 12T1R (Group 1) → 12T1R (Group 2). Despite the significantly increased channel dimension and signal complexity, the model still achieves a peak signal-to-noise ratio (PSNR) of 58.1492 dB and an L1 error of 4.7635, demonstrating stable prediction quality even in more challenging scenarios. The metrics reported in Table 1 are not intended as primary measures of the framework's effectiveness. Instead, these metrics confirm that the predicted virtual channels are physically plausible and retain important signal characteristics. The ultimate effectiveness of the virtual channel prediction is evaluated by its impact on downstream target detection performance.

[0088] Table 2. Average Precision (AP) of Raw 2T4R and Predicted 3T4R Configurations in the RADDet Dataset Configuration AP (%) AP_person (%) AP_bicycle (%) AP_car (%) AP_motor (%) AP_bus (%) AP_truck (%) 2T4R (Raw) 54.11 40.71 0.00 68.91 0.00 17.75 30.68 3T4R (Predicted) 59.04 45.55 0.00 73.83 0.00 7.01 38.91 Table 3. Average Recall (AR) under Raw 2T4R and Predictive 3T4R Configurations in the RADDet Dataset | Configuration | AR (%) | AR_person (%) | AR_bicycle (%) | AR_car (%) | AR_motor (%) | AR_bus (%) | AR_truck (%) | 2T4R (Raw) | 63.38 | 47.09 | 0.00 | 79.71 | 0.00 | 40.00 | 40.97 | 3T4R (Predictive) | 65.55 | 50.09 | 0.00 | 80.36 | 0.00 | 43.33 | 49.77 The category analysis shown in Tables 2 and 3 indicates that the performance improvements are primarily concentrated in the main target categories. For people and vehicles, both mean precision (AP) and mean recall (AR) have consistently improved, suggesting that the predicted transport channel measurements have enhanced the spatial representation of these common targets. The truck category also benefits from the enhanced representation, with significant improvements in both precision and recall.

[0089] Table 4. Number of samples for each object category in the RADDet dataset (total 10,194 frames) Category person bicycle motorcycle car busruck Number of objects 650 693 488 1700 421 737 87 Table 4 shows that the RADDet dataset suffers from severe class imbalance. As a result, a few categories, such as bicycles and motorcycles, have extremely few samples in the training set. Therefore, neither the raw nor the augmented data can effectively detect these categories, leading to zero AP and AR scores. The AP for the bus category decreases slightly, but the AR improves. This is mainly because the self-supervised channel prediction model is primarily trained on categories with sufficient samples (people and cars), resulting in uncertainty in predicting bus reflection signals with insufficient samples.

[0090] Table 5. Overall Average Precision (AP) and Average Recall (AR) for Multi-Step Channel Expansion on the RADDet Dataset | Configuration | AP (%) | AR (%) | |---|---|---|---| | 2T4R (Raw) | 54.11 | 63.38 | | 3T4R (Predicted) | 59.04 | 65.55 | | 4T4R (Predicted) | 50.71 | 56.35 Table 5 presents the results of multi-step channel extension. Compared to the single-step 3T4R configuration, the 4T4R configuration leads to a significant decrease in both mean precision and mean recall. This indicates that while single-step extrapolation can effectively improve angular resolution, consecutive prediction stages accumulate phase and amplitude errors, adversely affecting the spatial consistency of the reconstructed RAD representation. These results demonstrate that infinite-channel extrapolation is not feasible, while single-step virtual channel prediction provides a more reliable operating scheme for the RADDet dataset. Object detection performance was evaluated on 1,651 frames of data according to the standard RADinal detection protocol. The quantitative results under different channel configurations are summarized in Table 6: Table 6. Target detection performance configuration on the RADIAL dataset AP (%) AR (%) IoU (%) 12T2R (raw) 63.29 58.75 56.23 12T3R (predicted) 70.59 57.82 57.62 12T3R (raw) 71.63 58.41 57.85 12T4R (predicted) 66.67 53.94 58.03 12T4R (raw) 75.26 59.58 57.21 Compared to the original 12T2R configuration, the predicted 12T3R configuration achieved a significant improvement in average precision (AP), increasing from 63.29% to 70.59%. The intersection-over-union (IoU) ratio also improved from 56.23% to 57.62%, indicating more accurate spatial localization. Although the average recall (AR) decreased slightly, overall detection performance was significantly enhanced.

[0091] For reference, we also report the detection results obtained using actual 12T3R measurement data. The performance gap between the predicted 12T3R configuration and the actual 12T3R data is relatively small, indicating that our proposed virtual channel prediction framework can generate physically meaningful radar channels that are very close to the actual measurements.

[0092] We further investigated multi-step channel extension on the RADIAL dataset, extending the prediction configuration from 12T3R to 12T4R. As shown in Table 10, the predicted 12T4R configuration did not bring further performance improvements and its performance was significantly lower than the actual 12T4R measurement results.

[0093] Although the predicted 12T3R configuration represents a significant improvement over the original 12T2R configuration, the performance degradation of the predicted 12T4R indicates that reliable channel extrapolation is limited to a finite range. Beyond one prediction step, accumulated prediction uncertainty becomes dominant, adversely affecting subsequent detection performance. This observation further supports the choice of single-step virtual channel prediction as a stable and effective operating scheme.

[0094] The working principle of this invention is as follows: In this invention, the original multi-channel radar signal is used as input, and the mapping relationship between millimeter-wave radar channels is learned through self-supervised learning to predict and generate virtual channels that have not yet been acquired, thereby improving the radar angular resolution: Input channel x → Radar channel signal preprocessing and reassembly → Multi-scale residual feature extraction → Feature pyramid fusion and angular dimension upsampling → Virtual channel signal regression prediction → Generation of final virtual channel radar signal.

Claims

1. A method for predicting virtual channels in millimeter-wave radar based on self-supervised learning, characterized in that, The process includes the following steps: Step S1: Input the first channel signal from the millimeter-wave radar multi-channel signal into a deep convolutional network model, and use the multi-input multi-output precoding module in the deep convolutional network model to reconstruct the first channel signal to obtain reconstructed features; Step S2: Input the reconstructed features obtained in Step S1 into the backbone network of the deep convolutional network model for feature extraction to obtain multi-scale features; Step S3: ... The multi-scale features are input into the feature pyramid decoding network of the deep convolutional network model for upsampling. The upsampled high-level features are concatenated with the corresponding low-level features at the channel dimension. The concatenated multi-scale features are fused through convolutional blocks to obtain a radar feature map with high spatial resolution. Step S4: The radar feature map obtained in step S3 is input into the prediction head of the deep convolutional network model. After convolutional operation of the fused features and angular dimension prediction, the predicted virtual second channel signal is output. The prediction head includes multiple two-dimensional convolutional layers, batch normalization, and activation functions. Step S5: Using the second channel signal in the millimeter-wave radar multi-channel signal obtained in step S1 as the supervision target, the virtual second channel signal obtained in step S4 is self-supervised and trained to obtain a virtual channel prediction module. The second channel signal is input into the virtual channel prediction module, and the virtual third channel signal is output through the prediction head of the virtual channel prediction module.

2. The virtual channel prediction method as described in claim 1, characterized in that, Step S1: The first channel signal is a two-dimensional or three-dimensional tensor, which contains at least distance dimension and Doppler dimension information.

3. The virtual channel prediction method as described in claim 1, characterized in that, The multiple input multiple output precoding module in step S1 adopts a two-dimensional convolutional structure with a kernel size of "1, N_Tx" and an expansion rate of "1, N_Rx".

4. The virtual channel prediction method as described in claim 1, characterized in that, Step S2 The backbone network includes an initial convolutional layer and a multi-level residual feature extraction module. The initial convolutional layer uses a 3×3 convolutional kernel, a stride of 1, and padding of 1 for a two-dimensional convolution, and combines batch normalization and ReLU activation function. The multi-level residual feature extraction module is a bottleneck structure, which includes the following layers in sequence: first convolutional layer: 1×1 convolutional kernel, second convolutional layer: 3×3 convolutional kernel with stride of 1 or 2, third convolutional layer: 1×1 convolutional kernel. In the residual connection path, when the number of input channels is inconsistent with the number of output channels, a downsampling operation with a stride of 2 and a 1×1 convolution is introduced.

5. The virtual channel prediction method as described in claim 1, characterized in that, The feature pyramid decoding network described in step S3 includes multiple deconvolutional layers. The kernel size of the deconvolutional layer is 3×3, the stride is 2, the padding is 1, and the output padding is 1.

6. The virtual channel prediction method as described in claim 1, characterized in that, In the prediction head described in step S4, the first convolutional layer uses a 3×3 convolutional kernel, with 256 input channels and 144 output channels. The stride in the angular dimension can be adaptively set to 1 or 2 according to the input size, and the padding is 1. Batch normalization and ReLU activation are applied after the convolution. The second convolutional layer uses a 3×3 convolutional kernel, with 144 input channels and 96 output channels, a stride of 1, and padding of 1. Batch normalization and ReLU activation are applied after the second convolution. Subsequently, two 3×3 convolutional layers are used to maintain the number of channels at 96, further refining the angular dimension features. Finally, a single 3×3 convolution layer outputs the predicted virtual channel signal.

7. The virtual channel prediction method as described in claim 1, characterized in that, In the self-supervised training described in step S5, L1 loss, L2 loss, and / or structural similarity loss are introduced to constrain the consistency of the virtual second channel signal with the second channel signal in terms of amplitude distribution and spatial structure.

8. A high-angle resolution radar map generation method based on the virtual channel prediction method of any one of claims 1 to 7, characterized in that, The process includes the following steps: Step T1: The virtual third channel signal described in step S5 is spliced ​​with other channel signals of the millimeter-wave radar according to the channel dimension to form an extended radar array signal matrix; Step T2: The extended radar array signal matrix described in step T1 is subjected to Fast Fourier Transform (FFT) in the channel dimension to generate a range-angle RA map or a range-angle-Doppler RAD map; Step T3: The range-angle RA map or range-angle-Doppler RAD map described in step T2 is subjected to amplitude normalization and energy consistency correction to obtain a high-angle resolution radar map.

9. A system based on the virtual channel prediction method according to any one of claims 1 to 7, characterized in that, It includes the following modules: Radar channel signal preprocessing and reconstruction module: The first channel signal from the millimeter-wave radar multi-channel signal is input into a deep convolutional network model, and the first channel signal is reconstructed using the multi-input multi-output precoding module in the deep convolutional network model to obtain the reconstructed features; Multi-scale residual feature extraction module: The reconstructed features are input into the backbone network of the deep convolutional network model for feature extraction to obtain multi-scale features; Feature Pyramid Fusion and Angle Dimension Upsampling Module: The multi-scale features are input into the feature pyramid decoding network of the deep convolutional network model for upsampling. The upsampled high-level features are concatenated with the corresponding low-level features at the channel dimension. The concatenated multi-scale features are fused through convolutional blocks to obtain a radar feature map with high spatial resolution. Virtual Channel Signal Regression Prediction Module: The radar feature map is input into the prediction head of the deep convolutional network model. After convolutional operations of the fused features and angle dimension prediction, the predicted virtual second channel signal is output. The prediction head includes multiple two-dimensional convolutional layers, batch normalization, and activation functions. Self-Supervised Training and Loss Constraint Module: Using the second channel signal in the millimeter-wave radar multi-channel signal as the supervision target, the virtual second channel signal is self-supervised to obtain a virtual channel prediction module. The second channel signal is input into the virtual channel prediction module, and the virtual third channel signal is output through the prediction head of the virtual channel prediction module.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the virtual channel prediction method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Radar signal modulation identification method and device for self-supervised contrast mask reconstruction

    CN120687944A