Wall-penetrating radar high-resolution imaging method based on large language model
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING INST OF TECH
- Filing Date
- 2025-06-19
- Publication Date
- 2026-07-21
AI Technical Summary
Traditional through-wall radar imaging algorithms suffer from blurry imaging results due to limitations in antenna aperture and array spacing, making it difficult to directly represent target shape information and thus impossible to directly identify and use.
A high-resolution imaging method based on a large language model is adopted. This method involves constructing a dataset, building and training a high-resolution imaging network that includes an input adapter, a large language model, and an output adapter, freezing the pre-trained parameters, and introducing a low-rank matrix to achieve high-resolution imaging.
It enables the transformation from low-resolution radar images to high-resolution imaging results, allowing for the direct identification and use of target shape and contour information.
Smart Images

Figure CN120802255B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of radar signal processing technology, specifically relating to a high-resolution imaging method for through-wall radar based on a large language model. Background Technology
[0002] With the continuous advancement of urbanization, the demand for concealed space detection has increased significantly, playing an important role in fields such as counter-terrorism, security, rescue, and medical care.
[0003] Through-wall radar (TWR) uses low-frequency electromagnetic waves to penetrate buildings and image targets in obscured space, making it a primary technology for obscured space detection.
[0004] Current traditional through-wall radar imaging algorithms, due to limitations in antenna aperture and array spacing, produce blurry, spot-like images regardless of the target's shape, failing to directly represent target shape information and hindering direct identification and use. Therefore, researching high-resolution through-wall radar methods is a pressing issue in this field. Summary of the Invention
[0005] This invention proposes a high-resolution imaging method for through-wall radar based on Large Language Model (LLM). Addressing the problems of insufficient imaging resolution and difficulty in directly identifying and using imaging results, this method can obtain high-resolution imaging results from low-resolution radar images.
[0006] The technical solution for implementing the present invention is as follows:
[0007] In a first aspect, the present invention provides a high-resolution imaging method for through-wall radar based on a large language model, comprising: Step 1: Obtain the BP radar imaging results and corresponding optical images, and construct a dataset; Simulation imaging results acquisition: assuming an interval A mobile BP radar antenna is used to establish a wall-penetrating echo signal model for the antenna at different locations. Based on the model, echo data imaging results for different targets and different rotation angles of the targets are constructed. Acquisition of measured imaging results: by interval The BP radar antenna is moved to collect echo data from antennas at different positions to obtain radar imaging. The above process is repeated to obtain echo data imaging results for different targets and targets at different rotation angles. Step 2: Build and train a high-resolution imaging network based on a large language model; The network includes an input adapter, a large language model, and an output adapter; a low-rank adaptation module is set within the large language model. When training with the dataset, the query projection matrix and value projection matrix of each Transformer module in the pre-trained large language model are frozen, only the low-rank matrix in the low-rank adaptation module is trained, and the low-rank matrix is added to the query projection matrix and value projection matrix. Step 3: Input the BP radar image to be processed into the imaging network to obtain a high-resolution imaging result containing shape and contour information.
[0008] Optionally, the antenna through-wall echo signal model at different locations described in this invention is derived from the target reflected signal. Direct signal reflection from the wall Reflected signals between targets and noise The superposition of.
[0009] Optionally, this invention sets the radar transmitted signal as a stepped-frequency signal, with a starting frequency of [missing information]. The step frequency interval is ,Include If there are 10 frequency points, then the frequencies within the bandwidth are:
[0010] No. m The antenna position, the first k The received echo signal model for each frequency is represented as follows:
[0011] in, P It is the target quantity. W It is the number of direct reflection paths from the wall. R It is the number of multiple primary reflection paths on the target. It is the size of the imaging scene. , and They are the first w The wall's direct reflection path, the first p The r-th reflection path on the target and the r-th reflection path p, q The reflection coefficient of the path between targets.
[0012] Optionally, the present invention collects... M Echo data from antennas at different locations were analyzed, and IFFT transformation was performed on the echo data. Calculate the first m The antenna position and imaging area h Two-way latency per pixel Then the first h The imaging result of each pixel is: .
[0013] Optionally, the input adapter of the present invention encodes the BP radar image into a series of embedding vectors with the same dimension as the language embedding. The input adapter consists of 6 sequentially connected 3D convolutional layers and 1 linear layer, with each 3D convolutional layer followed by a BatchNorm layer and a LeakyReLU layer.
[0014] Optionally, the query projection matrix in the Transformer module of the large language model described in this invention... Sum projection matrix , ; and These are the query projection matrix and the value projection matrix after adding the low-rank matrix; ,
[0015] in, and It is a dimension of A low-rank matrix, and It is a dimension of A low-rank matrix, Set parameters.
[0016] Optionally, the low-rank matrix described in this invention and Use random initialization, low-rank matrix and Initialize to all zeros.
[0017] Optionally, the loss function during training in this invention for:
[0018] in, It is the MSE loss between the optical image and the generated image. The perceptual loss is calculated using a pre-trained VGG network. The loss is calculated based on cosine similarity. , and To set weights.
[0019] In a second aspect, the present invention provides a high-resolution imaging device for through-wall radar based on a large language model, comprising: The input adapter, large language model, and output adapter are trained using the methods described above. Input adapter for receiving radar images to generate embedding vectors ; Large language model for processing the embedding vectors The process is performed to generate a hidden state vector; Output adapters are used to reconstruct the output of large language models into high-resolution images that include shape and contour information.
[0020] Beneficial effects: This invention proposes a high-resolution imaging method for through-wall radar based on Large Language Model (LLM). A high-resolution imaging network based on LLM is designed, consisting of an input adapter, an output adapter, and an LLM containing a low-rank adaptation module. The input adapter encodes the radar image into embedding vectors that the LLM can process. These embedding vectors are then processed by the LLM and input to the output adapter to obtain the image result. Furthermore, the pre-training parameters of the LLM are frozen during this process, and a trainable low-rank matrix is introduced to reduce computational complexity. The final network can achieve high-resolution imaging of stationary targets, reconstructing the target's shape, contour, and other information, enabling direct identification and use. The network designed in this invention can obtain high-resolution imaging results from low-resolution radar images, reconstructing the target's shape, contour, and other information, allowing for direct identification and use. Attached Figure Description
[0021] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0022] Figure 1 This is a flowchart of an embodiment of the present invention; Figure 2 This is a schematic diagram of the signal scene in this invention; Figure 3 This is a schematic diagram of the network structure proposed in this invention; Figure 4 Here is an example of a simulation dataset of the present invention, wherein (a) is a simulation model diagram, (b) is a dataset label, (c) is a 3D BP radar chart, (d) is a top view, and (e) is a front view; Figure 5 This is a schematic diagram of the actual test scenario of the present invention, wherein (a) is a non-wall-penetrating scenario, (b) is a radar behind the wall in the wall-penetrating scenario, and (c) is a schematic diagram of target placement in the wall-penetrating scenario. Figure 6These are some examples of test results of the present invention on simulation and measured data, wherein (a) is the target imaged in the simulation and measured scenarios, (b) is the truth map, (c) is a three-dimensional map of the radar BP imaging result, (d) is a top view of the radar imaging result, (e) is a front view, (f) is a left view, and (g) is the generation result of the proposed method. Detailed Implementation
[0023] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0024] It should be noted that, in the absence of conflict, the following embodiments and features can be combined with each other; and, based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0025] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement the device and / or practice the method. Additionally, this device and / or method can be implemented using structures and / or functionalities other than one or more of the aspects set forth herein.
[0026] like Figure 1 As shown in the figure, an embodiment of this application provides a high-resolution imaging method for through-wall radar based on a large language model, the specific steps of which include: Step 1: Obtain the BP radar imaging results and corresponding optical images, and construct a dataset; Simulation imaging results acquisition: assuming an interval A mobile BP radar antenna is used to establish a wall-penetrating echo signal model for the antenna at different locations. Based on the model, echo data imaging results for different targets and different rotation angles of the targets are constructed. Acquisition of measured imaging results: by interval The BP radar antenna is moved to collect echo data from antennas at different positions to obtain radar imaging. The above process is repeated to obtain echo data imaging results for different targets and targets at different rotation angles. Step 2: Build and train a high-resolution imaging network based on a large language model; The network includes an input adapter, a large language model, and an output adapter; a low-rank adaptation module is set within the large language model. When training with the dataset, the query projection matrix and value projection matrix of each Transformer module in the pre-trained large language model are frozen, only the low-rank matrix in the low-rank adaptation module is trained, and the query projection matrix and value projection matrix are updated using the low-rank matrix. Step 3: Input the BP radar image to be processed into the imaging network to obtain a high-resolution imaging result containing shape and contour information.
[0027] The specific implementation methods for each of the above steps are described in detail below: Step 1: Obtain the BP radar imaging results and corresponding optical images (i.e., dataset labels) to construct the dataset; the specific process is as follows: Consider as Figure 2 The size shown includes a single reflection. Imaging scene, radar transmission signal It is a step frequency signal, with a starting frequency of The step frequency interval is ,Include If there are 10 frequency points, then the frequencies within the bandwidth are:
[0028] Assuming the wall thickness is The dielectric constant is , to interval Moving the antenna, then the first m The antenna positions are Considering the direct reflection and primary reflection signals from the target, the direct reflection signal from the wall, and the reflection signals between the targets, then the... m The antenna position, the first k The received echo signal at each frequency can be regarded as the target reflection signal. Direct signal reflection from the wall Reflected signals between targets and noise The superposition of can be represented as
[0029] in, P It is the target quantity. W It is the number of direct reflection paths from the wall. R This represents the number of multiple primary reflection paths on the target, where r indicates the reflection path number. , and They are the first w The wall's direct reflection path, the first p The r-th reflection path on the target and the r-th reflection path p, q The reflection coefficient of the path between targets. This represents the received signal obtained after the signal transmitted at the m-th antenna position is reflected by the p-th and q-th targets.
[0030] by Figure 2 Taking the target reflection path Path-A as an example, It is the first m The antenna to the first p Two-way delay for each target. It is the first w The two-way delay of the direct reflection path of a wall can be expressed as:
[0031]
[0032] in, It is the distance that electromagnetic waves travel within the wall. and It is the distance that electromagnetic waves travel in the air. It is the first w The length of the direct reflection path from the wall It is the speed of electromagnetic wave propagation in the wall, according to Snell's law, It can be represented as
[0033] in, It's the speed of light.
[0034] For 3D BP imaging, the imaging region is along xyz The axis is divided into a grid for data collection. M Echo data from antennas at different locations were analyzed, and IFFT transformation was performed on the echo data. Calculate the first m The antenna position and imaging area h Two-way latency per pixel Then the first h The imaging result of each pixel is:
[0035] Calculate the image of the entire region by measuring all pixels.
[0036] like Figure 4-5 As shown, repeat the above steps to collect echo data from different targets at different rotation angles and perform imaging to construct a complete simulation dataset. When constructing the simulation dataset, the formula above... To obtain the data through the above model, we conducted experimental tests to construct a complete experimental dataset. During the construction of the experimental dataset, the formula above... The echo data was obtained by IFFT transformation.
[0037] Step 2: Build and train a high-resolution imaging network based on a large language model This invention uses a deep learning model based on a large language model for high-resolution imaging, consisting of an input adapter, an output adapter, and an LLM containing a low-rank adaptation module. The overall processing flow is as follows: Figure 3 As shown.
[0038] During training, the 3D radar image is first processed by the input adapter and converted into an embedded representation. The embedded representation is then input into the LLM to obtain a vector-form output, which is then processed by the output adapter to obtain a high-resolution 2D image containing reconstructed detailed information. During this process, the pre-trained parameters of the LLM are frozen to conserve computational resources, and low-rank adaptation is introduced to maintain the LLM's learning ability.
[0039] Input Adapter: Since the input 3D radar image is 32×32×32 pixels, which is significantly different from the text input typically used in LLM (Low-Level Mesh) systems, LLM cannot directly process radar images. Therefore, this invention designs an input adapter to encode the 3D radar image into a series of embedding vectors with the same dimensions as the language embeddings. The input adapter consists of six sequentially connected 3D convolutional layers and one linear layer, capable of mapping the 3D radar image into a 4096×1 vector. Each 3D convolutional layer is followed by a BatchNorm layer and a LeakyReLU layer. The BatchNorm layer stabilizes training and improves network convergence, while the LeakyReLU layer introduces non-linearity to prevent gradient vanishing or descent. Through this process, the key features of the input radar image are extracted into the embedding vectors. In Chinese, similar text processing methods, It can be further processed by the Transformer module in LLM.
[0040] Low-rank adaptation: Large language models typically have billions of parameters. Updating all parameters during training would result in significant training time and computational resource consumption. To improve training efficiency and conserve computational resources, this invention introduces low-rank matrices to restrict training to a small number of parameters, i.e., low-rank adaptation (LoRA). LoRA injects trainable low-rank matrices into each layer of the original model and optimizes only these additional low-rank matrices. This invention uses Llama2 as the base large language model and queries the projection matrix in each Transformer module of the LLM. Sum projection matrix With the addition of an extra low-rank matrix, the self-attention mechanism in the Transformer can be represented as follows:
[0041] in, and These are the query matrix and the value matrix, respectively. and These are the query projection matrix and value projection matrix after adding the low-rank matrix. During training, only the following are updated: and Parameters, LLM pre-training parameters It is frozen, which can be represented as: ,
[0042] in, and It is a dimension of A low-rank matrix, and It is a dimension of A low-rank matrix. and Use random initialization and Initialized to all zeros. Because... Only 2.09% of the parameters were updated, thus saving computing resources and improving computing efficiency.
[0043] Output Adapter: After the embedding vector obtained through the input adapter is processed by LLM, the output still has the same dimension as the input embedding and is not the desired two-dimensional image. To obtain effective output results, this invention introduces an output adapter to bridge the gap between the LLM output and the pixel space of a two-dimensional image. Specifically, the output adapter is designed to reconstruct the LLM output into a meaningful and interpretable image. It consists of five transposed convolutional layers, a normalization layer, and a ReLU activation function, used to upsample the LLM output back to a 128×128 image, thereby ensuring the reliability of LLM in image generation.
[0044] Loss function: The loss function during training consists of three parts, which can be expressed as:
[0045] in, It is the MSE loss between the real label and the generated image, to ensure the accuracy and consistency of the generated image and label at the pixel level. The perceptual loss is calculated using a pre-trained VGG network (Visual Geometry Group Network, a perceptual feature extraction network whose inputs are the generated image and labels) to make the generated image more consistent with human visual perception. The loss is calculated based on cosine similarity to ensure that the generated image and the labeled image have overall structural consistency. , and It is a constant that ensures the generated image is accurate in detail, consistent in overall structure, and is also easier for the human eye to recognize directly.
[0046] Step 3: Input the BP radar image into the network to obtain a high-resolution imaging result containing shape and contour information.
[0047] Thus, a high-resolution imaging method for through-wall radar based on a large language model has been completed.
[0048] Example: Based on step 1, the simulated transmitted signal is set to a stepped-frequency signal, with an antenna step of 5cm and an imaging area of 3m×3m×4m. A table with different rotation angles is constructed within the simulation scenario. sim ), chair (Chair sim ), Human sim The dataset was designed to capture the transmitted signal as a stepped-frequency signal with a frequency interval of 2MHz and a frequency range of 1.7GHz-2.2GHz. The antenna array consisted of 10 transmitters and 10 receivers, approximately 40cm x 40cm in size. In a scenario without wall penetration, the distance between the radar and the target was 2m. In a scenario with wall penetration, the wall thickness was 0.2m, with the radar close to one side of the wall and the target 2m away from the other side. Data was collected from Chair-1 at different rotation angles. exp ), Chair-2 exp ), Human exp The dataset.
[0049] According to step 2, set the batch size to 4 and the initial learning rate to 2e-4 for the AdamW optimizer to train the high-resolution imaging network so that it can recover the shape of the target.
[0050] According to step 3, the simulated and measured radar imaging data are input into the trained imaging network to obtain high-resolution imaging results, and cosine similarity (COS Similarity), structural similarity index (SSIM), learned perceptual image patch similarity (LPIPS), intersection-over-union ratio (IoU), and surface miss percentage are used as evaluation indicators.
[0051] Figure 6Examples of imaging results from the network on simulated and measured data are presented. (a) shows the target image in the simulated and measured scenarios; (b) is the ground truth image; (c) is the 3D image of the radar BP imaging result; (d) is the top view of the radar imaging result; (e) is the front view; (f) is the left view; and (g) is the generation result of the proposed method. It can be seen that the 3D image and three-view image of the radar image have low resolution and cannot directly distinguish the shape information of the target type. The proposed method, however, can obtain high-resolution imaging results close to the labeled image, and can reconstruct the shape contour and other detailed information of the imaged target, allowing for direct identification and use by the human eye. Table 1 shows the results of cosine similarity, SSIM, LPIPS, IoU, and surface loss percentage on the test set. It can be seen that all indicators have been improved, demonstrating the effectiveness of the proposed high-resolution imaging method.
[0052] Table 1. Statistical comparison of results on the test set
[0053] In summary, this invention proposes a high-resolution imaging method for through-wall radar based on a large language model. The proposed model consists of an input adapter, an output adapter, and an LLM incorporating low-rank adaptation. Simulation and experimental results demonstrate that the proposed method achieves high-resolution imaging of stationary targets, yielding imaging results that can be directly identified by the human eye.
[0054] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the technical principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. A high-resolution imaging method for through-wall radar based on a large language model, characterized in that, include: Step 1: Obtain the BP radar imaging results and corresponding optical images, and construct a dataset, wherein the imaging results include simulated imaging and measured imaging; Step 2: Build and train a high-resolution imaging network based on a large language model; The network includes an input adapter, a large language model, and an output adapter; a low-rank adaptation module is set within the large language model. When training with the dataset, the query projection matrix and value projection matrix of each Transformer module in the pre-trained large language model are frozen, only the low-rank matrix in the low-rank adaptation module is trained, and the low-rank matrix is added to the query projection matrix and value projection matrix. Step 3: Input the BP radar image to be processed into the imaging network to obtain a high-resolution imaging result containing shape and contour information; The input adapter encodes the BP radar image into an embedding vector with the same dimension as the language embedding. The input adapter consists of six sequentially connected 3D convolutional layers and one linear layer, with each 3D convolutional layer followed by a BatchNorm layer and a LeakyReLU layer. The query projection matrix in the Transformer module of the large language model Sum projection matrix , ; and These are the query projection matrix and the value projection matrix after adding the low-rank matrix; , in, and It is a dimension of A low-rank matrix, and It is a dimension of A low-rank matrix, Set parameters; The output adapter consists of 5 transposed convolutional layers, a normalization layer, and a ReLU activation function.
2. The high-resolution imaging method for through-wall radar based on a large language model according to claim 1, characterized in that, Simulation imaging results acquisition: assuming an interval A mobile BP radar antenna is used to establish a wall-penetrating echo signal model for the antenna at different locations. Based on the model, echo data imaging results for different targets and different rotation angles of the targets are constructed. Acquisition of measured imaging results: by interval The BP radar antenna is moved to collect echo data from antennas at different locations to obtain radar imaging. The above process is repeated to obtain echo data imaging results for different targets and targets with different rotation angles.
3. The high-resolution imaging method for through-wall radar based on a large language model according to claim 1, characterized in that, The antenna through-wall echo signal model at different locations is based on the target reflected signal. Direct signal reflection from the wall Reflected signals between targets and noise The superposition of.
4. The high-resolution imaging method for through-wall radar based on a large language model according to claim 3, characterized in that, Assume the radar transmit signal is a stepped-frequency signal with a starting frequency of . The step frequency interval is ,Include If there are multiple frequency points, then the frequencies within the bandwidth are: No. m The antenna position, the first k The received echo signal model for each frequency is represented as follows: in, P It is the target quantity. W It is the number of direct reflection paths from the wall. R It is the number of multiple primary reflection paths on the target. It is the size of the imaging scene. , and They are the first w The wall's direct reflection path, the first p The r-th reflection path on the target and the r-th reflection path p, q The reflection coefficient of the path between targets.
5. The high-resolution imaging method for through-wall radar based on a large language model according to claim 2, characterized in that, collection M Echo data from antennas at different locations were analyzed, and IFFT transformation was performed on the echo data. Calculate the first m The antenna position and imaging area h Two-way latency per pixel Then the first h The imaging result of each pixel is: 。 6. The high-resolution imaging method for through-wall radar based on a large language model according to claim 1, characterized in that, The low-rank matrix and Use random initialization, low-rank matrix and Initialize to all zeros.
7. The high-resolution imaging method for through-wall radar based on a large language model according to claim 1, characterized in that, Loss function during training for: in, It is the MSE loss between the optical image and the generated image. The perceptual loss is calculated using a pre-trained VGG network. The loss is calculated based on cosine similarity. , and To set weights.
8. A high-resolution imaging device for through-wall radar based on a large language model, characterized in that, include: The input adapter, the large language model, and the output adapter are obtained by training using any one of the claims 1-7; Input adapter for receiving radar images to generate embedding vectors ; Large language model for processing the embedding vectors The process is performed to generate a hidden state vector; Output adapters are used to reconstruct the output of large language models into high-resolution images that include shape and contour information.