Near-field imaging clutter suppression method and device based on comparative learning

By using a near-field imaging clutter suppression method based on contrastive learning, a class-independent activation map is generated and clutter is suppressed. This solves the problem of inaccurate signal extraction in complex environments in traditional millimeter-wave near-field imaging, improves imaging quality and target detection capabilities, and enhances the adaptability and robustness of the radar system.

CN120871038AActive Publication Date: 2025-10-31HANGZHOU BOSER INFORMATION TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202511387875.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-26
Publication Date
2025-10-31
Estimated Expiration
2045-09-26

AI Technical Summary

Technical Problem

Traditional millimeter-wave near-field imaging methods struggle to accurately extract effective signals and features in complex environments, resulting in decreased clarity and contrast in the imaging results. This affects target detection and recognition, limiting their performance and reliability in practical applications.

Method used

A near-field imaging clutter suppression method based on contrastive learning is adopted. By acquiring echo data from multiple planes of the target, a three-dimensional imaging result is generated. A class-independent activation map is generated using a pre-trained near-field imaging clutter suppression network. Clutter suppression is performed based on the mask matrix, while preserving the target echo data.

Benefits of technology

In the absence of precise human body contour annotation, it effectively suppresses background clutter, enhances the target foreground signal, improves the imaging quality and target detection capability of millimeter-wave near-field imaging, and enhances the adaptability and robustness of radar system equipment in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120871038A_ABST
    Figure CN120871038A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of millimeter wave imaging, in particular to a near-field imaging clutter suppression method and device based on comparative learning, and the method comprises the steps: collecting the complete echo data of a target body in a plurality of planes parallel to a radar array plane, so as to obtain the imaging results of the target body in the plurality of planes, determining a three-dimensional imaging result of the target body and layered imaging data of the target body at different distances; inputting a hierarchical maximum projection result of the hierarchical imaging data into a pre-trained near-field imaging clutter suppression network based on comparative learning to output a category-independent activation graph of the foreground; and respectively executing clutter suppression processing on each layer of the layered imaging data based on a mask matrix generated by the category-independent activation graph so as to suppress clutter data in the layered imaging data and reserve effective target data. According to the invention, foreground signals and complex background clutters can be effectively distinguished, accurate separation of the foreground and the background is realized, and the expressive force and practicability of imaging are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of millimeter-wave imaging technology, and in particular to a near-field imaging clutter suppression method and apparatus based on contrastive learning. Background Technology

[0002] Among related technologies, millimeter-wave imaging technology, as a non-contact and harmless method of human body detection, has been widely used in security inspection, medical and other fields in recent years. Millimeter waves have good penetrating ability and can effectively penetrate common opaque media such as clothing and packages to detect hidden objects. At the same time, they are harmless to the human body, thus becoming an important technology for security inspections in places such as airports and subways.

[0003] However, in related technologies, traditional millimeter-wave near-field imaging methods are difficult to accurately extract effective signals and features in complex environments with strong background clutter interference, which often obscures or interferes with human target signals. Furthermore, since the feature representation of human targets is relatively weak, especially in diverse environments and under complex object occlusion, the clarity and contrast of the imaging results decrease, affecting the subsequent target detection and recognition effects. This limits the performance and reliability of millimeter-wave near-field imaging in practical applications, and urgently needs to be addressed. Summary of the Invention

[0004] This application provides a near-field imaging clutter suppression method and apparatus based on contrastive learning to address the shortcomings of traditional millimeter-wave near-field imaging methods in extracting effective signals and features in complex environments. Furthermore, due to the relatively weak feature representation of human targets, the clarity and contrast of imaging results in diverse environments and complex object occlusion scenarios are reduced, seriously affecting subsequent target detection and recognition effects and limiting the performance and reliability of millimeter-wave near-field imaging in practical applications.

[0005] The first aspect of this application provides a near-field imaging clutter suppression method based on contrastive learning, comprising the following steps: acquiring complete echo data of a target object on multiple planes parallel to the radar array plane to obtain imaging results of the target object on the multiple planes, determining the three-dimensional imaging result of the target object based on the imaging results of the multiple planes, and determining layered imaging data of the target object at different distances from the radar array based on the three-dimensional imaging results; inputting the layer maximum projection result corresponding to the layered imaging data into a pre-trained near-field imaging clutter suppression network based on contrastive learning to output a class-independent activation map of the foreground corresponding to the target object; generating a mask matrix based on the class-independent activation map, and performing clutter suppression processing on each layer of the layered imaging data according to the mask matrix to suppress other clutter data in the layered imaging data except for the target echo data, while retaining the target echo data in the layered imaging data.

[0006] Optionally, in one embodiment of this application, before inputting the layered maximum projection result corresponding to the layered imaging data into the pre-trained near-field imaging clutter suppression network based on contrastive learning, the method further includes: projecting the layered imaging data sequentially onto multiple color channels based on the multiple planes to generate a multi-channel two-dimensional image; and determining the layered maximum projection result based on the multi-channel two-dimensional image.

[0007] Optionally, in one embodiment of this application, the expression for the projection result of the layered maximum value is: , , , , in, This represents the projection result of the maximum value in the layer; The matrix represents the concatenation of elements along the channel dimension; Represents the set of natural numbers; This indicates the projection processing of the maximum value of the layer; Transitional parameters designed to facilitate the calculation process; Indicates the layer number where the layered data is located. RGB three channels, Indicates the total number of floors. This indicates that the index value is contained in the set along the distance direction. All tangent planes In the middle, select the maximum value among pixels with the same coordinates; This indicates the projection processing of the maximum value; These represent the number of layers, height, and width of the layered imaging, respectively. This represents layered imaging data.

[0008] Optionally, in one embodiment of this application, before acquiring the complete echo data of the target body on the plurality of planes parallel to the radar array plane, the method further includes: acquiring the original echo data of the plurality of typical bodies based on the radar array; obtaining the correction data of the radar array based on the original echo data; and performing matrix multiplication correction processing on the correction data and the original echo data of the target body to obtain the complete echo data.

[0009] Optionally, in one embodiment of this application, before inputting the projection result of the layered maximum value corresponding to the layered imaging data into the pre-trained near-field imaging clutter suppression network based on contrastive learning, the method further includes: multiplying the compressed vectors corresponding to the class-independent activation maps of the plurality of typical objects with the compressed vectors corresponding to the small target detection feature maps of the plurality of typical objects to obtain complementary foreground vectors and background vectors; constructing negative example pair loss, foreground positive example pair loss, and background positive example pair loss based on the foreground vectors and the background vectors; and constructing a total loss function for training the pre-trained near-field imaging clutter suppression network based on contrastive learning by combining the negative example pair loss, foreground positive example pair loss, and background positive example pair loss, so as to construct the pre-trained near-field imaging clutter suppression network based on contrastive learning using the total loss function.

[0010] Optionally, in one embodiment of this application, the expression for the total loss function is: , , , , in, Indicates the total loss. Indicates detection loss, Indicates the weight of the comparative loss. Indicates a negative example of loss. This indicates the positive example relative to the total loss. , These represent the foreground positive example loss and the background positive example loss, respectively. Indicates batch size, Indicates the first One foreground vector, Indicates the first One foreground vector, Indicates the first Background vectors, Indicates the first Background vectors.

[0011] A second aspect of this application provides a near-field imaging clutter suppression device based on contrastive learning, comprising: a conversion module, configured to acquire complete echo data of a target object on multiple planes parallel to the radar array plane, to obtain imaging results of the target object on the multiple planes, to determine a three-dimensional imaging result of the target object based on the imaging results of the multiple planes, and to determine layered imaging data of the target object at different distances from the radar array based on the three-dimensional imaging results; a first processing module, configured to input the layer maximum projection result corresponding to the layered imaging data into a pre-trained near-field imaging clutter suppression network based on contrastive learning, to output a class-independent activation map of the foreground corresponding to the target object; and a suppression module, configured to generate a mask matrix based on the class-independent activation map, and to perform clutter suppression processing on each layer of the layered imaging data according to the mask matrix, to suppress other clutter data in the layered imaging data except for the target echo data, and retain the target echo data in the layered imaging data.

[0012] Optionally, in one embodiment of this application, it further includes: a projection module, configured to project the layered imaging data sequentially onto multiple color channels based on the multiple planes to generate a multi-channel two-dimensional image before inputting the layered maximum projection result corresponding to the layered imaging data into the pre-trained near-field imaging clutter suppression network based on contrastive learning; and a determination module, configured to determine the layered maximum projection result based on the multi-channel two-dimensional image.

[0013] Optionally, in one embodiment of this application, the expression for the projection result of the layered maximum value is: , , , , in, This represents the projection result of the maximum value in the layer; The matrix represents the concatenation of elements along the channel dimension; Represents the set of natural numbers; This indicates the projection processing of the maximum value of the layer; Transitional parameters designed to facilitate the calculation process; Indicates the layer number where the layered data is located. They are three channels, RGB. Indicates the total number of floors. This indicates that the index value is contained in the set along the distance direction. All tangent planes In the middle, select the maximum value among pixels with the same coordinates; This indicates the projection processing of the maximum value; These represent the number of layers, height, and width of the layered imaging, respectively. This represents layered imaging data.

[0014] Optionally, in one embodiment of this application, it further includes: a data acquisition module, configured to acquire raw echo data of multiple typical objects based on the radar array before acquiring the complete echo data of the target object on the plurality of planes parallel to the radar array plane; an acquisition module, configured to acquire correction data of the radar array based on the raw echo data; and a correction module, configured to perform matrix multiplication correction processing on the correction data and the raw echo data of the target object to obtain the complete echo data.

[0015] Optionally, in one embodiment of this application, it further includes: a second processing module, configured to multiply the compressed vectors corresponding to the class-independent activation maps of the plurality of typical objects with the compressed vectors corresponding to the small target detection feature maps of the plurality of typical objects before inputting the projection result of the layered maximum value corresponding to the layered imaging data into the pre-trained near-field imaging clutter suppression network based on contrastive learning, so as to obtain complementary foreground vectors and background vectors; a first construction module, configured to construct negative example pair loss, foreground positive example pair loss and background positive example pair loss based on the foreground vector and the background vector; and a second construction module, configured to combine the negative example pair loss, foreground positive example pair loss and background positive example pair loss to construct a total loss function for training the pre-trained near-field imaging clutter suppression network based on contrastive learning, so as to construct the pre-trained near-field imaging clutter suppression network based on contrastive learning using the total loss function.

[0016] Optionally, in one embodiment of this application, the expression for the total loss function is: , , , , in, Indicates the total loss. Indicates detection loss, Indicates the weight of the comparative loss. Indicates a negative example of loss. This indicates the positive example relative to the total loss. , These represent the foreground positive example loss and the background positive example loss, respectively. Indicates batch size, Indicates the first One foreground vector, Indicates the first One foreground vector, Indicates the first Background vectors, Indicates the first Background vectors.

[0017] A third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the near-field imaging clutter suppression method based on contrastive learning as described in the above embodiments.

[0018] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described near-field imaging clutter suppression method based on contrastive learning.

[0019] A fifth aspect of this application provides a computer program product, including a computer program that, when executed, is used to implement the above-described contrastive learning-based near-field imaging clutter suppression method.

[0020] This application embodiment can input the projection result of the maximum value of each layer corresponding to the layered imaging data of the target into a pre-trained near-field imaging clutter suppression network based on contrastive learning. This network then generates a class-independent activation map and its mask matrix for the target, performing clutter suppression processing on each layer of the layered imaging data to suppress clutter signals in the complete echo data. Thus, even in the absence of precise human contour annotations, by fusing layered imaging processing with deep learning technology, constructing a class-independent activation map, and designing a contrastive learning loss, the ability to distinguish between human foreground signals and background clutter signals is improved. This effectively suppresses background clutter, enhances the target foreground signal, and accurately extracts effective signals and features, thereby completing the clutter suppression task on the basis of the detection process. This improves the imaging quality and target detection capability of millimeter-wave near-field imaging, and also enhances the adaptability and robustness of radar systems in complex environments, providing a solid technical foundation for the further application of millimeter-wave imaging technology in security, medical, and other fields. This solves the problems in related technologies, such as the difficulty in accurately extracting effective signals and features in complex environments with strong background clutter interference that often obscures or interferes with human target signals; and the fact that human target features are relatively weak, especially in diverse environments and under complex object occlusion, the clarity and contrast of the imaging results decrease, affecting the subsequent target detection and recognition effects, and limiting the performance and reliability of millimeter-wave near-field imaging in practical applications.

[0021] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description

[0022] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart of a near-field imaging clutter suppression method based on contrastive learning provided in an embodiment of this application; Figure 2 This is a schematic diagram of the structure of a clutter suppression module according to an embodiment of this application; Figure 3 This is an overall structural diagram of a near-field imaging clutter suppression network based on contrastive learning, according to an embodiment of this application. Figure 4 This is a schematic diagram of a clutter suppression process according to an embodiment of this application; Figure 5 This is a flowchart illustrating a near-field imaging clutter suppression method based on contrastive learning, according to one embodiment of this application. Figure 6This is a diagram illustrating the near-field imaging clutter suppression effect of one embodiment of this application; Figure 7 This is a schematic diagram of the structure of a near-field imaging clutter suppression device based on contrastive learning provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of this application.

[0023] Figure label: 10-Near-field imaging clutter suppression device based on contrastive learning; 100-Conversion module, 200-First processing module and 300-Suppression module; 801-Memory, 802-Processor and 803-Communication interface. Detailed Implementation

[0024] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.

[0025] The present application provides a near-field imaging clutter suppression method and apparatus based on contrastive learning, with reference to the accompanying drawings. Addressing the problems mentioned in the background section, the present application provides a near-field imaging clutter suppression method based on contrastive learning. In this method, the projection result of the maximum layer value corresponding to the layered imaging data of the target object is input into a pre-trained near-field imaging clutter suppression network based on contrastive learning. This network then generates a class-independent activation map of the target object and its mask matrix to perform clutter suppression processing on each layer of the layered imaging data, thereby suppressing clutter signals in the complete echo data. This breakthrough achieves improved discrimination between human foreground signals and background clutter in the absence of precise human contour annotations. By integrating layered imaging processing and deep learning techniques, constructing class-independent activation maps, and designing contrastive learning loss, it effectively suppresses background clutter, enhances target foreground signals, and accurately extracts effective signals and features. This allows for clutter suppression during the detection process, improving the imaging quality and target detection capabilities of millimeter-wave near-field imaging. It also enhances the adaptability and robustness of radar systems in complex environments, providing a solid technical foundation for further applications of millimeter-wave imaging technology in security, medical, and other fields. This breakthrough addresses the challenges of traditional millimeter-wave near-field imaging methods in complex environments with strong background clutter that often obscures or interferes with human target signals. Furthermore, the relatively weak feature representation of human targets, especially in diverse environments and with complex object occlusion, leads to decreased image clarity and contrast, affecting subsequent target detection and recognition, and limiting the performance and reliability of millimeter-wave near-field imaging in practical applications.

[0026] Before explaining the near-field imaging clutter suppression method based on contrastive learning in the embodiments of this application, the formula parameters involved in the embodiments of this application will be explained first. Table 1 is a table of formula parameter definitions for one embodiment of this application, which can be, but is not limited to, represented as follows: Table 1

[0027] Specifically, Figure 1 This is a flowchart illustrating a near-field imaging clutter suppression method based on contrastive learning, provided as an embodiment of this application.

[0028] like Figure 1 As shown, this near-field imaging clutter suppression method based on contrastive learning includes the following steps: In step S101, complete echo data of the target body in multiple planes parallel to the radar array plane are acquired to obtain the imaging results of the target body in multiple planes. The three-dimensional imaging results of the target body are determined based on the imaging results of multiple planes, and the layered imaging data of the target body at different distances from the radar array are determined based on the three-dimensional imaging results.

[0029] It can be understood that the multiple planes parallel to the radar array plane can be understood here as multiple "range layers" divided according to the range resolution in the three-dimensional imaging space of the radar array, and the planes (lateral planes) parallel to the radar array on each range layer.

[0030] Range resolution refers to the minimum distance at which a radar array can distinguish two adjacent targets in the range direction (i.e., the radial direction between the radar and the target). Higher resolution means more detailed objects can be distinguished. For example, a range resolution of 1 meter means the radar can distinguish two targets within 1 meter.

[0031] The target object here can be understood as the target scanned by the radar array. For example, in millimeter-wave body scanning, the human body is the target object.

[0032] In some embodiments, this application can acquire complete echo data of the target body on multiple planes parallel to the radar array plane to obtain the imaging results of the target body on multiple planes, and use the imaging results of each plane as the three-dimensional imaging results of the target body at different distances from the radar array.

[0033] To facilitate subsequent image processing, embodiments of this application can further perform operations such as amplitude preservation, range compression, and quantization on the three-dimensional imaging results to convert the three-dimensional imaging results into layered imaging data of the target object at different distances from the radar array. Here, layered imaging data can be understood as image data acquired and recorded in layers according to different levels.

[0034] For example, this application can set several distances in the imaging space according to the range resolution of the radar system, and use the range migration (RM) algorithm to calculate the imaging results at each distance relative to the plane parallel to the radar array, thus forming a three-dimensional imaging result. The discrete-condition RM algorithm can be expressed as:

[0035] in, This represents the imaging process of the RM algorithm; and These represent the two-dimensional forward Fourier transform and the three-dimensional inverse Fourier transform, respectively. This represents a one-dimensional interpolation operation in the spatial wavenumber dimension, which interpolates non-uniform data onto a uniform grid to facilitate subsequent Fourier transform operations. Represents the space wavenumber domain and Directional wave number and The spherical function is expressed as follows: ; This represents the reference distance from the entire imaging area to the array. These represent the coordinates of the transceiver elements in the planar array. Indicates the broadband wavenumber. This represents the raw echo data. It represents a complex exponential term.

[0036] Among them, three-dimensional imaging results Since it is in complex form, the embodiments of this application can be further processed by preserving amplitude information. Compressing numerical range operations and 8-bit quantization operation And by adjusting the data shape operation The number of layers, height, and width are obtained as follows: 、 and Layered imaging data as follows: .

[0037] It should be noted that the operation of retaining amplitude information... Compressing numerical range operations Quantitative Operation and operations to adjust data shape These operations are optional in practical applications. That is, in practical applications, those skilled in the art can choose not to perform these operations or only perform some of them. When no operations are performed, the three-dimensional imaging results can be directly used as layered imaging data. The embodiments in this application are for illustrative purposes only and do not constitute specific limitations.

[0038] The embodiments of this application can generate a three-dimensional imaging result of the target object based on the imaging results of the target object on multiple planes parallel to the radar array plane, and convert the three-dimensional imaging result into layered imaging data through operations that preserve amplitude information, compress numerical range, and quantize. This preserves the key features of the target object in the three-dimensional imaging as much as possible, converts the original massive floating-point data into a compact integer array, and can still clearly distinguish the amplitude difference between the target object and the background.

[0039] Optionally, in one embodiment of this application, before acquiring complete echo data of the target body on multiple planes parallel to the radar array plane, the method further includes: acquiring raw echo data of multiple typical bodies based on the radar array; obtaining correction data of the radar array based on the raw echo data; and performing matrix multiplication correction processing on the correction data and the raw echo data of the target body to obtain complete echo data.

[0040] In some embodiments, different radar devices may produce different errors when scanning targets due to differences in hardware, etc. In order to obtain more accurate echo data, this application can first obtain the correction data of the radar array, and then correct the actual echo data of the target based on the correction data to obtain the complete echo data of the target.

[0041] Specifically, in this application embodiment, the original echo data of multiple typical objects can be collected first through a radar array. Based on these original echo data, the correction data of the radar array can be obtained under certain conditions. Then, the obtained correction data and the original echo data of the target object are subjected to matrix multiplication correction processing to obtain the complete echo data of the target object.

[0042] Here, "typical object" refers to a target such as an object or human body scanned by the radar equipment. It is sufficient to collect its original echo data. This embodiment is only for illustrative purposes and does not impose any specific limitations.

[0043] For example, this application can collect raw echo data from objects and human bodies in a millimeter-wave human body security scanner (radar array device) that is actually in operation, and then obtain correction data based on the differences in radar system equipment (used to correct the deviations caused by the radar equipment itself) under strict conditions such as in a laboratory by professionals in this technical field.

[0044] Then, after acquiring the original echo data of the target body obtained by the millimeter-wave human body security scanner, the embodiments of this application can perform matrix multiplication correction on the correction data and the original echo data of the target body to obtain the complete three-dimensional echo data of the target body.

[0045] It should be noted that, in order to ensure the accuracy of the calibration results, it is recommended that the radar array (radar system equipment) scanning the target and the radar arrays of multiple typical targets be kept the same.

[0046] Among them, matrix multiplication correction refers to multiplying the correction data with the original echo data point by point through the mathematical operation of matrix multiplication, so as to eliminate the interference of radar equipment differences on the original echo signal (such as unifying the signal consistency of different radar channels).

[0047] The embodiments of this application can acquire the correction data of the radar array, and use the correction data to perform matrix multiplication with the actual original echo data of the target object collected by the radar array, thereby obtaining complete three-dimensional echo data that truly reflects the target object, providing a reliable data foundation for subsequent image processing.

[0048] Step S102: Input the projection result of the maximum value of the layer corresponding to the layered imaging data into the pre-trained near-field imaging clutter suppression network based on contrastive learning to output the category-independent activation map of the foreground corresponding to the target.

[0049] In other embodiments, after acquiring the layered imaging data, this application can input the projection result of the layered maximum value corresponding to the layered imaging data into a pre-trained near-field imaging clutter suppression network based on contrastive learning to obtain a category-independent activation map of the foreground corresponding to the target.

[0050] In this context, the layered maximum projection result can be understood as an improvement upon the maximum projection result, visualizing the echo data of each distance plane in the 3D data. Maximum projection is a commonly used 3D data visualization method in millimeter-wave imaging.

[0051] The category-independent activation map of the foreground corresponding to the target can be understood here as the category-independent activation map of the foreground where the target is located when the radar array scans the target.

[0052] The following section provides a further explanation of the layered maximum projection results and the pre-trained near-field imaging clutter suppression network based on contrastive learning in this application.

[0053] Optionally, in one embodiment of this application, before inputting the layered maximum projection result corresponding to the layered imaging data into a pre-trained near-field imaging clutter suppression network based on contrastive learning, the method further includes: projecting the layered imaging data sequentially onto multiple color channels based on multiple planes to generate a multi-channel two-dimensional image; and determining the layered maximum projection result based on the multi-channel two-dimensional image. The expression for the layered maximum projection result may, but is not limited to, be: , , , , in, This represents the projection result of the maximum value in the layer; The matrix represents the concatenation of elements along the channel dimension; Represents the set of natural numbers; This indicates the projection processing of the maximum value of the layer; Transitional parameters designed to facilitate the calculation process; Indicates the layer number where the layered data is located. RGB three channels, Indicates the total number of floors. This indicates that the index value is contained in the set along the distance direction. All tangent planes In the middle, select the maximum value among pixels with the same coordinates; This indicates the projection processing of the maximum value; These represent the number of layers, height, and width of the layered imaging, respectively. This represents layered imaging data.

[0054] In actual implementation, when obtaining the maximum layer projection result from the layered imaging data, this application may, but is not limited to, project the layered imaging data sequentially onto multiple color channels according to multiple planes parallel to the radar array plane to generate a multi-channel two-dimensional image, and then determine the maximum layer projection result based on the multi-channel two-dimensional image.

[0055] In this context, "multi-channel" refers to multiple color channels, typically the three primary colors of RGB: red, green, and blue. In practical applications, those skilled in the art can select and adjust the channels according to the specific circumstances. This embodiment is merely illustrative and does not impose any specific limitations.

[0056] Understandably, maximum projection is a simple and feasible method for visualizing 3D data and is widely used in millimeter-wave imaging. In traditional millimeter-wave imaging methods, the high-dimensionality of 3D data increases the complexity of data storage and processing. Directly processing 3D data is not only computationally intensive but also makes it difficult to fully utilize the spatial structural information of the target. Conventional maximum projection methods are prone to losing spatial details, affecting the effectiveness of subsequent feature extraction and clutter suppression.

[0057] Based on this, this application designs a two-dimensional image conversion scheme that combines layered imaging and layered maximum projection: by setting multiple distance planes in the imaging space, the three-dimensional imaging data is grouped and mapped to the three color channels of the image respectively, thereby achieving effective compression and expression of three-dimensional spatial information and generating a multi-channel two-dimensional image. This significantly reduces the complexity of data processing while providing a high-quality two-dimensional image foundation for subsequent feature extraction.

[0058] Specifically, the embodiments of this application can first utilize maximum value projection. By preserving the maximum modulus value in the distance dimension, three-dimensional data is compressed into one dimension, thereby generating a single-channel two-dimensional image. This process can be represented, but is not limited to, as follows: , , in, This indicates that the index value is contained in the set along the distance direction. All tangent planes In the selection process, the maximum value among pixels with the same coordinates is chosen.

[0059] Layered maximum projection This can be understood as a certain improvement to the maximum value projection, which projects each distance plane of the three-dimensional data onto the three color channels (RGB three primary colors: red, green, and blue) in turn to generate a three-channel two-dimensional image. This three-channel two-dimensional image is the result of the layered maximum value projection.

[0060] Compared to maximum projection, layered maximum projection can retain more spatial information while taking advantage of image format convenience. Its mathematical expression can be, but is not limited to, expressed as:

[0061]

[0062] in, This represents the projection result of the maximum value in the layer; The matrix represents the concatenation of elements along the channel dimension; Represents the set of natural numbers. Indicates the layer number where the layered data is located. RGB three channels, Indicates the total number of floors.

[0063] Furthermore, embodiments of this application can project the layered maximum value results. Input into a YOLO v8-based object detection network (e.g., object detection networks based on YOLO v8), projecting the results using hierarchical maximum values ​​(three-channel two-dimensional images). To detect targets, feature maps for small target detection are extracted using a multi-head mechanism in the target detection network. Then Input to category-independent activation graph generation network This will allow you to obtain a category-independent activation graph. Its mathematical expression can be, but is not limited to, the following:

[0064] Among them, the category-independent activation graph generation network in the embodiments of this application It can be, but is not limited to, composed of a single 2D convolutional layer (Conv 2D), a single 2D batch normalization layer (BN 2D), and a single sigmoid activation function layer, such as... Figure 2 As shown, Figure 2 The structure of a category-independent activation graph generation network according to one embodiment of this application is shown.

[0065] Optionally, in one embodiment of this application, before inputting the layer maximum projection result corresponding to the layered imaging data into the pre-trained near-field imaging clutter suppression network based on contrastive learning, the method further includes: multiplying the compressed vectors corresponding to the class-independent activation maps of multiple typical objects with the compressed vectors corresponding to the small target detection feature maps of multiple typical objects to obtain complementary foreground and background vectors; constructing negative example pair loss, foreground positive example pair loss, and background positive example pair loss based on the foreground and background vectors; and constructing a total loss function for training the pre-trained near-field imaging clutter suppression network based on contrastive learning by combining the negative example pair loss, foreground positive example pair loss, and background positive example pair loss, so as to construct the pre-trained near-field imaging clutter suppression network based on contrastive learning using the total loss function. The expression of the total loss function may, but is not limited to, be: , , , , in, Indicates the total loss. Indicates detection loss, Indicates the weight of the comparative loss. Indicates a negative example of loss. This indicates the positive example relative to the total loss. , These represent the foreground positive example loss and the background positive example loss, respectively. Indicates batch size, Indicates the first One foreground vector, Indicates the first One foreground vector, Indicates the first Background vectors, Indicates the first Background vectors.

[0066] Based on the descriptions of other embodiments, it will be understood that this application generates and outputs a category-independent activation map of the target object by inputting the projection result of the maximum layer value corresponding to the layered imaging data into a pre-trained near-field imaging clutter suppression network based on contrastive learning.

[0067] In some embodiments, the present application also implements the pre-trained near-field imaging clutter suppression network based on contrastive learning using class-independent activation maps.

[0068] In complex scenarios, millimeter-wave imaging targets are often mixed with background clutter, making it difficult for traditional feature extraction methods to accurately separate the foreground from the background, leading to decreased detection accuracy. While some target detection networks can extract certain spatial features, they struggle to effectively detect human contours in imaging images without labels or prior information. Furthermore, category-related feature extraction methods are easily affected by scene changes, reducing the model's generalization ability.

[0069] Therefore, this application designs a method for extracting human signals in millimeter-wave near-field imaging that combines category-independent activation map generation with a contrastive learning mechanism, and constructs a near-field imaging clutter suppression network based on this method. This method (the near-field imaging clutter suppression network based on contrastive learning) constructs foreground and background vectors using small target detection features and the generated activation map, and introduces positive and negative sample contrast loss to improve the network's sensitivity to human regions under label-deficient conditions, while simultaneously enhancing the model's ability to discriminate background clutter.

[0070] Specifically, the embodiments of this application can obtain complete echo data of multiple typical objects based on the original echo data of multiple typical objects on multiple planes parallel to the radar array plane, based on the same processing steps as the target object. Then, the imaging results of multiple typical objects on multiple planes are obtained based on the complete echo data of multiple typical objects, so as to generate three-dimensional imaging results of multiple typical objects based on the imaging results of multiple typical objects on multiple planes. Similarly, the three-dimensional imaging results of multiple typical objects are transformed into layered imaging data of multiple typical objects through operations of preserving amplitude information, compressing numerical range, and quantization.

[0071] Next, the layered imaging data of multiple typical objects are sequentially projected onto multiple color channels to obtain the layered maximum projection results corresponding to multiple typical objects. Then, the category-independent activation maps of multiple typical objects are obtained based on the layered maximum projection results corresponding to multiple typical objects.

[0072] Subsequently, embodiments of this application can combine category-independent activation maps of multiple typical entities with small target detection feature maps of multiple typical entities. After compressing into vectors, multiplying them yields complementary foreground vectors. With background vector In this context, the foreground vector can be understood as the echo signal that does not need to be suppressed, while the background vector is the echo signal that needs to be suppressed, thereby achieving accurate separation between the foreground and the background.

[0073] Taking radar equipment as a millimeter-wave human body security scanner, with the human body as a typical example, the foreground vector can be understood as the human body signal, and the background vector can be understood as the background clutter of the human body, etc.

[0074] As attached Figure 2 As shown, foreground vector With background vector The mathematical expressions for can be, but are not limited to, represented as follows: , in, This operation represents flattening an image into a vector; Indicates matrix transpose; It represents a vector that is all one.

[0075] Furthermore, to quantify the degree of separation between the foreground and background signals, for a batch size of... Based on the training data, this embodiment of the application can extract foreground vector groups. and background vector group Using the negative logarithmic relationship and the cosine distance of vectors, a negative pair loss is established. The mathematical expression can be, but is not limited to, represented as follows: , Similarly, to evaluate the similarity between foreground-foreground feature pairs and background-background feature pairs, a foreground positive pair loss is constructed. and background positive examples against loss And synthesize them into total positive examples against the loss. Its mathematical expression can be, but is not limited to, the following: , , Next, the embodiments of this application will be described in detail. As a contrastive learning loss and with The detection loss obtained by adding its coefficients to the detection network In this process, the total loss generated during network training is determined. Its mathematical expression can be, but is not limited to, the following:

[0076] in, This represents the weight of the contrastive learning loss.

[0077] By using the above loss to iteratively train and update the model parameters of the basic neural network, a near-field imaging clutter suppression network based on contrastive learning can be obtained. This improves the network's sensitivity to human regions in the absence of labels and enhances the model's ability to discriminate background clutter.

[0078] The basic neural network can be, but is not limited to, a fully connected neural network, a convolutional neural network, or a recurrent neural network, etc., as long as it can achieve the corresponding network function in the embodiments of this application. The embodiments of this application are only illustrative examples and are not specifically limited.

[0079] like Figure 3 As shown, Figure 3 This is an overall structural diagram of a near-field imaging clutter suppression network based on contrastive learning, according to one embodiment of this application. Layered imaging data is first processed by Layered Maximum Intensity Projection (LMIP) to generate a three-channel two-dimensional image, which is then input into the detector's backbone network and feature fusion network (Neck) to extract image features. Subsequently, the pre-trained near-field imaging clutter suppression network based on contrastive learning outputs a class-independent activation map of the foreground corresponding to the target object.

[0080] Further processing based on class-independent activation maps can achieve clutter suppression of layered imaging data. Finally, the clutter-suppressed layered imaging data is processed again through LMIP to generate visualization results, and the location of suspicious targets is marked according to the output of the detection head.

[0081] Step S103: Based on the class-independent activation map, generate a mask matrix, and perform clutter suppression processing on each layer of the layered imaging data according to the mask matrix to suppress other clutter data in the layered imaging data except for the target echo data, while retaining the target echo data in the layered imaging data.

[0082] Millimeter-wave imaging data is often interfered with by complex background clutter in practical applications. Traditional suppression methods struggle to effectively filter out non-target areas, resulting in a large amount of irrelevant information remaining in the imaging results, affecting subsequent identification and analysis. Furthermore, traditional suppression methods rely heavily on manual experience or simple thresholds for mask generation and region selection, making them difficult to adapt to changing real-world scenarios.

[0083] Based on this, after the near-field imaging clutter suppression network based on contrastive learning is trained, a mask matrix can be generated by the generated activation map, thereby guiding the network to selectively process the layered imaging data, effectively suppressing background clutter interference in non-human imaging target areas, and finally obtaining a clear imaging image that retains only human signals.

[0084] As one possible approach, after inputting the layer maximum projection result into a pre-trained contrastive learning-based near-field imaging clutter suppression network to generate a class-independent activation map, the pre-trained contrastive learning-based near-field imaging clutter suppression network generates a certain mask matrix based on this class-independent activation map. This mask matrix is ​​then used to perform clutter suppression processing on each layer of the layered imaging data, suppressing other clutter data in the layered imaging data except for the target echo data, while retaining the target echo data in the layered imaging data.

[0085] In this context, target echo data can be understood as the echo signal that needs to be retained. For example, in the echo data acquired by a millimeter-wave human body security scanner, the foremost human body signal can be understood as the target echo data, while the human body signal and other object signals in the background can be understood as other clutter data that needs to be suppressed.

[0086] Category-independent activation graph For example, a pre-trained near-field imaging clutter suppression network based on contrastive learning can utilize... Generate mask matrix Its mathematical expression can be, but is not limited to, expressed as:

[0087] in, This represents the binarization operation of an image. The threshold for binarization.

[0088] Next, the pre-trained near-field imaging clutter suppression network based on contrastive learning can perform clutter suppression processing on each layer of the layered imaging data using a mask matrix, and then perform layer maximum projection to finally obtain the clutter-suppressed imaging image. Its mathematical expression can be, but is not limited to, the following:

[0089] in, Used to specify whether clutter suppression processing is applied to each distance plane. Figure 4 This is a schematic diagram of a clutter suppression process according to an embodiment of this application, as shown below. Figure 4 As shown, Figure 4 The process of clutter suppression is demonstrated.

[0090] The final image obtained The image will contain only human body signals and not background clutter signals, thus achieving clutter suppression in near-field imaging.

[0091] It should be noted that the human body signal and background clutter signal mentioned here are only illustrative examples. In actual applications, the target echo data that needs to be retained can also be the echo signals of other objects that need to be retained. This can be determined by professionals in this field according to actual needs, and this application embodiment does not impose any restrictions.

[0092] The following is a detailed explanation of the near-field imaging clutter suppression method based on contrastive learning in this application, using a specific embodiment.

[0093] Figure 5 This is a flowchart of a near-field imaging clutter suppression method based on contrastive learning according to an embodiment of this application, as shown below. Figure 5 As shown: Step 1: Collect raw echo data of three-dimensional near-field imaging of the human body and objects from the laboratory environment. After preliminary correction based on differences in radar equipment, complete three-dimensional echo data is obtained. ; Step 2: Set several distances in the imaging space of the radar equipment, and use the RM algorithm to calculate the imaging results at each distance relative to the plane parallel to the radar array, forming a three-dimensional imaging result. Furthermore, through operations such as amplitude preservation, range compression, and quantization, the complex form of the three-dimensional imaging results is transformed. Converted into layered imaging data ; Step 3: Utilize projection at the maximum value Improved hierarchical maximum projection method Layered imaging data After being grouped by layer, each layer is projected onto the three color channels of the image, ultimately forming a multi-channel two-dimensional imaging image. ; Step 4: Image To detect targets, feature maps for small target detection are extracted using a multi-head mechanism of the target detection network. And input it into a class-independent activation graph generation network. This yields a category-independent activation graph of the foreground. ; Step 5: Train the network based on basic contrastive learning and activate the class-independent graphs. With small object detection feature map After flattening and combining, a foreground vector is generated. and background vector , representing the characteristics of human signals and background clutter, respectively. Then, a contrastive learning loss is constructed. This includes negative examples of loss between the foreground and background. And positive examples of foreground-foreground and background-background relationships on loss. By comparing learning loss With detection loss The weighted combination of these factors synthesizes the total loss function for network training. The network is trained using a contrastive learning model based on this loss, resulting in the final near-field imaging clutter suppression network based on contrastive learning. Step 6: Input the layered maximum projection result into the trained contrastive learning-based near-field imaging clutter suppression network to obtain an effective class-independent activation map. Then, based on the category-independent activation graph The size is adjusted by interpolation and threshold binarization is performed to generate a mask matrix. Then use this mask matrix Clutter suppression operations are performed on each layer of the layered imaging data. Finally, after layered maximum projection, the image with background clutter removed is obtained. This achieves suppression of the imaging image. Background clutter will preserve human signals. For example... Figure 6 As shown, Figure 6 This is a diagram illustrating the near-field imaging clutter suppression effect of one embodiment of this application.

[0094] The near-field imaging clutter suppression method based on contrastive learning proposed in this application can input the projection result of the maximum value of each layer of the layered imaging data of the target into a pre-trained near-field imaging clutter suppression network based on contrastive learning. This network then generates a class-independent activation map and its mask matrix for the target, performing clutter suppression processing on each layer of the layered imaging data to suppress clutter signals in the complete echo data. Thus, even in the absence of precise human contour annotation, by fusing layered imaging processing with deep learning technology, constructing a class-independent activation map, and designing a contrastive learning loss, the method improves the ability to distinguish between human foreground signals and background clutter signals, effectively suppressing background clutter, enhancing the target foreground signal, and accurately extracting effective signals and features. This completes the clutter suppression task on the basis of the detection process, improving the imaging quality and target detection capability of millimeter-wave near-field imaging, and enhancing the adaptability and robustness of radar systems in complex environments. It also provides a solid technical foundation for the further application of millimeter-wave imaging technology in security, medical, and other fields. This solves the problems in related technologies, such as the difficulty in accurately extracting effective signals and features in complex environments with strong background clutter interference that often obscures or interferes with human target signals; and the fact that human target features are relatively weak, especially in diverse environments and under complex object occlusion, the clarity and contrast of the imaging results decrease, affecting the subsequent target detection and recognition effects, and limiting the performance and reliability of millimeter-wave near-field imaging in practical applications.

[0095] Next, referring to the accompanying drawings, a near-field imaging clutter suppression device based on contrastive learning proposed according to an embodiment of this application is described.

[0096] Figure 7 This is a schematic diagram of the near-field imaging clutter suppression device based on contrastive learning according to an embodiment of this application.

[0097] like Figure 7 As shown, the near-field imaging clutter suppression device 10 based on contrastive learning includes: a conversion module 100, a first processing module 200, and a suppression module 300.

[0098] The conversion module 100 is used to collect complete echo data of the target body on multiple planes parallel to the radar array plane, so as to obtain the imaging results of the target body on multiple planes, determine the three-dimensional imaging results of the target body based on the imaging results of multiple planes, and determine the layered imaging data of the target body at different distances from the radar array based on the three-dimensional imaging results.

[0099] The first processing module 200 is used to input the projection result of the maximum value of the layer corresponding to the layered imaging data into a pre-trained near-field imaging clutter suppression network based on contrastive learning, so as to output the category-independent activation map of the foreground corresponding to the target.

[0100] The suppression module 300 is used to generate a mask matrix based on the class-independent activation map, and perform clutter suppression processing on each layer of the layered imaging data according to the mask matrix, so as to suppress other clutter data in the layered imaging data except for the target echo data, and retain the target echo data in the layered imaging data.

[0101] Optionally, in one embodiment of this application, it further includes a projection module and a determination module.

[0102] The projection module is used to project the layered imaging data onto multiple color channels sequentially based on multiple planes before inputting the projection result of the layered maximum value corresponding to the layered imaging data into the pre-trained near-field imaging clutter suppression network based on contrast learning, so as to generate a multi-channel two-dimensional image.

[0103] The determination module is used to determine the maximum projection result of the layer based on the multi-channel two-dimensional image.

[0104] Optionally, in one embodiment of this application, the expression for the projection result of the layered maximum value can be, but is not limited to, expressed as: , , , , in, This represents the projection result of the maximum value in the layer; The matrix represents the concatenation of elements along the channel dimension; Represents the set of natural numbers; This indicates the projection processing of the maximum value of the layer; Transitional parameters designed to facilitate the calculation process; Indicates the layer number where the layered data is located. RGB three channels, Indicates the total number of floors. This indicates that the index value is contained in the set along the distance direction. All tangent planes In the middle, select the maximum value among pixels with the same coordinates; This indicates the projection processing of the maximum value; These represent the number of layers, height, and width of the layered imaging, respectively. This represents layered imaging data.

[0105] Optionally, in one embodiment of this application, it further includes: a data acquisition module, an acquisition module, and a calibration module.

[0106] The acquisition module is used to acquire raw echo data of multiple typical objects based on the radar array before acquiring complete echo data of the target object on multiple planes parallel to the radar array plane.

[0107] The acquisition module is used to acquire the correction data of the radar array based on the raw echo data.

[0108] The correction module is used to perform matrix multiplication correction processing on the correction data and the original echo data of the target body to obtain complete echo data.

[0109] Optionally, in one embodiment of this application, it further includes: a second processing module, a first building module, and a second building module.

[0110] The second processing module is used to multiply the compressed vectors corresponding to the category-independent activation maps of multiple typical objects with the compressed vectors corresponding to the small target detection feature maps of multiple typical objects before inputting the projection results of the layered maximum values ​​corresponding to the layered imaging data into the pre-trained near-field imaging clutter suppression network based on contrastive learning, so as to obtain complementary foreground vectors and background vectors.

[0111] The first building module is used to construct negative example pair loss, foreground positive example pair loss, and background positive example pair loss based on foreground vector and background vector.

[0112] The second construction module is used to combine negative example pair loss, foreground positive example pair loss and background positive example pair loss to construct a total loss function for training a pre-trained contrastive learning-based near-field imaging clutter suppression network, so as to use the total loss function to construct the pre-trained contrastive learning-based near-field imaging clutter suppression network.

[0113] Optionally, in one embodiment of this application, the expression for the total loss function may, but is not limited to, be expressed as: , , , , in, Indicates the total loss. Indicates detection loss, Indicates the weight of the comparative loss. Indicates a negative example of loss. This indicates the positive example relative to the total loss. , These represent the foreground positive example loss and the background positive example loss, respectively. Indicates batch size, Indicates the first One foreground vector, Indicates the first One foreground vector, Indicates the first Background vectors, Indicates the first Background vectors.

[0114] It should be noted that the foregoing explanation of the near-field imaging clutter suppression method based on contrastive learning also applies to the near-field imaging clutter suppression device based on contrastive learning in this embodiment, and will not be repeated here.

[0115] The near-field imaging clutter suppression device based on contrastive learning proposed in this application can input the projection result of the maximum value of each layer corresponding to the layered imaging data of the target into a pre-trained near-field imaging clutter suppression network based on contrastive learning. This network then generates a class-independent activation map and its mask matrix for the target, performing clutter suppression processing on each layer of the layered imaging data to suppress clutter signals in the complete echo data. Thus, even in the absence of precise human contour annotation, by fusing layered imaging processing with deep learning technology, constructing a class-independent activation map, and designing a contrastive learning loss, the device improves the ability to distinguish between human foreground signals and background clutter signals, effectively suppressing background clutter, enhancing the target foreground signal, and accurately extracting effective signals and features. This completes the clutter suppression task on the basis of the detection process, improving the imaging quality and target detection capability of millimeter-wave near-field imaging, and enhancing the adaptability and robustness of radar systems in complex environments. It also provides a solid technical foundation for the further application of millimeter-wave imaging technology in security, medical, and other fields. This solves the problems in related technologies, such as the difficulty in accurately extracting effective signals and features in complex environments with strong background clutter interference that often obscures or interferes with human target signals; and the fact that human target features are relatively weak, especially in diverse environments and under complex object occlusion, the clarity and contrast of the imaging results decrease, affecting the subsequent target detection and recognition effects, and limiting the performance and reliability of millimeter-wave near-field imaging in practical applications.

[0116] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of this application. The electronic device may include: The memory 801, the processor 802, and the computer program stored on the memory 801 and capable of running on the processor 802.

[0117] When the processor 802 executes the program, it implements the near-field imaging clutter suppression method based on contrastive learning provided in the above embodiments.

[0118] Furthermore, electronic devices also include: Communication interface 803 is used for communication between memory 801 and processor 802.

[0119] The memory 801 is used to store computer programs that can run on the processor 802.

[0120] The memory 801 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0121] If the memory 801, processor 802, and communication interface 803 are implemented independently, then the communication interface 803, memory 801, and processor 802 can be interconnected via a bus to complete communication between them. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized into address buses, data buses, control buses, etc. For ease of representation, Figure 8 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0122] Optionally, in a specific implementation, if the memory 801, processor 802, and communication interface 803 are integrated on a single chip, then the memory 801, processor 802, and communication interface 803 can communicate with each other through an internal interface.

[0123] The processor 802 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.

[0124] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described near-field imaging clutter suppression method based on contrastive learning.

[0125] This application also provides a computer program product, including a computer program that can run computer instructions. When the computer instructions are executed by a processor, they implement the near-field imaging clutter suppression method based on contrastive learning provided in this application.

[0126] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0127] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "N" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0128] Any process or method described in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or N executable instructions for implementing custom logic functions or processes, and the scope of the preferred embodiments of this application includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of this application pertain.

[0129] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0130] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, it can be implemented using any one or more of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0131] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0132] Furthermore, the functional units in the various embodiments of this application can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0133] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of this application.

Claims

1. A near-field imaging clutter suppression method based on contrastive learning, characterized in that, Includes the following steps: Complete echo data of the target body in multiple planes parallel to the radar array plane are collected to obtain the imaging results of the target body in the multiple planes. The three-dimensional imaging results of the target body are determined based on the imaging results of the multiple planes, and the layered imaging data of the target body at different distances from the radar array are determined based on the three-dimensional imaging results. The projection result of the maximum layer value corresponding to the layered imaging data is input into a pre-trained near-field imaging clutter suppression network based on contrastive learning to output a category-independent activation map of the foreground corresponding to the target body. Based on the category-independent activation map, a mask matrix is ​​generated, and clutter suppression processing is performed on each layer of the layered imaging data according to the mask matrix to suppress other clutter data in the layered imaging data except for the target echo data, while retaining the target echo data in the layered imaging data.

2. The near-field imaging clutter suppression method based on contrastive learning according to claim 1, characterized in that, Before inputting the projection result of the layered maximum value corresponding to the layered imaging data into the pre-trained near-field imaging clutter suppression network based on contrastive learning, the method further includes: Based on the multiple planes, the layered imaging data is sequentially projected onto multiple color channels to generate a multi-channel two-dimensional image; The projection result of the layer maximum value is determined based on the multi-channel two-dimensional image.

3. The near-field imaging clutter suppression method based on contrastive learning according to claim 1, characterized in that, The expression for the projection result of the maximum value of the layer is: , , , , in, This represents the projection result of the maximum value in the layer; The matrix represents the concatenation of elements along the channel dimension; Represents the set of natural numbers; This indicates the projection processing of the maximum value of the layer; Transitional parameters designed to facilitate the calculation process; Indicates the layer number where the layered data is located. RGB three channels, Indicates the total number of floors. This indicates that the index value is contained in the set along the distance direction. All tangent planes In the middle, select the maximum value among pixels with the same coordinates; This indicates the projection processing of the maximum value; These represent the number of layers, height, and width of the layered imaging, respectively. This represents layered imaging data.

4. The near-field imaging clutter suppression method based on contrastive learning according to claim 1, characterized in that, Before acquiring the complete echo data of the target body in the plurality of planes parallel to the radar array plane, the method further includes: Based on the radar array, raw echo data of multiple typical objects are collected; The correction data of the radar array is obtained based on the raw echo data; The corrected data and the original echo data of the target body are subjected to matrix multiplication correction to obtain the complete echo data.

5. The near-field imaging clutter suppression method based on contrastive learning according to claim 4, characterized in that, Before inputting the projection result of the layered maximum value corresponding to the layered imaging data into the pre-trained near-field imaging clutter suppression network based on contrastive learning, the method further includes: The compressed vectors corresponding to the category-independent activation maps of the multiple typical objects are multiplied with the compressed vectors corresponding to the small target detection feature maps of the multiple typical objects to obtain complementary foreground and background vectors. Based on the foreground vector and the background vector, construct the negative example pair loss, the foreground positive example pair loss, and the background positive example pair loss; The negative example pair loss, foreground positive example pair loss, and background positive example pair loss are combined to construct a total loss function for training the pre-trained contrastive learning-based near-field imaging clutter suppression network, so as to construct the pre-trained contrastive learning-based near-field imaging clutter suppression network using the total loss function.

6. The near-field imaging clutter suppression method based on contrastive learning according to claim 5, characterized in that, The expression for the total loss function is: , , , , in, Indicates the total loss. Indicates detection loss, Indicates the weight of the comparative loss. Indicates a negative example of loss. This indicates the positive example relative to the total loss. , These represent the foreground positive example loss and the background positive example loss, respectively. Indicates batch size, Indicates the first One foreground vector, Indicates the first One foreground vector, Indicates the first Background vectors, Indicates the first Background vectors.

7. A near-field imaging clutter suppression device based on contrastive learning, characterized in that, include: The conversion module is used to collect complete echo data of the target body on multiple planes parallel to the radar array plane, so as to obtain the imaging results of the target body on the multiple planes, generate a three-dimensional imaging result of the target body based on the imaging results of the multiple planes, and convert the three-dimensional imaging result into layered imaging data of the target body at different distances from the radar array. The processing module is used to input the projection result of the maximum layer value corresponding to the layered imaging data into a pre-trained near-field imaging clutter suppression network based on contrastive learning, so as to output the category-independent activation map of the foreground corresponding to the target body. The suppression module is used to generate a mask matrix based on the category-independent activation map, and perform clutter suppression processing on each layer of the layered imaging data according to the mask matrix, so as to suppress other clutter data in the layered imaging data except for the target echo data, and retain the target echo data in the layered imaging data.

8. An electronic device, characterized in that, include: The memory, the processor, and the computer program stored in the memory and executable on the processor, the processor executing the program to implement the near-field imaging clutter suppression method based on contrastive learning as described in any one of claims 1-6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the near-field imaging clutter suppression method based on contrastive learning as described in any one of claims 1-6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed, it is used to implement the near-field imaging clutter suppression method based on contrastive learning as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Ground penetrating radar road disease image prediction visualization method based on category activation mapping

    CN117152083A

  • Pulse Doppler radar target detection method and system based on background contrast attention mechanism

    CN120178232A

  • Near-field sparse imaging method based on spatial convolution

    CN120446950A

  • Techniques to process layers of a three-dimensional image using one or more neural networks

    US20210374384A1