Millimeter wave radar dangerous article detection method and device based on a mixture model
By constructing a hybrid model that combines a 3D convolutional neural network and a Transformer module, the problems of privacy leakage and insufficient recognition accuracy in the detection of dangerous items by millimeter-wave radar are solved, achieving efficient and safe identification of dangerous items and improving the safety of public places.
Patent Information
- Application Number
- CN202510464484.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-11-28
- Estimated Expiration
- 2045-04-14
AI Technical Summary
Existing millimeter-wave radar methods for detecting hazardous materials have issues such as privacy risks and insufficient identification accuracy, especially in complex environments where it is difficult to effectively utilize radar signals for hazardous material identification.
A hybrid model-based approach is adopted, combining a 3D convolutional neural network and a Transformer module. The hybrid model is constructed through a local energy 3D convolutional module, an embedding layer, a global energy Transformer module, and a global average pooling layer, and trained using the Polyloss function to achieve efficient feature extraction and classification of radar data.
It improves the accuracy and security of hazardous material identification, avoids privacy leaks, enhances the model's recognition performance in complex environments, and is suitable for security in public places.
Smart Images

Figure CN120370309B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to radar target detection technology, and in particular to a millimeter wave radar dangerous article detection method and device based on a hybrid model. BACKGROUND
[0002] With the increasing demand for public safety, dangerous article identification has become one of the key technologies in the field of security inspection. Currently, X-ray scanning technology is mainly used to check whether passengers carry weapons or explosives and other dangerous goods, but due to its ionizing radiation to the human body, the safety is questioned. As a new type of non-contact detection method, millimeter wave radar technology can penetrate clothing, but does not produce ionizing radiation, is harmless to the human body, and is not affected by factors such as light and smoke, and has all-weather detection capability. In addition, millimeter wave radar can accurately detect objects carried openly on the human body, and has great advantages in dangerous article detection.
[0003] Millimeter wave radar detection of danger usually uses its reflected energy data for target identification, but this method faces challenges. Because the radar reflected energy data lacks detailed image information, only abstract energy data or energy images are provided, resulting in difficulties in distinguishing and identifying precision based on deep learning feature extraction.
[0004] Currently, the dangerous object identification method based on millimeter wave radar mainly includes two types. The first type is an identification method based on millimeter wave radar imaging, which can effectively detect the outline and structure of the target and generate high-resolution three-dimensional images. However, millimeter wave radar imaging technology needs to process a large amount of target human image data, has poor privacy and complex data processing. In addition, this method is easily affected by the reflecting surface and multipath effect, which may cause the image quality to decrease, thereby affecting the detection accuracy of small objects. The second type of method is an identification method based on reflected energy data, which only relies on the energy data reflected by the radar for signal processing. Compared with image recognition based on millimeter wave radar imaging, this method avoids generating image data of the target human body, so it has good privacy protection, but the multiple convolution operations of this method may pay too much attention to local features and ignore global features, especially when there is a lot of noise carrying personnel information, the performance of this method is limited.
[0005] Therefore, how to effectively utilize millimeter wave radar signals for dangerous article identification, avoid privacy leakage and overcome the feature extraction difficulty, and improve the identification accuracy, has become a technical problem to be solved. SUMMARY
[0006] In view of the defects of the prior art, the present application provides a millimeter wave radar dangerous article detection method and device based on a hybrid model.
[0007] To achieve the above object, the technical scheme adopted by the present application is as follows:
[0008] In one aspect, the present application provides a mixed model-based millimeter wave radar dangerous goods detection method, comprising the following steps:
[0009] Obtaining original radar data, preprocessing the original radar data, and obtaining three-dimensional radar data;
[0010] Constructing a mixed model, the mixed model comprising three local energy three-dimensional convolution modules connected in sequence, an embedding layer, three global energy Transformer modules, a global average pooling layer, and a prediction head; the local energy three-dimensional convolution module is used for high-level feature extraction of the input three-dimensional radar data, and the extracted feature map is input into the embedding layer; the embedding layer is used for cutting the feature map into several small blocks and embedding it into a fixed-dimensional vector to form an embedding vector and input into the global energy Transformer module; the global energy Transformer module is used for fusing local features and global features according to the input embedding vector using a self-attention mechanism to improve the feature expression capability; the global average pooling layer is used to extract the global information of the input feature and input into the prediction head; the prediction head classifies and processes the input feature;
[0011] Training the mixed model with a Polyloss function as the target loss function to obtain a trained mixed model, the Polyloss function being
[0012]
[0013] wherein, is a standard cross-entropy loss; is the predicted probability of the target category; is a weight parameter for balancing the cross-entropy and the polynomial loss;
[0014] Using the trained mixed model to detect dangerous goods in practice and output the detection result.
[0015] Further, the three-dimensional radar data is obtained according to the following steps:
[0016] Reshaping the original radar signal into a four-dimensional format covering the fast time, slow time, azimuth angle, and elevation angle dimensions;
[0017] Performing fast Fourier transform on the fast time dimension and the slow time dimension in sequence to generate a range-Doppler matrix;
[0018] Based on the range-Doppler matrix, effectively identifying the targets with motion characteristics in the radar data;
[0019] Compensate the phase error caused by the target motion to obtain a compensated range-Doppler matrix;
[0020] The compensated range-Doppler matrix is subjected to fast Fourier transform in the fast time, azimuth and elevation angle dimensions respectively, and then subjected to non-coherent accumulation in the slow time dimension to generate a three-dimensional data matrix of range-azimuth-elevation, i.e. three-dimensional radar data.
[0021] Further, the method further comprises:
[0022] The distances, speeds and angles of the target points in the three-dimensional data matrix of range-azimuth-elevation are clustered using a DBSCAN algorithm, and a threshold is set according to the speed and step distance of human motion to accurately locate the target position.
[0023] Based on the target position, a three-dimensional matrix of 24x36x10 is cropped from the center of the detected range, azimuth and elevation, and the three-dimensional matrix is used as the three-dimensional radar data.
[0024] Further, a constant false alarm rate algorithm is used to effectively identify the target with motion characteristics in the radar data.
[0025] Further, the compensation of the phase error caused by the target motion specifically comprises:
[0026] Compensation of half of the Doppler phase shift estimated in the velocity fast Fourier transform result.
[0027] Further, the local energy three-dimensional convolution module comprises two 3x3x3 3D convolution layers and one 1x1x1 3D convolution layer. The data input into the local energy three-dimensional convolution module is divided into two paths. One path of data is sequentially subjected to two 3x3x3 3D convolution layers. Each 3x3x3 3D convolution layer extracts features from the data, and the extracted features are subjected to batch normalization and then output after being subjected to non-linear transformation by a ReLU activation function. The other path of data is subjected to feature extraction by a 1x1x1 3D convolution layer, and the extracted features are subjected to batch normalization and then output after being subjected to non-linear transformation by a ReLU activation function. The output features of the two paths are fused by an addition operation, and the fused feature map is subjected to non-linear transformation by a ReLU activation function.
[0028] Further, the global energy Transformer module comprises a multi-head self-attention layer and a one-dimensional convolution layer; the input embedding vector is transmitted to the multi-head self-attention layer after layer normalization and linear projection; each self-attention layer in the multi-head self-attention layer calculates different weighted sums in parallel according to different attention heads, generates attention output, and transmits the attention output to the one-dimensional convolution layer for feature extraction after layer normalization; the extracted features are subjected to nonlinear transformation through a ReLU activation function, and then output after regularization by a dropout layer; the output of the dropout layer is added to the attention output through a residual connection and then output.
[0029] Further, the standard cross-entropy loss is calculated according to the following formula:
[0030]
[0031] wherein, is an indicator function of the true class; is the predicted probability of the model for the class c .
[0032] In another aspect, the application provides a mixed model-based millimeter wave radar dangerous goods detection device, comprising:
[0033] A first module is configured to acquire original radar data, pre-process the original radar data, and obtain three-dimensional radar data;
[0034] A second module is configured to construct a mixed model, wherein the mixed model comprises three local energy three-dimensional convolution modules, an embedding layer, three global energy Transformer modules, a global average pooling layer and a prediction head connected in sequence; the local energy three-dimensional convolution module is configured to extract high-level features from the input three-dimensional radar data and input the extracted feature map to the embedding layer; the embedding layer is configured to cut the feature map into small blocks and embed them into fixed-dimensional vectors to form embedding vectors and input the embedding vectors to the global energy Transformer module; the global energy Transformer module is configured to fuse local features and global features by using a self-attention mechanism according to the input embedding vectors, so as to improve the feature expression capability; the global average pooling layer is configured to extract global information of the input features and input the global information to the prediction head; and the prediction head is configured to perform classification processing on the input features.
[0035] A third module is configured to train the mixed model by taking a Polyloss function as a target loss function, so as to obtain a trained mixed model, wherein the Polyloss function is
[0036]
[0037] wherein, is a standard cross-entropy loss; is a predicted probability of a target class; is a weight parameter for balancing cross-entropy and polynomial loss;
[0038] A fourth module is configured to use the trained hybrid model to perform actual detection on dangerous goods and output a detection result.
[0039] Compared with the prior art, the beneficial technical effects of the present application are that:
[0040] The method and device for detecting dangerous goods based on a hybrid model provided by the present application can construct a three-dimensional convolutional neural network and a Transformer hybrid model, do not need to generate images, directly input the data collected by a radar after preprocessing into the constructed hybrid model for prediction, and comprehensively use the three-dimensional convolutional neural network to extract local features and the Transformer module to capture global features in the hybrid model, so that the hybrid model has good accuracy in the dangerous object recognition task.
[0041] Meanwhile, the Polyloss function is used to replace the traditional cross-entropy loss function to train the hybrid model, so that the model pays more attention to samples with classification errors or low confidence, effectively modulates the characteristics of the task and the data set, and achieves better training effect.
[0042] The present application can efficiently, safely and privately detect dangerous goods, and significantly improves the safety guarantee capability of public places. BRIEF DESCRIPTION OF DRAWINGS
[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor on the basis of the drawings shown.
[0044] Figure 1 A flow chart of the method for detecting dangerous goods based on a hybrid model provided by an embodiment is shown in the figure.
[0045] Figure 2 A radar data collection experiment scene diagram provided by an embodiment is shown in the figure, wherein, Figure 2 (a) is a diagram of an experimenter walking with a gun, Figure 2 (b) is a diagram of an experimenter walking without carrying an object, Figure 2 (c) is a diagram of an experimenter walking with a mobile phone, Figure 2 (d) is a diagram of a detection object;
[0046] Figure 3 A mixed network diagram provided for an embodiment;
[0047] Figure 4 A local energy three-dimensional convolution module diagram provided for an embodiment;
[0048] Figure 5 A global energy Transformer module diagram provided for an embodiment. DETAILED DESCRIPTION
[0049] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0050] Referring to Figure 1 An embodiment provides a mixed model-based millimeter wave radar dangerous goods detection method, including the following steps:
[0051] Obtaining original radar data, preprocessing the original radar data to obtain three-dimensional radar data;
[0052] Constructing a mixed model, the mixed model including three local energy three-dimensional convolution modules connected in sequence, an embedding layer, three global energy Transformer modules, a global average pooling layer and a prediction head; the local energy three-dimensional convolution module is used for high-level feature extraction on the input three-dimensional radar data, and the extracted feature map is input into the embedding layer; the embedding layer is used for cutting the feature map into small pieces and embedding them into fixed-dimensional vectors to form embedding vectors and input them into the global energy Transformer module; the global energy Transformer module is used for fusing local features and global features by using a self-attention mechanism according to the input embedding vectors to improve the feature expression capability; the global average pooling layer is used for extracting the global information of the input features and inputting them into the prediction head; and the prediction head classifies the input features;
[0053] Training the mixed model with a Polyloss function as a target loss function to obtain a trained mixed model, the Polyloss function being
[0054]
[0055] wherein, is a standard cross-entropy loss; is the predicted probability of the target category; is a weight parameter for balancing the cross-entropy and the polynomial loss;
[0056] The trained hybrid model is used for actual detection of dangerous goods, and a detection result is output.
[0057] The three local energy three-dimensional convolution modules gradually expand the channel dimension of the features from 32, 64 to 128, so as to extract rich high-dimensional feature information. Specifically, the three local energy three-dimensional convolution modules gradually expand the channel dimension of the features from 32, 64 channels to 128 channels, gradually construct various local features at different scales, so as to fully extract rich high-dimensional feature information, capture sufficient spatial information before the feature information is transmitted to the Transformer module, and avoid model redundancy; using more than three local energy three-dimensional convolution modules may cause overfitting or require more training data; the three local energy three-dimensional convolution modules are selected in the application, so that a high-efficiency balance between performance and efficiency of high-dimensional feature information is achieved, the performance and efficiency of extracting high-dimensional feature information are considered, and the real-time performance and practicality of the application are enhanced.
[0058] The three global energy Transformer modules are respectively configured with 2, 4 and 8 attention heads, which can capture global features at different scales and levels. This multi-head attention mechanism not only enhances the understanding of the model to the global dependency relationship, but also improves the distinguishing ability of the features, so that the model can still maintain high recognition performance in a complex environment. Specifically, at the low layer (the first Transformer module, 2 attention heads), the model mainly focuses on the basic features and local information; at this time, a small number of attention heads can effectively capture the relevant global information. As the model depth increases, the number of attention heads increases, which can help the model capture more detailed global dependencies and higher-level features. However, increasing the number of attention heads will also increase the computational overhead. With the increase of one attention head, the computational complexity will increase, which will cause a large increase in training time and memory consumption. Too many attention heads may cause model redundancy, which will affect the training efficiency and generalization ability of the model. The application sets 2, 4 and 8 attention heads respectively, which considers the computational efficiency and model performance at the same time, saves the training time and reduces the memory consumption while maintaining the optimal model performance, so that the application can consider timeliness and accuracy in practical application, and also enhances the real-time performance and practicality of the application.
[0059] Referring to Figure 2 In an embodiment, a toy pistol is used and the surface is wrapped with tin paper to simulate the reflection characteristics of the gun, referring to Figure 2(d), the experimenter carries three types of objects, which are a gun, a mobile phone and no object. The radar data collection process is carried out in the laboratory. The millimeter wave radar device is fixedly arranged in front of the laboratory. The experimenter carries any one of the above three types of objects and randomly walks in the range of 0.6 m to 4 m in front of the radar, as shown in Figs. Figure 2 As shown in (a)-(c), data measurement is performed. The measured radar data set contains 29661 frames of single-target radar three-dimensional data, wherein the data of carrying a gun is 9994 frames, the data of carrying a mobile phone is 9728 frames, and the data of no object is 9939 frames.
[0060] In an embodiment, the following steps are used to preprocess the above radar data set to obtain three-dimensional radar data:
[0061] The original radar signal is reshaped into a four-dimensional format covering the fast time, slow time, azimuth angle and elevation angle dimensions; by reshaping the original radar signal into a four-dimensional format, structured space-time information is provided for subsequent signal processing.
[0062] The fast time dimension and the slow time dimension are sequentially subjected to fast Fourier transform to generate a range-Doppler matrix;
[0063] Based on the range-Doppler matrix, the target with motion characteristics in the radar data is effectively identified;
[0064] The phase error caused by target motion is compensated to obtain a compensated range-Doppler matrix;
[0065] After the compensated range-Doppler matrix is subjected to fast Fourier transform in the fast time, azimuth angle and elevation angle dimensions, non-coherent accumulation is performed on the slow time dimension to generate a range-azimuth-elevation three-dimensional data matrix, i.e., three-dimensional radar data.
[0066] In a preferred embodiment, a constant false alarm rate algorithm is used to effectively identify the target with motion characteristics in the radar data, and the constant false alarm rate algorithm can further improve the accuracy of target detection.
[0067] In an embodiment, in the angle estimation stage, especially when angle estimation is performed on a MIMO virtual array, the phase error caused by target motion needs to be compensated; in a time division multiplexing (TDM) scheme, due to the switching delay between the transmitters, the phase difference caused by motion is a factor that must be considered. After target detection is completed, the original data is phase compensated for the moving target, thereby ensuring the accuracy of the signal. In this embodiment, the phase error correction is performed by compensating for half of the Doppler phase shift estimated in the fast Fourier transform result.
[0068] In a preferred embodiment, the above preprocessing further includes:
[0069] The distance, speed and angle of the target point in the three-dimensional data matrix of distance-azimuth-elevation angle are clustered using the DBSCAN algorithm, and the threshold is set according to the human motion speed and step distance to accurately locate the target position;
[0070] Based on the target position, a 24x36x10 three-dimensional matrix is cropped with the detected distance, azimuth and elevation angle as the center, and the three-dimensional matrix is used as three-dimensional radar data.
[0071] Each frame of data of the three-dimensional radar data is taken as a training sample and input into the hybrid model.
[0072] Referring to Figure 3 , the hybrid model comprises three local energy three-dimensional convolution modules, an embedding layer, three global energy Transformer modules, a global average pooling layer and a prediction head connected in sequence; the local energy three-dimensional convolution module is used for high-level feature extraction of the input three-dimensional radar data, and the extracted feature map is input into the embedding layer; the embedding layer is used for cutting the feature map into small blocks and embedding it into a fixed dimension vector to form an embedding vector and input into the global energy Transformer module; the global energy Transformer module is used for fusing local features and global features according to the input embedding vector using a self-attention mechanism to improve the feature expression ability; the global average pooling layer is used for extracting the global information of the input feature and inputting it into the prediction head; the prediction head classifies the input feature; specifically, the global average pooling layer uses a feedforward neural network to further process and classify the output feature; since the global average pooling layer reduces the parameter amount, the complexity of the model is reduced, thereby preventing overfitting to a certain extent, so that the model has better generalization ability when processing data; at the same time, the global average pooling layer reduces the dimension of the output feature through the pooling operation, so that the model has stronger robustness to local changes of the input data.
[0073] Referring to Figure 4In an embodiment, the local energy three-dimensional convolution module includes two 3x3x3 three-dimensional convolution layers, a 1x1x1 three-dimensional convolution layer, data input into the local energy three-dimensional convolution module is divided into two paths, one path of data sequentially passes through the two 3x3x3 three-dimensional convolution layers, each 3x3x3 three-dimensional convolution layer extracts features from the data, the extracted features are batch normalized, and then output after nonlinear transformation by a ReLU activation function; the other path of data passes through the 1x1x1 three-dimensional convolution layer to extract features, the extracted features are batch normalized, and then output after nonlinear transformation by a ReLU activation function; the output features of the two paths are fused by an addition operation (residual connection), and the fused feature map is output after nonlinear transformation by a ReLU activation function. The feature information of the two paths is fused, thereby enhancing the expression ability of the three-dimensional convolutional neural network.
[0074] With reference to Figure 5 In an embodiment, the global energy Transformer module includes a multi-head self-attention layer and a one-dimensional convolution layer; the input embedding vector is transmitted to the multi-head self-attention layer after layer normalization and linear projection, each self-attention layer in the multi-head self-attention layer calculates different weighted sums in parallel according to different attention heads to generate attention output, the generated attention output can simultaneously focus on multiple different parts of the input features, thereby enhancing the modeling ability of the model for complex dependency relationships; the attention output is transmitted to the one-dimensional convolution layer for feature extraction after layer normalization, the extracted features are nonlinearly transformed by a ReLU activation function, and then output after regularization by a dropout layer; the output of the dropout layer is added to the attention output by a residual connection and then output.
[0075] In a preferred embodiment, the local energy three-dimensional convolution module and the global energy Transformer module remain consistent with the above, in this embodiment, through the combination of the local energy three-dimensional convolution module and the global energy Transformer module, wherein the local energy three-dimensional convolution module is responsible for extracting local spatial features, and the global energy Transformer module models global information through a self-attention mechanism, the local features extracted by the local energy three-dimensional convolution module provide high-quality input for the global energy Transformer module, and the global energy Transformer module further enhances the global correlation between features, avoids the limitations of relying only on local features, and ensures that the model can fully utilize local spatial features and global dependencies. This combination enables the model to effectively capture local features and enhance the understanding of cross-region dependencies, thereby improving the accuracy and robustness of target recognition in radar data. At the same time, the global energy Transformer module also enables the model to handle larger range dependencies, making the model suitable for target recognition in complex scenarios, such as some practical application scenarios, where the motion trajectory of an object may span multiple regions of radar data, and traditional convolution methods may have difficulty capturing such cross-region dependencies, while the self-attention mechanism of the global energy Transformer module can effectively associate information from different regions, thereby improving overall recognition effectiveness.
[0076] In an embodiment, the Polyloss function is used as the target loss function, and the hybrid model is trained by the SGD optimizer, the learning rate is dynamically adjusted by the cosine annealing method, the initial learning rate is set to 1x10 -2 and gradually reduced to 1x10 -6 by the cosine annealing method. The batch size is set to 32, and in the multi-classification task, the hybrid model is trained using the PolyLoss function to comprehensively optimize the prediction performance of the model in each class. The Polyloss function is
[0077] ;
[0078] The standard cross-entropy loss is calculated according to the following formula:
[0079]
[0080] where, is the indicator function of the true class; is the predicted probability of the model for class c . At this time, the Polyloss function is:
[0081] .
[0082] The Polyloss function adopted by the present application is used to train the mixed model instead of the traditional cross-entropy loss function. The polynomial correction term is added on the basis of the traditional cross-entropy loss function, so that the model gives higher attention to the samples with classification errors or confidence disclosure, can effectively modulate the characteristics of the task and the data set, and achieve better training effect.
[0083] In an embodiment, the dimension of the input data is 24x36x10, and the dimension of the output data is 3x1; the training set and the validation set are divided according to the ratio of 7:3, the optimizer is selected as the adam optimizer, the learning rate is dynamically modulated, and the learning rate is gradually reduced from 0.01 to 1x10 -4 ; the total training period is set to 100, and the recall rate is used to evaluate the performance of the mixed model, so as to accurately evaluate the performance of the mixed model.
[0084] In actual testing, for the three cases of carrying a gun, carrying a mobile phone and no object, the recall rates of the mixed model are 85.8%, 87.5% and 97.6% respectively, which shows that the present application has high accuracy and practicability in public carrying object detection.
[0085] Referring to Table 1, the average precision and recall rate of a model in the prior art and the model described in the present application are compared. Compared with the model in the prior art, the model described in the present application improves the average precision and recall rate by 24.7% and 6.9% respectively.
[0086] Table 1 Average precision and recall rate of two models
[0087]
[0088] In an embodiment, an ablation experiment is performed to verify the advantages of the model described in the present application. In this embodiment, a local energy three-dimensional convolution module is used as the backbone network, and the Transformer module is not introduced. Referring to Table 2, the results show that the model proposed in the present application improves the average precision and recall rate by 3.9% and 3.1% respectively compared with the model without introducing the Transformer module.
[0089] Table 2 Performance comparison of ablation experiment
[0090]
[0091] In an embodiment, 5 samples are randomly selected for prediction, and the prediction results are shown in Table 3. The inference time of single frame data is 0.26 seconds in terms of inference speed, which shows a fast prediction speed. The speed shows that the model described in the present application has high efficiency in real-time application and can meet the needs of most actual scenes.
[0092] Table 3 Predicted estimates for five samples
[0093]
[0094] The application effectively balances precision and efficiency, further improving feasibility and practicality.
[0095] In another embodiment, a mixed model-based millimeter wave radar dangerous goods detection device is provided, comprising:
[0096] A first module is configured to obtain original radar data, pre-process the original radar data, and obtain three-dimensional radar data.
[0097] A second module is configured to construct a mixed model, wherein the mixed model comprises three local energy three-dimensional convolution modules, an embedding layer, three global energy Transformer modules, and a global average pooling layer connected in sequence; the local energy three-dimensional convolution module is configured to extract high-level features from the input three-dimensional radar data and input the extracted feature map to the embedding layer; the embedding layer is configured to cut the feature map into small blocks and embed them into fixed-dimensional vectors to form embedding vectors and input them to the global energy Transformer module; the global energy Transformer module is configured to fuse local features and global features using a self-attention mechanism based on the input embedding vectors to improve feature expression capability; and the global average pooling layer is configured to extract global information of the input features and input them to a prediction head; the prediction head is configured to classify and process the input features.
[0098] A third module is configured to train the mixed model using a Polyloss function as a target loss function to obtain a trained mixed model, wherein the Polyloss function is
[0099]
[0100] wherein, is a standard cross-entropy loss; is a predicted probability of a target category; is a weight parameter for balancing cross-entropy and polynomial loss;
[0101] A fourth module is configured to use the trained mixed model to perform actual detection on dangerous goods and output a detection result.
[0102] The remaining matters of the application are known technologies.
[0103] The technical features of the above embodiments can be combined arbitrarily, and to make the description concise, not all possible combinations of the technical features in the above embodiments are described, however, as long as the combinations of the technical features do not exist contradictory, they should be considered as the scope of the present application.
[0104] The above-described embodiments are merely illustrative for several embodiments of the present application, and the description is relatively specific and detailed, but should not be understood as limiting the scope of the application. It should be noted that for those skilled in the art, several modifications and improvements can be made without departing from the concept of the present application, which are all within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
[0105] The above-described is only the preferred embodiment of the present application, and is not used to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. within the spirit and principle of the present application should be included in the protection scope of the present application.
Claims
1. A millimeter-wave radar method for detecting hazardous materials based on a hybrid model, characterized in that, Includes the following steps: The raw radar data is acquired and preprocessed to obtain three-dimensional radar data. A hybrid model is constructed, comprising three locally energetic 3D convolutional modules, an embedding layer, three globally energetic Transformer modules, a global average pooling layer, and a prediction head, all connected sequentially. The locally energetic 3D convolutional modules perform high-level feature extraction on the input 3D radar data and input the extracted feature maps into the embedding layer. The embedding layer segments the feature maps into small blocks and embeds them into vectors of fixed dimensions, forming embedding vectors, which are then input into the globally energetic Transformer modules. The globally energetic Transformer modules, based on the input embedding vectors, utilize a self-attention mechanism to fuse local and global features, thereby enhancing feature representation capabilities. The global average pooling layer is used to extract global information of the input features and then input it into the prediction head; the prediction head performs classification processing on the input features. The hybrid model is trained using the Polyloss function as the target loss function to obtain a trained hybrid model. in, Standard cross-entropy loss; The predicted probability of the target category; These are weighting parameters used to balance cross-entropy and polynomial loss; The trained hybrid model is used to detect hazardous materials and output the detection results.
2. The millimeter-wave radar hazardous materials detection method based on a hybrid model as described in claim 1, characterized in that, Obtain 3D radar data by following these steps: The original radar signal was reshaped into a four-dimensional format encompassing fast time, slow time, azimuth, and elevation dimensions; Perform Fast Fourier Transform sequentially on the fast time dimension and the slow time dimension to generate the distance-Doppler matrix; Based on the range-Doppler matrix, targets with motion characteristics in radar data can be effectively identified; The phase error caused by the target motion is compensated to obtain the compensated range-Doppler matrix; After the compensated range-Doppler matrix is subjected to fast Fourier transform in the fast time, azimuth, and elevation dimensions, it is then incoherently accumulated in the slow time dimension to generate a three-dimensional data matrix of range-azimuth-elevation, i.e., three-dimensional radar data.
3. The millimeter-wave radar hazardous materials detection method based on a hybrid model as described in claim 2, characterized in that, Also includes: The DBSCAN algorithm is used to cluster the distance, velocity, and angle of target points in the three-dimensional data matrix of distance-azimuth-elevation angle, and thresholds are set according to human movement speed and stride distance to accurately locate the target position. Based on the target location, a 24×36×10 three-dimensional matrix is cropped with the detected distance, azimuth, and elevation angle as the center, and the three-dimensional matrix is used as three-dimensional radar data.
4. The millimeter-wave radar hazardous materials detection method based on a hybrid model as described in claim 2, characterized in that, The constant false alarm rate algorithm effectively identifies targets with motion characteristics in radar data.
5. The millimeter-wave radar method for detecting hazardous materials based on a hybrid model as described in claim 2, characterized in that, The specific steps to compensate for the phase error caused by the target motion are as follows: The compensation speed is half of the estimated Doppler phase shift in the fast Fourier transform result.
6. The millimeter-wave radar method for detecting hazardous materials based on a hybrid model as described in claim 1, characterized in that, The local energy 3D convolutional module includes two 3×3×3 3D convolutional layers and one 1×1×1 3D convolutional layer. The data input to the local energy 3D convolutional module is divided into two paths. One path passes through the two 3×3×3 3D convolutional layers sequentially. Each 3×3×3 3D convolutional layer extracts features from the data. After batch normalization of the extracted features, a nonlinear transformation is performed using the ReLU activation function before output. The other path passes through the 1×1×1 3D convolutional layer for feature extraction. After batch normalization of the extracted features, a nonlinear transformation is performed using the ReLU activation function before output. The two output features are fused by an addition operation, and the fused feature map is then nonlinearly transformed using the ReLU activation function before output.
7. The millimeter-wave radar method for detecting hazardous materials based on a hybrid model as described in claim 1, characterized in that, The global energy Transformer module includes a multi-head self-attention layer and a one-dimensional convolutional layer. The input embedding vector is normalized and linearly projected through layers before being transmitted to the multi-head self-attention layer. Each self-attention layer in the multi-head self-attention layer calculates different weighted sums in parallel according to different attention heads to generate attention outputs. The attention outputs are normalized through layers before being transmitted to the one-dimensional convolutional layer for feature extraction. The extracted features are nonlinearly transformed using the ReLU activation function and then regularized through a dropout layer before being output. The output of the dropout layer is added to the attention output through a residual connection before being output.
8. The millimeter-wave radar method for detecting hazardous materials based on a hybrid model as described in claim 1, characterized in that, The standard cross-entropy loss Calculate using the following formula: in, An indicator function for the true category; For the model to class c The predicted probability.
9. A millimeter-wave radar hazardous materials detection device based on a hybrid model, characterized in that, include: The first module is used to acquire raw radar data, preprocess the raw radar data, and obtain three-dimensional radar data. The second module is used to construct a hybrid model, which includes three locally energetic 3D convolutional modules, an embedding layer, three globally energetic Transformer modules, a global average pooling layer, and a prediction head, connected in sequence. The locally energetic 3D convolutional modules are used to perform high-level feature extraction on the input 3D radar data and input the extracted feature maps into the embedding layer. The embedding layer is used to cut the feature maps into several small blocks and embed them into vectors of fixed dimensions to form embedding vectors, which are then input into the globally energetic Transformer modules. The globally energetic Transformer modules are used to fuse local and global features using a self-attention mechanism based on the input embedding vectors to improve feature representation capabilities. The global average pooling layer is used to extract global information of the input features and then input it into the prediction head; the prediction head performs classification processing on the input features. The third module is used to train the hybrid model using the Polyloss function as the target loss function, resulting in a trained hybrid model. The Polyloss function is... in, Standard cross-entropy loss; The predicted probability of the target category; These are weighting parameters used to balance cross-entropy and polynomial loss; The fourth module is used to perform actual detection of dangerous items using a trained hybrid model and output the detection results.
Citation Information
Patent Citations
Millimeter wave radar dynamic gesture recognition method applied to interference environment
CN116794602A
Millimeter wave radar fall detection method based on Transform and multi-scale convolutional neural network
CN118050703A