A target recognition and positioning method based on deep learning

By constructing a range Doppler map and performing manual annotation and feature extraction, and using a deep learning network for target recognition and positioning, the problem of the difficulty in integrating target recognition and positioning functions in radar is solved, and efficient target recognition and positioning effects are achieved.

CN120428192BActive Publication Date: 2025-09-30ANHUI DIWAN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510940595.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-09
Publication Date
2025-09-30
Estimated Expiration
2045-07-09

AI Technical Summary

Technical Problem

In existing radar technology, target recognition and positioning functions are difficult to integrate, resulting in increased workload and possible target mismatches. In addition, the deep learning model has poor interpretability, making it difficult to independently improve the accuracy of recognition and positioning.

Method used

By acquiring echo signals to construct a range Doppler map, manual annotation and feature extraction are performed to form a feature map, candidate frames are constructed and positive samples and negative samples are separated. Convolutional neural networks are used for feature extraction and fully connected neural networks are trained to achieve target recognition and positioning.

Benefits of technology

It achieves rapid integration of radar target recognition and positioning, improves the integration and performance of data streams, and improves the accuracy and efficiency of recognition and positioning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120428192B_ABST
    Figure CN120428192B_ABST
Patent Text Reader

Abstract

The present invention provides a target recognition and positioning method based on deep learning, which relates to the field of radar detection technology. The present invention constructs a range Doppler map according to the echo signal, performs feature extraction after manual annotation to form a feature map, and then constructs a candidate frame to separate the features in the feature map, forms positive samples and negative samples based on the relationship between the candidate frame and the feature frame, separates the key features, constructs a target discrimination model with the thinking of processing pictures, utilizes the correspondence between the feature map and the range Doppler, constructs a positioning model with the thinking of a regression model, and determines the target point of the range Doppler model, thereby realizing rapid positioning of the target. The present invention realizes target recognition and positioning by visualizing the echo signal first, and then realizing target recognition and positioning through the discrimination model and the positioning model, thereby integrating data streams and improving the performance of radar in target recognition and positioning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of radar detection technology, and specifically to a target recognition and positioning method based on deep learning. Background Art

[0002] Radar technology for target detection is becoming increasingly mature. The echo signal can be processed more accurately and converted into a range Doppler map, and the target can be located based on the range Doppler map. However, it is generally based on traditional calculation methods, which to a certain extent affects the accuracy and efficiency. With the development of deep learning technology, end-to-end recognition and positioning have been achieved, which has improved the data processing capability to a certain extent. However, the recognition function and positioning function are difficult to integrate. Moreover, due to the poor interpretability of deep learning models, it is difficult to make independent improvements in the middle. Therefore, it is necessary to strengthen the processing of signals.

[0003] In the prior art, publication number CN111985349A discloses a radar received signal type classification and identification method and system, which obtains a radar received signal; performs moving target detection processing on the received signal to obtain a three-dimensional range Doppler plane; processes the three-dimensional range Doppler plane to obtain a top-view feature map of the three-dimensional range Doppler plane, converts the top-view feature map into a grayscale image and performs binarization processing to obtain a binary feature map; inputs the binary feature map into a pre-trained neural network-based radar scene signal processing model, and outputs a classification and identification result.

[0004] Publication number CN119780905A discloses a radar tracking and positioning method, device, and medium combined with micro-Doppler harmonics. The method obtains state information of a tracked target at the previous moment, and determines the main peak distance unit prediction range and the main peak Doppler frequency prediction range of the tracked target at the current moment based on the state information and the state transition equation; determines multiple micro-Doppler harmonic secondary peaks of the tracked target based on the main peak distance unit prediction range, the main peak Doppler frequency prediction range, and the offset of the micro-Doppler harmonic secondary peak of the tracked target relative to the main peak; determines the observation information of the main peak of the tracked target at the current moment based on the multiple micro-Doppler harmonic secondary peaks; and determines the position of the tracked target at the current moment based on the observation information of the main peak.

[0005] Although the public documents achieve the classification and positioning of targets, they belong to two separate systems. In actual applications, two completely independent learning models need to be designed, which increases the workload. Moreover, when the two parts work independently, there may be a mismatch between the two targets, and the data flow cannot be shared to the greatest extent.

[0006] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention

[0007] The purpose of the present invention is to provide a target recognition and positioning method based on deep learning to solve the problems raised in the above background technology.

[0008] To achieve the above object, the present invention provides the following technical solutions:

[0009] A target recognition and positioning method based on deep learning, the specific steps include:

[0010] Step 1: Acquire the echo signal and convert it into a complex signal matrix and perform Fourier transform to obtain the complex amplitude of the complex signal matrix for each frequency unit of each distance unit;

[0011] Step 2: The complex amplitude is positively computed to form an amplitude matrix. The amplitude matrix is ​​converted into a two-dimensional range-Doppler map. At the same time, a ground truth box that reflects the positioning is manually selected. The range-Doppler map and the ground truth box are input into a convolutional neural network for feature extraction to form a feature map and feature box.

[0012] Step 3: Three categories of candidate boxes are formed based on the area of ​​the feature box and arranged on the feature map. The intersection-over-union (IoU) of each candidate box and the feature box is obtained. All candidate boxes are divided into positive samples and negative samples according to the IoU threshold and input into the fully connected neural network for training to obtain the category discrimination model.

[0013] Step 4: Obtain the candidate box of the result obtained in the category discrimination model, mark it as the target box, use the center offset vector between the target box and the feature box as the label, input the target box and internal features into the fully connected neural network model for training, and obtain the positioning model;

[0014] Step 5: The radar acquires echoes in real time and forms a feature map. Three types of candidate boxes are arranged in the feature map and input into the discrimination model to determine the target category. The target box and internal features are obtained and input into the positioning model. The center offset vector and positioning identification box are obtained to realize target positioning.

[0015] Furthermore, historical data of the radar is obtained, the historical data including received echo signals and the target category identified for each echo signal, the target category being the type of aircraft detectable by the radar. The sampling rate, number of single pulse sampling points, and number of pulses of the radar are obtained, the echo signal is mixed to obtain a corresponding complex signal, the complex signal is filtered using a fixed bandwidth filter of the radar, the sampling points within each pulse are weighted using a Hamming window function, the weighted complex signal is normalized, and a complex signal matrix is ​​obtained, based on the following formula:

[0016] ;

[0017] in, is the complex signal matrix, Represents the first Pulse No. The signal value of the sampling point, Retrieve variable for pulse number, , , is the total number of pulses in the echo signal, Retrieve variable for the number of sample points in a single pulse, , , is the number of sampling points for a single pulse.

[0018] Furthermore, an intra-pulse Fourier transform is performed on each pulse in the complex signal matrix, where one row of the complex signal matrix represents one pulse. The intra-pulse Fourier transform is based on the following formula:

[0019] ;

[0020] in, Indicates the In the pulse The complex magnitude of the distance units, Represents the first In the pulse The signal value of the sampling point, Retrieve variable for pulse number, , , is the total number of pulses in the echo signal, Retrieve variable for distance cell number, , , is the total number of distance units, , Retrieve variable for the number of sample points in a single pulse, , , is the number of sampling points for a single pulse, is the Fourier transform basis function.

[0021] Furthermore, the complex amplitude matrix is ​​subjected to an inter-pulse Fourier transform according to the following formula:

[0022] ;

[0023] in, Indicates the The distance unit The complex amplitude of frequency units, is the retrieval variable of the distance unit, , , is the total number of frequency units, , is the total number of pulses in the echo signal, Indicates the In the pulse The complex magnitude of the distance units, Retrieve variable for distance cell number, , , is the total number of distance units, Retrieve variable for pulse number, , , is the Fourier transform base.

[0024] Furthermore, the complex amplitudes formed by the inter-pulse Fourier transform are positively calculated to form an amplitude matrix, according to the following formula:

[0025] ;

[0026] in, Indicates the magnitude matrix The distance unit The amplitude of a frequency unit, Indicates the The distance unit The complex amplitude of frequency units, is the retrieval variable of the distance unit, , , is the total number of frequency units, Retrieve variable for distance cell number, , , is the total number of distance units;

[0027] Convert the amplitude matrix to a range-Doppler map using the following logic:

[0028] The number of pixels in the range Doppler image is the same as the number of elements in the amplitude matrix. The number of the pixel points in the range Doppler image is the same as the coordinate number of the elements in the amplitude matrix. The pixel value of each pixel point in the range Doppler image is the amplitude value of the corresponding position in the amplitude matrix. The horizontal axis of the range Doppler image represents the distance, and the vertical axis of the range Doppler image represents the frequency shift.

[0029] Furthermore, each range Doppler image is manually framed. The frame selection range is the part of the range Doppler image where the pixel value is significantly higher than the surrounding part. The manually determined frame is marked as the real frame, the type of the real frame is marked, and the type is the category of the target. It is recorded through one-hot encoding, and the height and width of the real frame are marked at the same time. The range Doppler image is input into the convolutional neural network to obtain the feature map. The convolution kernel size of the first convolutional layer of the convolutional neural network is 3x3, the number is 32, the step size is 1, and the output size is , the first pooling layer has a step size of 2 and an output size of , the convolution kernel size of the second convolution layer is 3x3, the number is 64, and the output size is , the stride of the second pooling layer is 2, and the output size is , the convolution kernel size of the third convolution layer is 3x3, the number is 128, and the output size is , the feature map size output by the final feature output layer is , where the feature map contains the true box after convolution, and the true box at this time is marked as the feature box.

[0030] Furthermore, the anchor box is set on the feature map, and the logic is as follows:

[0031] Get the feature frame with the largest area, calibrate its width and height as the first candidate frame, and reduce it according to its width and height, reducing the width and height by one-third respectively. If the width or height after reduction is not an integer, round it up. The frame formed by the width and height at this time is calibrated as the second candidate frame. Reduce the width and height of the first candidate frame by two-thirds. If the width or height after reduction is not an integer, round it up. The frame formed by the width and height at this time is calibrated as the third candidate frame;

[0032] The three types of candidate boxes are arranged on the feature map respectively. The arrangement logic of the candidate boxes of each type is as follows:

[0033] The candidate box is first located in the upper left corner to form the first position, and then it is translated one grid to the right to form the second position. In this way, the first row of candidate boxes is formed, and the rightmost candidate box is located in the upper right corner. Then all the candidate boxes in the first row are translated down one grid to form the second row of candidate boxes, and then translated down one grid to form the third row of candidate boxes. In this way, an arrangement that completely covers the feature map is formed.

[0034] Furthermore, in a single feature map, the intersection-and-union (IoU) of each candidate box and the feature box is determined separately, and an IoU threshold is set. The candidate boxes and their in-box features that are greater than or equal to the IoU threshold are calibrated as positive samples, and the candidate boxes and their in-box features that are lower than the IoU threshold are calibrated as negative samples. The positive samples are labeled with the labels of the true boxes to which they belong, and the negative samples are labeled as background. All positive and negative samples and their corresponding labels are summarized and trained in a fully connected neural network to obtain a category discrimination model.

[0035] Furthermore, the candidate box of the determination result obtained in the output layer of the category discrimination model is obtained and marked as the target box. The logic of obtaining the candidate box of the determination result is:

[0036] A single feature map has multiple positive samples. Each positive sample will output the probability of this sample on each target category in the fully connected layer in the category discrimination model. The judgment result of this discrimination model is obtained, and the positive sample with the highest probability on the target category consistent with the judgment result is obtained. The candidate box corresponding to this positive sample is the target box;

[0037] The center offset vectors of the target box and the feature box in each feature map are compared respectively. Each feature map is labeled with the corresponding center offset. The target box and internal features are input into the fully connected neural network model for training to obtain the positioning model.

[0038] Furthermore, an echo is obtained by radar, and a range Doppler map is obtained through step 1. This is input into the convolutional neural network of step 2 to form a feature map. Three types of candidate boxes are arranged in the feature map. Each candidate box and the internal feature box are input into the distinction model to determine the target category and obtain the target box. The target box and the internal features are input into the positioning model to obtain the center offset. The target box is adjusted according to the center offset vector to obtain the positioning recognition box. The positioning recognition box is restored according to the convolutional neural network to obtain the distance point. The restoration logic is as follows:

[0039] In the process of converting the range Doppler map into a feature map in the convolutional neural network, the size of the range Doppler map is changed, but the internal relative relationship does not change. The relative position of the fixed point in the feature map is the same as the relative position in the range Doppler map. Therefore, according to the relative position of the center point in the target frame in the feature map, a specific point in the range Doppler map is corresponding to this point, and this point is marked as a range point.

[0040] The distance to the target is obtained based on the position of the range point in the range-Doppler map.

[0041] Compared with the prior art, the present invention has the following beneficial effects:

[0042] The present invention constructs a range Doppler map based on the echo signal, manually annotates the range Doppler map and then extracts features to form a feature map, then constructs a candidate frame to separate the features in the feature map, forms positive samples and negative samples based on the relationship between the candidate frame and the feature frame, separates the key features, constructs a target discrimination model with the thinking of processing pictures, utilizes the correspondence between the feature map and the range Doppler, constructs a positioning model in the form of a regression model and determines the target point of the range Doppler model, thereby realizing rapid positioning of the target. The present invention first visualizes the echo signal and then realizes target recognition and positioning through the discrimination model and the positioning model, thereby integrating data streams and improving the performance of the radar in target recognition and positioning. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 Schematic diagram of the overall method flow of the present invention;

[0044] Figure 2 This is a schematic diagram of manually marking a true frame on a range-Doppler map according to the present invention. DETAILED DESCRIPTION

[0045] In order to make the objectives, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below with reference to specific embodiments.

[0046] It should be noted that, unless otherwise defined, the technical or scientific terms used in the present invention should have the usual meanings understood by people with ordinary skills in the field to which the present invention belongs. The "first", "second" and similar words used in the present invention do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative position relationships. When the absolute position of the object being described changes, the relative position relationship may also change accordingly.

[0047] Example:

[0048] See also Figure 1 , the present invention provides a technical solution:

[0049] A target recognition and positioning method based on deep learning, the specific steps include:

[0050] Step 1: Acquire the echo signal and convert it into a complex signal matrix and perform Fourier transform to obtain the complex amplitude of the complex signal matrix for each frequency unit of each distance unit;

[0051] The step 1 includes the following:

[0052] Step 101: Obtain historical data of the radar, including received echo signals and the target category identified for each echo signal, where the target category is the type of aircraft detectable by the radar. Obtain the sampling rate, number of single pulse sampling points, and number of pulses of the radar, mix the echo signals to obtain the corresponding complex signals, filter the complex signals using a fixed-bandwidth filter of the radar, weight the sampling points within each pulse using a Hamming window function, and normalize the weighted complex signals to obtain a complex signal matrix. The formula is as follows:

[0053] ;

[0054] in, is the complex signal matrix, Represents the first Pulse No. The signal value of the sampling point, Retrieve variable for pulse number, , , is the total number of pulses in the echo signal, Retrieve variable for the number of sample points in a single pulse, , , is the number of sampling points for a single pulse.

[0055] By acquiring historical data and performing mixing, fixed-bandwidth filtering, Hamming window weighting, and normalization on the echo signal, it is possible to ensure that noise and other interference are effectively removed from the original signal, thereby ensuring that the complex signal matrix truly and stably reflects the sampling point information within each pulse and reducing the risk of data distortion caused by noise. In addition, this step provides high-quality input data for subsequent intra-pulse Fourier transform processing, ensuring that subsequent frequency domain transforms can accurately capture the physical characteristics of the target. At the same time, it serves as a bridge for data preprocessing in the entire system, allowing each subsequent step to be smoothly connected and improving the overall robustness and accuracy of the system.

[0056] Step 102: Perform an intra-pulse Fourier transform on each pulse in the complex signal matrix. A row of the complex signal matrix represents a pulse. The intra-pulse Fourier transform is based on the following formula:

[0057] ;

[0058] in, Indicates the In the pulse The complex magnitude of the distance units, Represents the first In the pulse The signal value of the sampling point, Retrieve variable for pulse number, , , is the total number of pulses in the echo signal, Retrieve variable for distance cell number, , , is the total number of distance units, , Retrieve variable for the number of sample points in a single pulse, , , is the number of sampling points for a single pulse, is the Fourier transform basis function.

[0059] is a complex magnitude, specifically reflecting the After the pulse is Fourier transformed, The frequency domain response over a range unit represents the time domain data collected from the real echo signal. The comprehensive contribution of the frequency domain components. In physical terms, represents different pulses, that is, multiple pulse samples due to the existence of target echo in actual radar system, and Corresponding to the distance resolution unit divided by Fourier transform, it reflects the energy information contained in the signal of each distance unit, so it can be regarded as the echo intensity or amplitude of the target at different distances. Internal summation variable represents the continuous sampling points in a single pulse, and its changes reflect the changes of the signal over time in the actual environment, while It means in the In the pulse The signal value at each sampling moment reflects the radar echo signal in the real environment. With independent variables The relationship is through the complex exponential function It plays the role of "rotating" the phase of each sampling point in the independent variable, thereby mapping the time domain signal to the frequency domain and capturing the characteristic information of a specific distance unit. In other words, the independent variable After the values ​​at different sampling points are superimposed, the energy distribution of each distance unit is reflected by multiplying them with different frequency factors. Therefore, when the signal strength at certain moments in X(m,k) is relatively high, according to the influence of the corresponding frequency components, It will also show a larger amplitude at the corresponding distance unit, and this superposition effect accurately expresses the energy distribution characteristics of the actual echo signal at different distances, thereby achieving effective analysis of the target distance information. The signal amplitude in the sample is enhanced or the signal is more concentrated in a specific sampling area. The modulus (i.e. amplitude) of The change of the value, different frequency factors will change the effect of phase superposition of different sampling points, resulting in differences in the response strength of the signal energy at different distance units. The entire Fourier transform process accurately decomposes the continuous signal in the time domain into multiple discrete frequency domain components to facilitate subsequent target detection and analysis.

[0060] A Fourier transform is performed on each pulse, converting the time domain data into the frequency domain and generating complex amplitude information corresponding to each distance unit. This not only directly demodulates the original sampling point data into recognizable distance information, but also lays an accurate data foundation for the subsequent inter-pulse Fourier transform. By capturing accurate distance unit information, the ability to analyze distance components during target detection is effectively improved, and the natural transition between time information and frequency domain information is achieved in the internal logic, which promotes the guarantee of data consistency and accuracy at each stage in the overall signal processing process.

[0061] Step 103: Perform an inter-pulse Fourier transform on the complex amplitude matrix according to the following formula:

[0062] ;

[0063] in, Indicates the The distance unit The complex amplitude of frequency units, is the retrieval variable of the distance unit, , , is the total number of frequency units, , is the total number of pulses in the echo signal, Indicates the In the pulse The complex magnitude of the distance units, Retrieve variable for distance cell number, , , is the total number of distance units, Retrieve variable for pulse number, , , is the Fourier transform base.

[0064] Indicates in The first distance unit is obtained after the multi-pulse data is subjected to the inter-pulse Fourier transform. The complex amplitude of frequency units. Reflects the pulse data obtained after the intra-pulse Fourier transform The frequency domain response after fusion in the inter-pulse dimension, where represents different pulse trains, and It corresponds to the frequency (or Doppler) unit obtained by Fourier transform between pulses. In a real environment, represents the signal energy of each pulse in the corresponding distance unit, and The corresponding frequency information is directly related to the target's motion state (Doppler effect), so In fact, the comprehensive results reveal the frequency shift changes caused by the movement of the target at different distances. is the inner dependent variable after transformation from time domain to frequency domain, and the complex exponential factor It plays the role of adjusting the phase of different pulse data to achieve coherent superposition, so that each pulse data forms a representative overall response in the frequency domain. The amplitude and phase of not only describe the characteristics of a single pulse, but their changes will directly affect the The size and phase of the pulse signal are combined to reveal the cumulative information; when some pulse signals have a strong response at a certain distance unit, through this superposition, their The contribution of will also be enhanced accordingly, showing a larger amplitude at this frequency unit. The change of the complex exponential rotation angle will adjust the phase interference effect of each pulse data, making the overall output In different The energy distribution on the surface of the target is different. This change relationship can reflect the speed or Doppler information of the target. Its amplitude varies with The response changes with the increase or decrease of the independent variable The size of the The energy accumulation effect reflects the target motion characteristics obtained by the coherent superposition of pulses.

[0065] By using the inter-pulse Fourier transform, the pulse data obtained in the previous processing step are integrated into a unified frequency shift dimension, and then the Doppler information closely related to the target movement speed is extracted. This not only enriches the dimension of the signal characteristics, allowing the distance information and speed information to be captured synchronously, but also plays an important role in connecting multiple pulse data in the entire signal processing chain. Through this conversion, the entire system can more comprehensively reflect the target characteristics, providing a solid information foundation for the subsequent generation of intuitive range-Doppler maps, while promoting the precise connection of data from a single pulse to the overall feature map.

[0066] Step 2: The complex amplitude is positively computed to form an amplitude matrix. The amplitude matrix is ​​converted into a two-dimensional range-Doppler map. At the same time, a ground truth box that reflects the positioning is manually selected. The range-Doppler map and the ground truth box are input into a convolutional neural network for feature extraction to form a feature map and feature box.

[0067] The step 2 includes the following:

[0068] Step 201: The complex amplitudes formed by the inter-pulse Fourier transform are positively computed to form an amplitude matrix according to the following formula:

[0069] ;

[0070] in, Indicates the magnitude matrix The distance unit The amplitude of a frequency unit, Indicates the The distance unit The complex amplitude of frequency units, is the retrieval variable of the distance unit, , , is the total number of frequency units, Retrieve variable for distance cell number, , , is the total number of distance units;

[0071] The complex amplitude Take the modulus to form the amplitude matrix . Characterized in the distance unit and The amplitude of the signal corresponding to each frequency unit reflects the energy distribution of the actual radar signal in these two dimensions. It is a complex value obtained by Fourier transform between pulses, which contains phase and amplitude information. Corresponding to the target's velocity information (Doppler effect), and is related to the target distance; therefore, through the modulo operation, Converting these complex signals into pure positive energy distribution data facilitates subsequent analysis and display, such as forming an intuitive range Doppler map. Contains the comprehensive information of each pulse data after coherent superposition, and its amplitude is affected by the complex exponential weighting in each previous step, the intensity of each pulse signal and phase interference; this makes directly reflects the superposition of all these factors in the target's energy distribution, so when When the amplitude of each complex variable in is large, the corresponding will also increase, and when the phases cancel each other out, When the amplitude decreases, By calculating the absolute value, the formula fully strips away the phase information, retaining only the amplitude information, achieving a transition from complex signals to intuitive energy representation and ensuring that all results are non-negative, facilitating subsequent image processing, manual calibration, and target detection analysis.

[0072] Convert the amplitude matrix to a range-Doppler map using the following logic:

[0073] The number of pixels in the range Doppler image is the same as the number of elements in the amplitude matrix. The number of the pixel points in the range Doppler image is the same as the coordinate number of the elements in the amplitude matrix. The pixel value of each pixel point in the range Doppler image is the amplitude value of the corresponding position in the amplitude matrix. The horizontal axis of the range Doppler image represents the distance, and the vertical axis of the range Doppler image represents the frequency shift.

[0074] By taking the positive value of the complex amplitude data output by the inter-pulse Fourier transform to form an amplitude matrix, and converting its units into decibels for easy visual comparison, it is further converted into a two-dimensional range Doppler map. This not only makes the data visually intuitive and easy to understand, but also ensures that the energy distribution of each distance unit and frequency shift unit is clearly presented. This process enables subsequent manual labeling and feature extraction of convolutional neural networks to be carried out under standardized and high signal-to-noise ratio data input, ensuring the consistency of physical quantity conversion in the entire signal processing chain, and achieving multiple enhancements in data accuracy, stability and intuitive expression for the entire system.

[0075] Step 203: Manually select each range Doppler image. The selection range is the part of the range Doppler image where the pixel value is significantly higher than the surrounding area. The manually determined selection box is marked as the real box, the type of the real box is marked, and the type is the category of the target. It is recorded through one-hot encoding, and the height and width of the real box are marked at the same time. The range Doppler image is input into the convolutional neural network to obtain the feature map. The convolution kernel size of the first convolutional layer of the convolutional neural network is 3x3, the number is 32, the step size is 1, and the output size is , the first pooling layer has a step size of 2 and an output size of , the convolution kernel size of the second convolution layer is 3x3, the number is 64, and the output size is , the stride of the second pooling layer is 2, and the output size is , the convolution kernel size of the third convolution layer is 3x3, the number is 128, and the output size is , the feature map size output by the final feature output layer is , where the feature map contains the true box after convolution, and the true box at this time is marked as the feature box.

[0076] As a preferred embodiment, when the complex signal is converted into a range Doppler map, the detected target will show different reaction intensities in specific frequency and distance dimensions. On the range Doppler map, the amplitude (i.e., pixel value) of the position is significantly higher than the surrounding environment. Select this location, and the center of the selection box represents the detected target, such as Figure 2 As shown, when the interference increases, the amplitude of the detected target is not prominent enough, but it can still be manually judged and identified. This process is a common technical feature used by those skilled in the art, so it will not be described in detail in this embodiment. The "real frame" referred to here in this embodiment is a rectangular frame that can reflect the detected target and is manually drawn in the range Doppler map by those skilled in the art based on their experience.

[0077] As a preferred embodiment, the convolutional neural network maintains the spatial information in the range Doppler map, so that the feature map maintained after the network has a direct correspondence with the specific physical position in the original image. For example, the center position of the target in the original image will also appear in the same relative position in the feature map. Therefore, on the one hand, the deep convolution layer can capture local features such as low-level edges and textures, and evolve layer by layer into high-level abstract features that can represent the real frame target. On the other hand, these data designed after convolution, pooling and feature mapping facilitate the subsequent steps of candidate frame arrangement, positive and negative sample judgment and center offset correction for accurate target classification and positioning.

[0078] By manually selecting the range Doppler map, calibrating the target area (i.e., the true box) and recording the target category and box size information, an accurate supervision signal is provided for the subsequent convolutional neural network to perform feature extraction. Manual box selection ensures that the target area that is obvious in the data can be accurately captured. At the same time, the target category is annotated using one-hot encoding to make the label information clear and unambiguous during network training. This process not only refines the position expression of the target in the image, but also provides an accurate reference for the subsequent matching between candidate boxes and feature boxes and sample generation, realizing an effective transition from raw signals to feature extraction, and laying a solid foundation for the accuracy of the overall target detection and positioning system.

[0079] Step 3: Three categories of candidate boxes are formed based on the area of ​​the feature box and arranged on the feature map. The intersection-over-union (IoU) of each candidate box and the feature box is obtained. All candidate boxes are divided into positive samples and negative samples according to the IoU threshold and input into the fully connected neural network for training to obtain the category discrimination model.

[0080] The step 3 includes the following:

[0081] Step 301: Set the anchor box on the feature map. The logic is as follows:

[0082] Get the feature frame with the largest area, calibrate its width and height as the first candidate frame, and reduce it according to its width and height, reducing the width and height by one-third respectively. If the width or height after reduction is not an integer, round it up. The frame formed by the width and height at this time is calibrated as the second candidate frame. Reduce the width and height of the first candidate frame by two-thirds. If the width or height after reduction is not an integer, round it up. The frame formed by the width and height at this time is calibrated as the third candidate frame;

[0083] The three types of candidate boxes are arranged on the feature map respectively. The arrangement logic of the candidate boxes of each type is as follows:

[0084] The candidate box is first located in the upper left corner to form the first position, and then it is translated one grid to the right to form the second position. In this way, the first row of candidate boxes is formed, and the rightmost candidate box is located in the upper right corner. Then all the candidate boxes in the first row are translated down one grid to form the second row of candidate boxes, and then translated down one grid to form the third row of candidate boxes. In this way, an arrangement that completely covers the feature map is formed.

[0085] By setting anchor boxes on the feature map, the first candidate box is generated using the feature box with the largest area, and then the third and second candidate boxes are reduced respectively. The three types of candidate boxes are evenly covered in the entire feature map according to a certain arrangement logic. This not only takes into account the diversity of different target sizes, but also effectively captures target information of different scales and positions. The systematic arrangement of candidate boxes can fully ensure that every area in the image has the possibility of being detected, providing sufficient alternative data for the subsequent positive and negative sample judgment based on the intersection-over-union ratio. At the same time, it also plays a key role in matching each candidate box with the real target box, closely combining global feature information with local target features, thereby improving the detection accuracy and robustness of the category discrimination model.

[0086] Step 302: In a single feature map, determine the IoU of each candidate box and the feature box respectively, set an IoU threshold, calibrate the candidate boxes and the features within the boxes that are greater than or equal to the IoU threshold as positive samples, calibrate the candidate boxes and the features within the boxes that are lower than the IoU threshold as negative samples, mark the positive samples with the labels of the true boxes to which they belong, and mark the negative samples as background, summarize all positive and negative samples and the corresponding labels, train them in a fully connected neural network, and obtain a category discrimination model.

[0087] As a preferred embodiment, the discrimination model is essentially to classify feature images through a fully connected neural network model, which is a common technical feature in this field, such as the image classification method, device, computer equipment and storage medium with publication number CN111159450A, which has disclosed the purpose of achieving classification by generating judgment probabilities through image features through a fully connected neural network model. Therefore, the construction of this model can adopt the same construction method, which will not be repeated here.

[0088] By calculating the intersection-over-union ratio between the candidate box and the feature box, and dividing the candidate box into positive samples (i.e., target candidate boxes) and negative samples (background candidate boxes) based on the set threshold, these samples and their labels are then input into the fully connected neural network for training. This not only achieves fine screening of candidate areas, but also enhances the network's ability to distinguish between targets and backgrounds. This process not only provides clear classification training labels in local areas, ensuring the minimization of target detection and positioning errors during network learning, but also overall connects the early candidate box setting and subsequent positioning process, laying a solid data foundation for the rapid and accurate judgment of the category discrimination model, and improving the adaptability and detection robustness of the entire system in complex scenarios.

[0089] Step 4: Obtain the candidate box of the result obtained in the category discrimination model, mark it as the target box, use the center offset vector between the target box and the feature box as the label, input the target box and internal features into the fully connected neural network model for training, and obtain the positioning model;

[0090] The step 4 includes the following:

[0091] Obtain the candidate box of the judgment result obtained in the output layer of the category discrimination model and mark it as the target box. The logic of obtaining the candidate box of the judgment result is:

[0092] A single feature map has multiple positive samples. Each positive sample will output the probability of this sample on each target category in the fully connected layer in the category discrimination model. The judgment result of this discrimination model is obtained, and the positive sample with the highest probability on the target category consistent with the judgment result is obtained. The candidate box corresponding to this positive sample is the target box;

[0093] The center offset vectors of the target box and the feature box in each feature map are compared respectively. Each feature map is labeled with the corresponding center offset. The target box and internal features are input into the fully connected neural network model for training to obtain the positioning model.

[0094] As a preferred embodiment, for each target frame, the convolution features in the frame are converted into a one-dimensional vector through pooling measures. The one-dimensional vector structure is closely related to the energy distribution, local pattern and other features of the target in the radar echo, which is the key data for subsequent center offset prediction. The center offset is horizontally displaced. and vertical displacement Indicates that the input layer dimension is the same as the one-dimensional vector dimension, the number of hidden layers is three, the number of neurons in each layer is 128, the activation function uses the ReLU function, the output layer is a fully connected layer, and the output nodes are two, respectively. and ,The loss function adopts the mean square error, and the Adam optimizer default settings are used to optimize the fully connected neural network, with an initial learning rate of 0.001 and 86 training iterations. In each training iteration, the network updates the parameters through back propagation, thereby continuously narrowing the gap between the predicted value and the true offset.

[0095] By obtaining the judgment results obtained by the output layer of the category discrimination model, the final target frame is further determined, and the center offset vectors of the target frame and the feature frame are compared, so that the positioning model is established using the position information to ensure the accuracy of the target frame position in actual detection; this process realizes the effective connection between target detection and fine positioning, and combines the probability information obtained through the fully connected neural network with the spatial position information. It not only corrects the possible position information deviation, but also improves the robustness of the overall detection system in precise target positioning, so that the center position of each target can be reasonably mapped and accurately calibrated in the entire image, creating technical conditions and data support for the subsequent high-precision target tracking in practical applications.

[0096] Step 5: The radar acquires echoes in real time and forms a feature map. Three types of candidate boxes are arranged in the feature map and input into the discrimination model to determine the target category. The target box and internal features are obtained and input into the positioning model. The center offset vector and positioning identification box are obtained to realize target positioning.

[0097] The step 5 includes the following contents:

[0098] Acquire the echo through the radar, and obtain the range Doppler map through step 1. Input the convolutional neural network in step 2 to form a feature map. Arrange three types of candidate boxes in the feature map. Input each candidate box and the internal feature box into the discrimination model to determine the target category and obtain the target box. Input the target box and the internal features into the positioning model to obtain the center offset. Adjust the target box according to the center offset vector to obtain the positioning recognition box. The positioning recognition box is restored according to the convolutional neural network to obtain the distance point. The restoration logic is as follows:

[0099] In the process of converting the range Doppler map into a feature map in the convolutional neural network, the size of the range Doppler map is changed, but the internal relative relationship does not change. The relative position of the fixed point in the feature map is the same as the relative position in the range Doppler map. Therefore, according to the relative position of the center point in the target frame in the feature map, a specific point in the range Doppler map is corresponding to this point, and this point is marked as a range point.

[0100] The distance to the target is obtained based on the position of the range point in the range-Doppler map.

[0101] As a preferred embodiment, judging the distance to the target by a specific point in the distance Doppler map is a common technical feature in this field. The method for calculating the distance is clearly stated in "Radar Principles and Applications" published by Higher Education Press, which will not be described in detail here.

[0102] The real-time echo signal obtained from the radar, the feature map extracted by the convolutional neural network, and the target frame adjusted by the candidate frame classification and positioning model are comprehensively utilized. Finally, the actual distance of the target is determined by mapping the adjusted positioning recognition frame back to the original range Doppler map, forming a closed-loop detection and positioning system. This process not only ensures the intrinsic spatial relationship and data consistency of the signal after passing through each processing module, but also organically connects each processing link, realizing a seamless connection from the original radar signal to high-precision target detection, classification and positioning, thereby improving the real-time performance and accuracy of the system. In the overall solution, it enhances the comprehensive performance and engineering practical value of deep learning methods in radar target detection applications.

[0103] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.

[0104] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed by hardware or software depends on the specific application and design constraints of the technical solution.

[0105] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, and may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment as needed.

[0106] The above is only a specific implementation method of the present application, but the scope of protection of the present application is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in this application, which should be covered by the scope of protection of the present application.

Claims

1. A target recognition and positioning method based on deep learning, characterized in that: The specific steps include: Step 1: Acquire the echo signal and convert it into a complex signal matrix and perform Fourier transform to obtain the complex amplitude of the complex signal matrix for each frequency unit of each distance unit; Step 2: The complex amplitude is positively computed to form an amplitude matrix. The amplitude matrix is ​​converted into a two-dimensional range-Doppler map. At the same time, a ground truth box that reflects the positioning is manually selected. The range-Doppler map and the ground truth box are input into a convolutional neural network for feature extraction to form a feature map and feature box. Convert the amplitude matrix to a range-Doppler map using the following logic: The number of pixels in the range-Doppler image is the same as the number of elements in the amplitude matrix, the number of the pixel points in the range-Doppler image is the same as the coordinate number of the element in the amplitude matrix, the pixel value of each pixel point in the range-Doppler image is the amplitude value of the corresponding position in the amplitude matrix, the horizontal axis of the range-Doppler image represents the distance, and the vertical axis of the range-Doppler image represents the frequency shift; Step 3: Three categories of candidate boxes are formed based on the area of ​​the feature box and arranged on the feature map. The intersection-over-union (IoU) of each candidate box and the feature box is obtained. All candidate boxes are divided into positive samples and negative samples according to the IoU threshold and input into the fully connected neural network for training to obtain the category discrimination model. Step 4: Obtain the candidate box of the result obtained in the category discrimination model, mark it as the target box, use the center offset vector between the target box and the feature box as the label, input the target box and internal features into the fully connected neural network model for training, and obtain the positioning model; Step 5: The radar acquires echoes in real time and forms a feature map. Three types of candidate boxes are arranged in the feature map and input into the discrimination model to determine the target category. The target box and internal features are obtained and input into the positioning model. The center offset vector and positioning identification box are obtained to realize target positioning.

2. The target recognition and positioning method based on deep learning according to claim 1, characterized in that: Obtain historical data of the radar, including received echo signals and the target category identified for each echo signal, where the target category is the type of aircraft detectable by the radar. Obtain the sampling rate, number of single pulse sampling points, and number of pulses of the radar. Mix the echo signals to obtain the corresponding complex signal. Filter the complex signal using a fixed-bandwidth filter of the radar. Weight the sampling points within each pulse using a Hamming window function. Normalize the weighted complex signal to obtain a complex signal matrix, based on the following formula: in, is the complex signal matrix, Represents the first Pulse No. The signal value of the sampling point, Retrieve variable for pulse number, , , is the total number of pulses in the echo signal, Retrieve variable for the number of sample points in a single pulse, , , is the number of sampling points for a single pulse.

3. The target recognition and positioning method based on deep learning according to claim 2, characterized in that: Perform an intra-pulse Fourier transform on each pulse in the complex signal matrix. A row of the complex signal matrix represents a pulse. The intra-pulse Fourier transform is based on the following formula: in, Indicates the In the pulse The complex magnitude of the distance units, Represents the first In the pulse The signal value of the sampling point, Retrieve variable for pulse number, , , is the total number of pulses in the echo signal, Retrieve variable for distance cell number, , , is the total number of distance units, , Retrieve variable for the number of sample points in a single pulse, , , is the number of sampling points for a single pulse, is the Fourier transform basis function.

4. The method for target recognition and positioning based on deep learning according to claim 3, characterized in that: The complex amplitude matrix is ​​subjected to the inter-pulse Fourier transform according to the following formula: in, Indicates the The distance unit The complex amplitude of frequency units, is the retrieval variable of the distance unit, , , is the total number of frequency units, , is the total number of pulses in the echo signal, Indicates the In the pulse The complex magnitude of the distance units, Retrieve variable for distance cell number, , , is the total number of distance units, Retrieve variable for pulse number, , , is the Fourier transform base.

5. The method for target recognition and positioning based on deep learning according to claim 4, characterized in that: The complex amplitudes formed by the Fourier transform between pulses are positively calculated to form an amplitude matrix, based on the following formula: in, Indicates the magnitude matrix The distance unit The amplitude of a frequency unit, Indicates the The distance unit The complex amplitude of frequency units, is the retrieval variable of the distance unit, , , is the total number of frequency units, Retrieve variable for distance cell number, , , is the total number of distance units.

6. The method for target recognition and positioning based on deep learning according to claim 5, characterized in that: Each range Doppler image is manually framed. The frame selection range is the part of the range Doppler image where the pixel value is significantly higher than the surrounding area. The manually determined frame is marked as the real frame, the type of the real frame is marked, and the type is the category of the target. It is recorded through unique hot encoding, and the height and width of the real frame are marked at the same time. The range Doppler image is input into the convolutional neural network to obtain the feature map. The convolution kernel size of the first convolutional layer of the convolutional neural network is 3x3, the number is 32, the step size is 1, and the output size is , the first pooling layer has a step size of 2 and an output size of , the convolution kernel size of the second convolution layer is 3x3, the number is 64, and the output size is , the stride of the second pooling layer is 2, and the output size is , the convolution kernel size of the third convolution layer is 3x3, the number is 128, and the output size is , the feature map size output by the final feature output layer is , where the feature map contains the true box after convolution, and the true box at this time is marked as the feature box.

7. The method for target recognition and positioning based on deep learning according to claim 6, characterized in that: The logic for setting the anchor box on the feature map is as follows: Get the feature frame with the largest area, calibrate its width and height as the first candidate frame, and reduce it according to its width and height, reducing the width and height by one-third respectively. If the width or height after reduction is not an integer, round it up. The frame formed by the width and height at this time is calibrated as the second candidate frame. Reduce the width and height of the first candidate frame by two-thirds. If the width or height after reduction is not an integer, round it up. The frame formed by the width and height at this time is calibrated as the third candidate frame; The three types of candidate boxes are arranged on the feature map respectively. The arrangement logic of the candidate boxes of each type is as follows: The candidate box is first located in the upper left corner to form the first position, and then it is translated one grid to the right to form the second position. In this way, the first row of candidate boxes is formed, and the rightmost candidate box is located in the upper right corner. Then all the candidate boxes in the first row are translated down one grid to form the second row of candidate boxes, and then translated down one grid to form the third row of candidate boxes. In this way, an arrangement that completely covers the feature map is formed.

8. The method for target recognition and positioning based on deep learning according to claim 7, characterized in that: In a single feature map, the intersection-and-union (IoU) of each candidate box and the feature box is determined separately, and an IoU threshold is set. Candidate boxes and their in-box features that are greater than or equal to the IoU threshold are calibrated as positive samples, and candidate boxes and their in-box features that are lower than the IoU threshold are calibrated as negative samples. Positive samples are labeled with the labels of the true boxes to which they belong, and negative samples are labeled as background. All positive and negative samples and their corresponding labels are summarized and trained in a fully connected neural network to obtain a category discrimination model.

9. The method for target recognition and positioning based on deep learning according to claim 8, characterized in that: Obtain the candidate box of the judgment result obtained in the output layer of the category discrimination model and mark it as the target box. The logic of obtaining the candidate box of the judgment result is: A single feature map has multiple positive samples. Each positive sample will output the probability of this sample on each target category in the fully connected layer in the category discrimination model. The judgment result of this discrimination model is obtained, and the positive sample with the highest probability on the target category consistent with the judgment result is obtained. The candidate box corresponding to this positive sample is the target box; The center offset vectors of the target box and the feature box in each feature map are compared respectively. Each feature map is labeled with the corresponding center offset. The target box and internal features are input into the fully connected neural network model for training to obtain the positioning model.

10. The method for target recognition and positioning based on deep learning according to claim 9, characterized in that: Acquire the echo through the radar, and obtain the range Doppler map through step 1. Input the convolutional neural network in step 2 to form a feature map. Arrange three types of candidate boxes in the feature map. Input each candidate box and the internal feature box into the discrimination model to determine the target category and obtain the target box. Input the target box and the internal features into the positioning model to obtain the center offset. Adjust the target box according to the center offset vector to obtain the positioning recognition box. The positioning recognition box is restored according to the convolutional neural network to obtain the distance point. The restoration logic is as follows: In the process of converting the range Doppler map into a feature map in the convolutional neural network, the size of the range Doppler map is changed, but the internal relative relationship does not change. The relative position of the fixed point in the feature map is the same as the relative position in the range Doppler map. Therefore, according to the relative position of the center point in the target frame in the feature map, a specific point in the range Doppler map is corresponding to this point, and this point is marked as a range point. The distance to the target is obtained based on the position of the range point in the range-Doppler map.