Radar-assisted single target tracking and holder linkage control method

By performing denoising and time-frequency analysis of radar signals, combining convolutional neural networks and LSTM models for target detection and trajectory prediction, and using deep reinforcement learning to optimize the gimbal control strategy, the problems of misdetection and missed detection in the existing technology are solved, and the robustness and accuracy of target tracking are improved.

CN119986632AInactive Publication Date: 2025-05-13CHENGDU SHUXI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510104565.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the face of complex signal noise or rapid movement of the target, the prior art is prone to misdetection and misdetection, which affects the accuracy of target trajectory prediction.

Method used

The radar signal data is denoised and enhanced by wavelet transformation and adaptive filtering algorithms, and time-frequency domain analysis is performed by combining short-time Fourier transform and continuous wavelet transformation to extract the target's time-frequency characteristic data. Then, object detection is performed using a convolutional neural network and multimodal data fusion model, trajectory prediction is performed using the LSTM model, and correction is performed by Kalman filter. Finally, deep reinforcement learning algorithm is used to optimize the control strategy of the gimbal.

Benefits of technology

It significantly improves the robustness and accuracy of radar-assisted target tracking in complex environments, reduces the risk of misdetection and misdetection, and ensures that the gimbal can quickly respond to the complex dynamic behavior of the target.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119986632A_ABST
    Figure CN119986632A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of holder control, and relates to a radar-assisted single target tracking and holder linkage control method. Denoising and enhancing the radar signal data through wavelet transform and an adaptive filtering algorithm; performing time-frequency domain analysis by using short-time Fourier transform and continuous wavelet transform, and extracting more accurate target time-frequency characteristic data; by combining a convolutional neural network with a multi-modal data fusion model, accurately extracting related information of the target from the time-frequency domain feature data; learning a historical track of the target through an LSTM model, predicting a future motion track of the target, and correcting and updating a prediction result through a Kalman filter; dynamically optimizing a control strategy of the holder according to the updated state of the target and the current state of the holder by using a deep reinforcement learning algorithm; the problem of inaccurate control caused by false detection, missing detection and rapid movement of the target is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of pan-tilt control technology, and more specifically, to a radar-assisted single target tracking and pan-tilt linkage control method. Background Art

[0002] In modern military, aviation, drone control and other fields, accurate target tracking and pan-tilt linkage control technology are crucial, especially in dynamic and complex environments. Although traditional target tracking technology can provide a certain degree of accuracy in static or relatively simple environments, when faced with situations where the target moves at a fast speed, has an irregular trajectory or is subject to external interference, the existing technology still faces great challenges in terms of accuracy, robustness and real-time performance.

[0003] The radar signal acquisition process is affected by many factors, including signal noise, environmental interference, weather conditions, etc., which makes it very difficult to extract accurate information about the target from the original radar signal. Traditional target tracking methods may suffer from false detection, missed detection, or tracking loss when faced with complex signal noise or rapid target movement, greatly affecting the accuracy of target trajectory prediction.

[0004] In addition, due to the strong uncertainty of the target's motion, especially in some cases of rapid speed changes or turns, relying solely on classic Kalman filtering or traditional tracking algorithms is often difficult to cope with complex scenarios, and there are problems such as trajectory estimation error accumulation and difficulty in error correction. Summary of the invention

[0005] The present invention provides a radar-assisted single target tracking and pan-tilt linkage control method, which is intended to solve the current technical problems of possible misdetection and missed detection when facing complex signal noise or rapid target movement.

[0006] The radar-assisted single target tracking and pan / tilt linkage control method comprises the following steps:

[0007] Step 1: Collect the original radar signal data, and denoise and enhance the collected radar signal data based on wavelet transform and adaptive filtering algorithm; then, based on the denoised and enhanced radar signal data, use short-time Fourier transform and continuous wavelet transform to perform time-frequency domain analysis to obtain time-frequency domain feature data;

[0008] Step 2: Use the obtained time-frequency domain feature data and the multimodal data fusion model based on convolutional neural network to analyze the time-frequency domain feature data to obtain the target detection information, including distance, speed, acceleration, azimuth and type;

[0009] Step 3: Based on the target detection information and the target historical trajectory information, the LSTM model is used to predict the target trajectory to obtain the target's future position and trajectory. The Kalman filter is then used to correct and update the target position in combination with the target's future position and trajectory to obtain the updated target position and state estimate.

[0010] Step 4: Use a deep reinforcement learning algorithm to observe the state of the target and the state of the gimbal based on the updated target position and state estimation and the current gimbal state, perform strategy learning, and optimize the gimbal control strategy;

[0011] Step 5: The gimbal control system generates corresponding control instructions based on the optimized control strategy to adjust the gimbal movement; and updates the actual state of the gimbal and feeds it back to the deep reinforcement learning algorithm for further learning and optimization.

[0012] In the present invention, the radar signal data is denoised and enhanced by wavelet transform and adaptive filtering algorithm, which effectively reduces the interference of noise; short-time Fourier transform and continuous wavelet transform are used for time-frequency domain analysis to extract more accurate target time-frequency feature data; convolutional neural network (CNN) is combined with multimodal data fusion model to accurately extract relevant information of the target from the time-frequency domain feature data, including distance, speed, acceleration, azimuth, etc., thereby improving the accuracy of target detection and reducing the risk of false detection and missed detection; the historical trajectory of the target is learned by LSTM model to predict the future motion trajectory of the target, and the prediction result is corrected and updated by Kalman filter, thereby achieving more accurate target detection and less error detection and missed detection. Accurate target position estimation; using a deep reinforcement learning algorithm, dynamically optimize the control strategy of the gimbal according to the target's update status and the current state of the gimbal, to ensure that the gimbal can respond quickly to the rapid changes of the target, thereby effectively coping with the complex dynamic behavior of the target; the gimbal control system generates control instructions according to the optimized control strategy, accurately adjusts the gimbal position, and feeds back the actual state of the gimbal to the deep reinforcement learning model to form a closed-loop learning, continuously improving the system's adaptability and control accuracy; based on this, the present invention significantly improves the robustness and accuracy of radar-assisted target tracking in complex environments, and effectively solves the problems of inaccurate control caused by false detection, missed detection, and rapid target movement.

[0013] Preferably, the specific steps of denoising and enhancing the radar signal are as follows:

[0014] Wavelet transform: Discrete wavelet transform is used to decompose the original radar signal data into low-frequency components and high-frequency components;

[0015] Local noise estimation: A local noise estimation method is used to calculate the noise level of each component that is heavier than the high-frequency component;

[0016] Soft threshold denoising: Soft threshold processing is performed on the high-frequency components after noise estimation:

[0017] D k [n] denoised =sgn(D k [n])·max(|D| k [n]|-λ k ,0);

[0018]

[0019] Where: represents the wavelet threshold of the kth layer; γ represents the constant factor; represents the noise estimate obtained by wavelet decomposition at the kth layer; D k [n] represents the high-frequency component of the k-th wavelet decomposition; D k [n] denoised Represents the high-frequency component after soft threshold denoising;

[0020] Adaptive filter output:

[0021] y[n]=w T [n]·D k [n] denoised ;

[0022] Where: w[n] represents the weight vector of the adaptive filter, which is a parameter that changes with time and is used to adjust the weight vector according to the input signal D k [n] denoised To adjust the output; w T [n]·D k [n] denoised represents the dot product operation, which represents the linear combination of the weight vector and the input signal;

[0023] Error calculation:

[0024] e[n]=D k [n] denoised -y[n];

[0025] Where: e[n] represents the error of the adaptive filter; y[n] represents the output of the filter;

[0026] Weight update:

[0027] w[n]+1]=w[n]+μ k ·e[n]·D k [n] denoised ;

[0028] Where: w[n+1] represents the updated weight vector; μ krepresents the step size factor; w[n] represents the weight vector before updating;

[0029] Updated high frequency section:

[0030] D k [n] denoised,updated =D k [n] denoised -w T [n]·D k [n] denoised ;

[0031] Where: D k [n] denoised,updated Represents the high-frequency component after adaptive filtering;

[0032] The final signal synthesis:

[0033] y[n]'=A k [n]+D k [n] denoised,updated ;

[0034] Where: y[n]' represents the final denoised signal; A k [n] represents the low-frequency component obtained by wavelet transform.

[0035] Preferably, the compensation factor is dynamically adjusted based on the noise estimation value obtained by the k-th layer wavelet decomposition:

[0036]

[0037] Where: α represents the adjustment factor; Represents the noise estimate obtained by the k-th layer wavelet decomposition.

[0038] Preferably, the specific steps of performing time-frequency domain analysis based on the denoised and enhanced radar signal data are as follows:

[0039] Short-time Fourier transform: The denoised and enhanced radar data is divided into several small blocks, and Fourier transform is performed in each time window through a Gaussian window to obtain a two-dimensional complex matrix containing the amplitude and phase information of the signal at different time points and frequency points;

[0040] Instantaneous frequency extraction of short-time Fourier transform: Calculate the instantaneous frequency of the signal based on the obtained two-dimensional complex matrix:

[0041]

[0042] Where: arg(X(τ,f)) represents the signal phase; f inst (τ) represents the instantaneous frequency of the short-time Fourier transform; X(τ,f) represents the two-dimensional complex matrix output by the short-time Fourier transform; represents the derivative with respect to the time parameter τ;

[0043] Instantaneous amplitude extraction: the loss amplitude is the amplitude value of the short-time Fourier transform amplitude spectrum;

[0044] Wavelet transform: Continuous wavelet transform is used to perform time-frequency analysis on the denoised and enhanced radar data to obtain wavelet transform coefficients;

[0045] Frequency feature extraction: Based on the wavelet transform coefficients obtained by continuous wavelet transform, the frequency response at different scales is calculated to obtain the frequency features;

[0046] Instantaneous frequency of wavelet transform: Based on the time-frequency spectrum of wavelet transform, the instantaneous frequency is obtained by analyzing the wavelet transform phase transformation of the signal:

[0047]

[0048] Where: represents the instantaneous frequency of wavelet transform; arg(W y[n]' (a, b)) represents the phase part of wavelet transform; represents the derivative of the phase with respect to the translation parameter b;

[0049] Energy spectrum extraction: The energy spectrum of the signal is obtained by squaring and summing the wavelet transform coefficients of different scales;

[0050] Spectrum center frequency: Extract spectrum center frequency and spectrum bandwidth based on Fourier transform results;

[0051] Integration of time-frequency domain features: The extracted features are integrated into a time-frequency domain feature vector, where the time-frequency domain feature vector includes: instantaneous frequency and instantaneous amplitude of short-time Fourier transform, instantaneous frequency, energy spectrum, spectrum center frequency and spectrum bandwidth of wavelet transform.

[0052] Preferably, the multimodal data fusion model includes an input layer, a time domain feature separation module, a convolutional neural network module, a self-attention mechanism, a multi-layer perceptron fusion module and an output layer;

[0053] The input layer is used to input the extracted time-frequency domain features;

[0054] The time domain feature separation module establishes an independent input channel for the features extracted by short-time Fourier transform and continuous wavelet transform, and each input channel is sent to a convolutional neural network module for feature extraction;

[0055] The convolutional neural network module includes multiple convolutional layers and multiple pooling layers. The features of each input channel sent to the convolutional neural network module are processed to output the local feature representation of each modality;

[0056] The self-attention mechanism performs weighted fusion based on the output of the convolutional neural network module to obtain a weighted feature representation of each convolutional neural network;

[0057] The multi-layer perceptron fusion module fuses all feature representations weighted by the self-attention mechanism and outputs final target information, including target detection information, wherein the target detection information includes distance, speed, acceleration, azimuth and type.

[0058] Preferably, step 3 obtains an updated target state estimate by fusing the prediction result of the LSTM model and the prediction state of the Kalman filter:

[0059]

[0060] Where: represents the predicted state of the Kalman filter; represents the predicted value of the LSTM model; H represents the measurement matrix; ω represents a weight factor between 0 and 1; represents the updated target position and state estimate.

[0061] Preferably, the deep reinforcement learning algorithm in step 4 is as follows:

[0062] State definition: define the state vector Where P t represents the updated target position estimate; represents the updated target state estimate; C t Indicates the position of the PTZ; Indicates the speed status of the gimbal;

[0063] Action space definition: Use the gimbal angle and angular velocity as the control action A t =[Δθ,Δv θ ], where Δθ represents the change in the gimbal angle; Δv θ Indicates the change in the angular velocity of the gimbal;

[0064] Reward function:

[0065]

[0066] Where: ‖e P ‖ represents the target position error; represents the target velocity error; |Δθ| and |Δv θ | represents the change in the angle and angular velocity of the gimbal; α, β, γ, δ represent weight coefficients;

[0067] in:

[0068] eP =P t -C t ;

[0069]

[0070] Where: k1 and k2 represent control gains, which adjust the weights of position error and speed error respectively; k3 and k4 represent control gains; express rate of change over time;

[0071] Q value update: Update the Q value function Q(S) based on the Q-learning algorithm t ,A t ):

[0072]

[0073] Where: R t represents the reward at the current moment; η represents the learning rate; γ' represents the discount factor; A' represents the next action;

[0074] The control action of the gimbal is selected based on the deep reinforcement learning algorithm, and the control action is sent to the gimbal control system to adjust the angle and angular velocity of the gimbal.

[0075] Preferably, the control action of the pan / tilt is selected by using an ε-greedy strategy to select the optimal action from the Q value, according to S t And the corresponding Q value function, select the action corresponding to the maximum Q value.

[0076] The beneficial effects of the present invention include:

[0077] In the present invention, the radar signal data is denoised and enhanced by wavelet transform and adaptive filtering algorithm, which effectively reduces the interference of noise; short-time Fourier transform and continuous wavelet transform are used for time-frequency domain analysis to extract more accurate target time-frequency feature data; convolutional neural network (CNN) is combined with multimodal data fusion model to accurately extract relevant information of the target from the time-frequency domain feature data, including distance, speed, acceleration, azimuth, etc., thereby improving the accuracy of target detection and reducing the risk of false detection and missed detection; the historical trajectory of the target is learned by LSTM model to predict the future motion trajectory of the target, and the prediction result is corrected and updated by Kalman filter, thereby achieving more accurate target detection and less error detection and missed detection. Accurate target position estimation; using a deep reinforcement learning algorithm, dynamically optimize the control strategy of the gimbal according to the target's update status and the current state of the gimbal, to ensure that the gimbal can respond quickly to the rapid changes of the target, thereby effectively coping with the complex dynamic behavior of the target; the gimbal control system generates control instructions according to the optimized control strategy, accurately adjusts the gimbal position, and feeds back the actual state of the gimbal to the deep reinforcement learning model to form a closed-loop learning, continuously improving the system's adaptability and control accuracy; based on this, the present invention significantly improves the robustness and accuracy of radar-assisted target tracking in complex environments, and effectively solves the problems of inaccurate control caused by false detection, missed detection, and rapid target movement. BRIEF DESCRIPTION OF THE DRAWINGS

[0078] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0079] Figure 1 An overall step block diagram provided for an embodiment of the present invention.

[0080] Figure 2 A schematic diagram of the structure of a multimodal data fusion model provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0081] In order to make the technical problems, technical solutions and beneficial effects to be solved by the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0082] See also Figure 1 As shown, the video content intelligent analysis and statistical method supported by the large model includes the following steps:

[0083] The radar-assisted single target tracking and pan / tilt linkage control method comprises the following steps:

[0084] Step 1: Collect the original radar signal data, and denoise and enhance the collected radar signal data based on wavelet transform and adaptive filtering algorithm; then, based on the denoised and enhanced radar signal data, use short-time Fourier transform and continuous wavelet transform to perform time-frequency domain analysis to obtain time-frequency domain feature data;

[0085] The specific steps of denoising and enhancing the radar signal are as follows:

[0086] Wavelet transform: Discrete wavelet transform is used to decompose the original radar signal data into low-frequency components and high-frequency components;

[0087]

[0088] Where: A k [n] represents the low-frequency part obtained by the k-th wavelet decomposition; D k [n] represents the high-frequency component of the k-th wavelet decomposition; L represents the number of wavelet decomposition layers;

[0089] Local noise estimation: A local noise estimation method is used to calculate the noise level of each component that is heavier than the high-frequency component;

[0090]

[0091] Where: N represents the number of sample points; D k [n] represents the high-frequency component of the k-th wavelet decomposition; represents the noise estimate obtained by the k-th layer wavelet decomposition;

[0092] Soft threshold denoising: Soft threshold processing is performed on the high-frequency components after noise estimation:

[0093] D k [n] denoised =sgn(D k [n])·max(|D k [n]|-λ k ,0);

[0094]

[0095] Where: represents the wavelet threshold of the kth layer; γ represents the constant factor; represents the noise estimate obtained by wavelet decomposition at the kth layer; D k [n] represents the high-frequency component of the k-th wavelet decomposition; D k [n] denoised Represents the high-frequency component after soft threshold denoising;

[0096] Adaptive filter output:

[0097] y[n]=w T [n]·D k [n] denoised ;

[0098] Where: w[n] represents the weight vector of the adaptive filter, which is a parameter that changes with time and is used to adjust the weight vector according to the input signal D k [n] denoised To adjust the output; w T [n]·D k [n] denoised represents the dot product operation, which represents the linear combination of the weight vector and the input signal;

[0099] Error calculation:

[0100] e[n]=D k [n] denoised -y[n];

[0101] Where: e[n] represents the error of the adaptive filter; y[n] represents the output of the filter;

[0102] Weight update:

[0103] w[n+1]=w[n]+μ k ·e[n]·D k [n] denoised ;

[0104] Where: w[n+1] represents the updated weight vector; μ k represents the step size factor; w[n] represents the weight vector before updating;

[0105] Updated high frequency section:

[0106] D k [n] denoised,updated =D k [n] denoised -w T [N]·D k [n] denoised ;

[0107] Where: D k [n] denoised,updated Represents the high-frequency component after adaptive filtering;

[0108] The final signal synthesis:

[0109] y[n]'=A k [n]+D k [n]denoised,updated ;

[0110] Where: y[n]' represents the final denoised signal; A k [n] represents the low-frequency component obtained by wavelet transform.

[0111] The compensation factor is dynamically adjusted based on the noise estimate obtained by wavelet decomposition of the kth layer:

[0112]

[0113] Where: α represents the adjustment factor; Represents the noise estimate obtained by the k-th layer wavelet decomposition.

[0114] The specific steps for time-frequency domain analysis based on denoised and enhanced radar signal data are as follows:

[0115] Short-time Fourier transform: Divide the signal y[n]' into several small blocks and pass it through a Gaussian window to perform Fourier transform in each time window. Set the window function to w[n], the window function length to L, and the starting position of the window to n0:

[0116]

[0117] Where: X(τ, f) represents the short-time frequency of the signal y[n]' at time τ and frequency f, that is, the two-dimensional complex matrix output by the short-time Fourier transform; w[n-τ] is the window function, τ represents the center position of the window; f represents the frequency; n represents the time index; X(τ, f) is a two-dimensional complex matrix, which contains the amplitude and phase information of the signal at different time points and frequency points;

[0118] Instantaneous frequency extraction of short-time Fourier transform: Calculate the instantaneous frequency of the signal based on the obtained two-dimensional complex matrix:

[0119]

[0120] Where: arg(X(τ,f)) represents the signal phase; f inst (τ) represents the instantaneous frequency of the short-time Fourier transform; X(τ,f) represents the two-dimensional complex matrix output by the short-time Fourier transform; represents the derivative with respect to the time parameter τ;

[0121] Instantaneous amplitude extraction: The instantaneous amplitude is the amplitude value of the short-time Fourier transform amplitude spectrum;

[0122] Wavelet transform: Continuous wavelet transform is used to perform time-frequency analysis on the denoised and enhanced radar data to obtain wavelet transform coefficients;

[0123]

[0124] Where: W y[n]' (a, b) represents the coefficients of wavelet transform, which reflects the local time-frequency characteristics of the signal at scale a and time b; represents the conjugate of the complex mother wavelet, a represents the scale factor, and b represents the time position;

[0125] Frequency feature extraction: Based on the wavelet transform coefficients obtained by continuous wavelet transform, the frequency response at different scales is calculated to obtain the frequency features;

[0126] Instantaneous frequency of wavelet transform: Based on the time-frequency spectrum of wavelet transform, the instantaneous frequency is obtained by analyzing the wavelet transform phase transformation of the signal:

[0127]

[0128] Where: represents the instantaneous frequency of wavelet transform; arg(W y[n]' (a, b)) represents the phase part of wavelet transform; represents the derivative of the phase with respect to the translation parameter b;

[0129] Energy spectrum extraction: The energy spectrum of the signal is obtained by squaring and summing the wavelet transform coefficients of different scales;

[0130] E a =Σ b |W y[n]' (a,b)| 2 ;

[0131] Where: E a It represents the energy on scale a, reflecting the energy distribution of the signal within the frequency range;

[0132] Spectrum center frequency: Extract spectrum center frequency and spectrum bandwidth based on Fourier transform results;

[0133]

[0134] Where: f represents frequency; f c represents the center frequency;

[0135]

[0136] Where: B w Indicates frequency bandwidth;

[0137] Integration of time-frequency domain features: The extracted features are integrated into a time-frequency domain feature vector, where the time-frequency domain feature vector includes: instantaneous frequency, instantaneous amplitude, spectrum center frequency and spectrum bandwidth of short-time Fourier transform, instantaneous frequency and energy spectrum of wavelet transform.

[0138] In this embodiment, the radar signal is decomposed into a low-frequency part and a high-frequency part by wavelet transform. On this basis, the local noise estimation and soft threshold denoising method are used to remove the noise, thereby retaining the true characteristics of the signal. The adaptive filtering algorithm further adjusts the filter weight according to the characteristics of the input signal, and dynamically optimizes the denoising effect of the high-frequency component. Secondly, by adjusting the compensation factor based on the noise estimation value, it is possible to more finely cope with different noise levels, optimize the denoising effect, and improve the signal quality. This dynamic adjustment mechanism can ensure that the denoising process maintains high efficiency and stability under different signal environments.

[0139] By dividing the signal into small segments and performing Fourier transform, the time-frequency characteristics of the signal can be analyzed in real time, and the instantaneous frequency and amplitude information of the signal at different times and frequencies can be extracted; the short-time Fourier transform is particularly suitable for the non-stationary characteristics of the signal, and can effectively reveal the dynamic change characteristics of the target; the continuous wavelet transform can accurately capture the local characteristics of the radar signal at different scales, especially when the signal mutates or changes, it can better reflect its instantaneous frequency and energy distribution, and provide an important basis for subsequent target analysis; therefore, this embodiment not only improves the single target tracking accuracy, but also enhances the robustness and adaptability of the system in complex environments.

[0140] Step 2: Use the obtained time-frequency domain feature data and the multimodal data fusion model based on convolutional neural network to analyze the time-frequency domain feature data to obtain the target detection information, including distance, speed, acceleration, azimuth and type;

[0141] See also Figure 2 As shown, the multimodal data fusion model includes an input layer, a time domain feature separation module, a convolutional neural network module, a self-attention mechanism, a multi-layer perceptron fusion module and an output layer;

[0142] Constructing feature maps For each feature (such as instantaneous frequency, instantaneous amplitude, spectral center frequency, etc.), we can regard them as two-dimensional time-frequency data, and then construct a three-dimensional feature map; store the time-frequency representation of each feature as an independent two-dimensional matrix, and finally stack these matrices together to form a three-dimensional feature map with multiple channels:

[0143] Assuming that we extract 4 features, the feature map can be expressed as:

[0144]

[0145] Where: H represents the number of time frames; W represents the number of frequency points; 4 represents the number of channels (each channel corresponds to a feature: instantaneous frequency, instantaneous amplitude, spectrum center frequency, spectrum bandwidth);

[0146] Similar to the STFT feature, each feature obtained from the wavelet transform is also mapped to two-dimensional time-frequency data, and they are stacked into a multi-channel feature map. For details, see above. The final features are as follows:

[0147]

[0148] Where: H represents the number of time points; W represents the number of wavelet scales; 2 represents the number of channels (each channel corresponds to a feature: instantaneous frequency and energy spectrum).

[0149] By combining the above, we can get the input time-frequency domain features:

[0150] The input layer is used to input the extracted time-frequency domain features;

[0151] The time domain feature separation module establishes an independent input channel for the features extracted by short-time Fourier transform and continuous wavelet transform, and each input channel is sent to a convolutional neural network module for feature extraction;

[0152] In one possible implementation of this embodiment, each feature is clearly distinguished by adding an identifier to the data; for example, an extra dimension is added to the input data to mark the STFT features and CWT features in the current input data; and the input matrix is ​​provided with a channel identifier for distinction.

[0153] STFT feature channel: The features extracted by STFT (such as instantaneous frequency, instantaneous amplitude, etc.) are fed into a dedicated convolutional neural network (CNN) module, which will focus on extracting high-level time-frequency information from the STFT features;

[0154] CWT feature channel: The features extracted by CWT (such as instantaneous frequency, energy spectrum, etc.) are fed into another convolutional neural network (CNN) module; this module will focus on extracting information with high time-frequency resolution from the CWT features.

[0155] The convolutional neural network module includes multiple convolutional layers and multiple pooling layers. The features of each input channel sent to the convolutional neural network module are processed to output the local feature representation of each modality;

[0156] An exemplary STFT feature CNN module structure is as follows:

[0157] First convolutional layer: input The convolution kernel size is 3×3; the number of output channels is 32; the activation function is ReLU;

[0158] Second convolutional layer: input The convolution kernel size is 3×3; the number of output channels is 64; the activation function is ReLU;

[0159] Maximum pooling layer: the pooling size is 2×2;

[0160] The third convolutional layer: input The convolution kernel size is 3×3; the number of output channels is 128; the activation function is ReLU;

[0161] Fourth convolutional layer: input The convolution kernel size is 3×3; the number of output channels is 256; the activation function is ReLU;

[0162] Fully connected layer: Input (Flattened features), output a vector containing the target feature representation, that is, the local feature representation of the target;

[0163] CWT feature CNN module structure:

[0164] First convolutional layer: input The convolution kernel size is 3×3; the number of output channels is 32; the activation function is ReLU;

[0165] Second convolutional layer: input The convolution kernel size is 3×3; the number of output channels is 64; the activation function is ReLU;

[0166] Maximum pooling layer: the pooling size is 2×2;

[0167] The third convolutional layer: input The convolution kernel size is 3×3; the number of output channels is 128; the activation function is ReLU;

[0168] Fourth convolutional layer: input The convolution kernel size is 3×3; the number of output channels is 256; the activation function is ReLU;

[0169] Fully connected layer: Input (Flattened features), output a vector containing the target feature representation, that is, the local feature representation of the target;

[0170] In the exemplary technical solution given above, the structures of the CWT feature CNN module and the STFT feature CNN module are the same, but this is not a limitation of the present invention. In the present invention, the CNN modules of the two may adopt different structures.

[0171] The self-attention mechanism performs weighted fusion based on the output of the convolutional neural network module to obtain a weighted feature representation of each convolutional neural network;

[0172] Exemplary: The output feature vector from each convolutional neural network module is:

[0173] Map input features to queries, keys, and values ​​through learned weight matrices;

[0174]

[0175] Where: W Q , W K , W V represents the learnable weight matrix; Q, K, and V represent the corresponding The vector representation of Query, Key and Value; Q', K', V' represent the corresponding Vector representation of Query, Key and Value;

[0176] Calculate the attention score based on the similarity between Query and Key:

[0177]

[0178] Where: d k Indicates the dimensions of Query and Key; K T Indicates the transposition of Key;

[0179] The Value is weighted and summed by the attention score to obtain the weighted feature representation:

[0180]

[0181] Among them: the Softmax operation normalizes the scores to ensure that the sum of all weights is 1;

[0182] Based on the above logic, we can get and They respectively represent the feature representation of the output results of the CWT feature CNN module and the STFT feature CNN module after the self-attention mechanism is applied;

[0183] The multi-layer perceptron fusion module fuses all feature representations weighted by the self-attention mechanism and outputs final target information, including target detection information, wherein the target detection information includes distance, speed, acceleration, azimuth and type.

[0184] Example: The weighted features from STFT and CWT are weighted summed to obtain the fused feature representation:

[0185]

[0186] Where: λ represents the hyperparameter; F att represents the fused feature representation;

[0187] Exemplarily, the specific structure of the multi-layer perceptron fusion module is as follows:

[0188] First fully connected layer: O1 = ReLU (W1F att +b1); where: represents the weight matrix; d att Represents the input feature F att The dimension of ; d1 represents the output dimension; Represents the bias term; ReLU represents the activation function; represents the output of the first fully connected layer;

[0189] Second fully connected layer: O2 = ReLU (W2O1 + b2); where: represents the weight matrix; d2 represents the output dimension of the second fully connected layer; Represents the bias term; ReLU represents the activation function; represents the output of the second fully connected layer;

[0190] The third fully connected layer: O final =W3O2+b3; wherein: represents the weight matrix; d out Indicates the dimension of the output, corresponding to the number of attributes of target detection; represents the bias term; O final Represents the final output;

[0191] Output layer: The dimension of the output layer is d out , corresponding to the number of target attributes that need to be predicted; the predicted targets include distance, speed, acceleration, azimuth and type, so d out The value of is 5, corresponding to five attributes;

[0192] The distance is predicted by a regression task and a continuous value is output; the speed is predicted by a regression task and a continuous value is output; the acceleration is predicted by a regression task and a continuous value is output; the azimuth is predicted by a regression task and is an angle value in the range of [0,360]; the type is the category of the target, and a category label is output by the classification task, which is normalized by the Softmax activation function and converted into a probability distribution of each category.

[0193] In this embodiment, the self-attention mechanism can automatically learn and emphasize important information in different modal features, suppress redundant and irrelevant information, and thus improve the feature expression ability of the model; by weighted fusion of the features extracted by STFT and CWT, the model can focus on the most discriminative features and improve the precision and accuracy of target detection; through the specially designed CNN modules (STFT feature CNN and CWT feature CNN), the local information of time-frequency domain features can be effectively extracted; the convolution operation can capture subtle signal changes in the local area, and the pooling layer helps to reduce the feature dimension and extract more discriminative high-level features. This process helps to reduce the complexity of the input data while maintaining the effective extraction of key information; based on the weighted features processed by the self-attention mechanism, the multi-layer perceptron module is used to further process the fused features and finally output the detection information of the target; through layer-by-layer nonlinear transformation, MLP can further abstract the features and explore the deep relationship between the features, thereby improving the accuracy of target detection.

[0194] The loss function for the regression task uses the mean square error loss function, and the classification task uses the cross entropy loss function.

[0195] Step 3: Based on the target detection information and the target historical trajectory information, the LSTM model is used to predict the target trajectory to obtain the target's future position and trajectory. The Kalman filter is then used to correct and update the target position in combination with the target's future position and trajectory to obtain the updated target position and state estimate.

[0196] Exemplary: The specific structure of the LSTM model is as follows:

[0197] Input layer: First, you need to define the shape of the input. Since the input is time series data, it contains multiple time steps. The data dimension of each time step is 8+T. Assuming that the time step is N, the input shape is (None, 8+T), where None represents a variable time step length.

[0198] The data dimensions are described as follows:

[0199] The data we input includes distance, azimuth, type, acceleration and speed; the distance is 1 value; the speed is 3 values, which are the speeds in the three directions of x, y and z axis; the acceleration is also 3 values, which are the accelerations in the three directions of x, y and z axis; the azimuth is 1 value; the type T represents the category, which is represented by one-hot encoding; therefore, the data dimension of each time step is 8+T;

[0200] In this embodiment, we add type as input data, because in large target tracking, the target type will still have an impact on the motion behavior, especially in specific scenarios (for example, different types of flying objects or vehicles); understanding the target type helps the LSTM model adjust the modeling of the dynamic characteristics of the target; for example, the acceleration or speed change trends of cars and pedestrians are different; therefore, the model can adjust its prediction strategy according to different types of targets.

[0201] LSTM layer: The LSTM layer includes two LSTM layers; each LSTM layer contains 128 units, and the first layer is used to process the data output by the input layer, return the sequence (return_sequences = True), and use the output of each time step as the input of the next layer;

[0202] The second LSTM layer is used to extract features from the output of the previous LSTM layer and finally output the predicted target future state;

[0203] Add a Dropout layer after the LSTM layer to discard a certain proportion of neurons; the Dropout ratio is 10%-40%;

[0204] Fully connected layer: The output of LSTM is mapped to the predicted values ​​of the target's position, speed, and acceleration through full connection. The dimension of the output layer is 9, and the x, y, and z coordinates of the future position, as well as the values ​​of speed and acceleration on the x, y, and z coordinates are predicted;

[0205] The specific implementation steps of the Kalman filter are as follows:

[0206] State transfer equation:

[0207]

[0208] Where: represents the Kalman filter's prediction of the state at time k; A represents the state transfer matrix; represents the estimated state at the previous moment; B represents the input matrix; u k represents control input;

[0209] For example, our state transition matrix is ​​defined as follows:

[0210]

[0211] Where: Δt represents the time step; the first line indicates the position x k Determined by the influence of position, velocity and acceleration; the second line shows the velocity v k Determined by the effects of velocity and acceleration; the third line shows the acceleration a k Not affected by other factors;

[0212] Control input matrix B: acceleration a k represents the control input, and the control input matrix B is defined as:

[0213]

[0214] The above matrix B represents the influence of acceleration on position, velocity and acceleration state. For example, the control input acceleration affects the state of the target through the time step; where the control input u k =a k ;

[0215] Prediction error covariance matrix: Assuming that the process noise matrix Q reflects the noise introduced by inaccurate acceleration control or other unmodeled factors, the prediction error covariance matrix in the state transfer equation is calculated as:

[0216] P k|k-1 =A·P k-1 ·A T +Q;

[0217] Where: P k|k-1 represents the prediction error covariance matrix, which indicates the uncertainty of the target state prediction; P k-1 represents the estimated error covariance matrix of the previous moment; Q represents the process noise covariance matrix;

[0218] Update phase: Combine the prediction results of the Kalman filter with the actual measurement data and update them through the Kalman gain; assuming that we measure the position and velocity (or other physical quantities) of the target, these measurements constitute the measurement vector z k , Kalman gain K k Used to weigh the difference between the predicted value and the measured value and update the final state estimate:

[0219] The Kalman gain calculation formula is as follows:

[0220] K k =P k|k-1 ·H T ·(H·P k|k-1 ·H T +R) -1 ;

[0221] Where: R represents the measurement noise covariance matrix; H T The device represents the measurement matrix; H represents the measurement matrix, which maps the state space to the measurement space. Assuming that only the position and velocity are measured (for example, z k =[x k ,v k ]), then H is:

[0222]

[0223] Then according to the residual and the Kalman gain K k , update the target state estimate:

[0224]

[0225] Step 3 obtains the updated target state estimate by fusing the prediction result of the LSTM model and the prediction state of the Kalman filter:

[0226]

[0227] Where: represents the predicted state of the Kalman filter; Represents the predicted value of the LSTM model; ω represents a weight factor between 0 and 1; represents the updated target position and state estimate.

[0228] It should be noted here that when fusing the prediction results of the Kalman filter and the LSTM model, we use the prediction state As the basis for fusion, the final output of the Kalman filter is The result is the result after the measurement data is updated. As the basis for fusion, it will be strongly influenced by the measurement data currently updated by the Kalman filter, thereby ignoring the prediction of the LSTM model; therefore, the predicted state is adopted As the basis for fusion, it helps to maintain the independence of the Kalman filter's prediction of future states and the LSTM model's prediction of future trajectories during fusion, avoiding over-reliance on the correction results of measured data;

[0229] Based on this, it can be seen that the update state of the above-mentioned Kalman filter may not be necessary, but it is determined for specific application scenarios; for example, if the application scenario relies on high-frequency sensor data or the environment changes rapidly, such as autonomous driving, then the subsequent update stage is also necessary; if the application scenario has less measured data or changes slowly.

[0230] In this embodiment, the target trajectory is predicted by the LSTM model, and combined with the correction of the target state by the Kalman filter, accurate prediction of the target trajectory and real-time state update can be achieved; LSTM is good at processing time series data and can capture the movement law of the target from the historical trajectory, and the Kalman filter optimizes the prediction of LSTM by considering the noise and uncertainty of the system, thereby effectively improving the accuracy of target state estimation.

[0231] Step 4: Use a deep reinforcement learning algorithm to observe the state of the target and the state of the gimbal based on the updated target position and state estimation and the current gimbal state, perform strategy learning, and optimize the gimbal control strategy;

[0232] The deep reinforcement learning algorithm in step 4 is as follows:

[0233] State definition: define the state vector Where P t represents the updated target position estimate; represents the updated target state estimate; C t Indicates the position of the PTZ; Indicates the speed status of the gimbal;

[0234] Action space definition: Use the gimbal angle and angular velocity as the control action A t =[Δθ,Δv θ ], where Δθ represents the change in the gimbal angle; Δv θ Indicates the change in the angular velocity of the gimbal;

[0235] Reward function:

[0236]

[0237] Where: ||e P || represents the target position error; represents the target velocity error; |Δθ| and |Δv θ | represents the change in the angle and angular velocity of the gimbal; α, β, γ, δ represent weight coefficients;

[0238] in:

[0239] e P =P t -C t ;

[0240]

[0241] Where: k1 and k2 represent control gains, which adjust the weights of position error and speed error respectively; k3 and k4 represent control gains; express rate of change over time;

[0242] Q value update: Update the Q value function Q(S) based on the Q-learning algorithm t ,A t ):

[0243]

[0244] Where: Rt represents the reward at the current moment; η represents the learning rate; γ' represents the discount factor; A' represents the next action;

[0245] The control action of the gimbal is selected based on the deep reinforcement learning algorithm, and the control action is sent to the gimbal control system to adjust the angle and angular velocity of the gimbal.

[0246] The control action of the pan / tilt is selected by using the ε-greedy strategy to select the optimal action from the Q value according to S t And the corresponding Q value function, select the action corresponding to the maximum Q value.

[0247] In this embodiment, by penalizing the target position error e P and speed error Ensure that the gimbal can accurately track the position and speed of the target; if the tracking error of the gimbal is too large, it will be subject to a higher penalty, thereby guiding the algorithm to select a more precise control strategy; secondly, by penalizing the change in the gimbal angle and angular velocity, the gimbal is encouraged to adopt smooth control actions to avoid excessive oscillation or unstable control behavior, which helps to reduce the inertia effect in the gimbal control and ensure the stability and accuracy of the system.

[0248] Step 5: The gimbal control system generates corresponding control instructions based on the optimized control strategy to adjust the gimbal movement; and updates the actual state of the gimbal and feeds it back to the deep reinforcement learning algorithm for further learning and optimization.

[0249] The gimbal control system feeds back the actual state of the gimbal (for example, the current angle and angular velocity) to the deep reinforcement learning algorithm; based on the comparison between the actual execution status and the predicted state, the deep reinforcement learning algorithm can continue to adjust the Q-value function to optimize the control strategy of the gimbal. How to specifically adjust the Q-value function and optimize the control strategy of the gimbal, based on the above-mentioned technical solution, is a conventional technical means in this field and will not be elaborated in this application.

[0250] In the present invention, the radar signal data is denoised and enhanced by wavelet transform and adaptive filtering algorithm, which effectively reduces the interference of noise; short-time Fourier transform and continuous wavelet transform are used for time-frequency domain analysis to extract more accurate target time-frequency feature data; convolutional neural network (CNN) is combined with multimodal data fusion model to accurately extract relevant information of the target from the time-frequency domain feature data, including distance, speed, acceleration, azimuth, etc., thereby improving the accuracy of target detection and reducing the risk of false detection and missed detection; the historical trajectory of the target is learned by LSTM model to predict the future motion trajectory of the target, and the prediction result is corrected and updated by Kalman filter, thereby achieving more accurate target detection and less error detection and missed detection. Accurate target position estimation; using a deep reinforcement learning algorithm, dynamically optimize the control strategy of the gimbal according to the target's update status and the current state of the gimbal, to ensure that the gimbal can respond quickly to the rapid changes of the target, thereby effectively coping with the complex dynamic behavior of the target; the gimbal control system generates control instructions according to the optimized control strategy, accurately adjusts the gimbal position, and feeds back the actual state of the gimbal to the deep reinforcement learning model to form a closed-loop learning, continuously improving the system's adaptability and control accuracy; based on this, the present invention significantly improves the robustness and accuracy of radar-assisted target tracking in complex environments, and effectively solves the problems of inaccurate control caused by false detection, missed detection, and rapid target movement.

[0251] The above are only preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application.

Claims

1. A radar-assisted single target tracking and pan / tilt linkage control method, characterized in that: The following steps are involved: Step 1: Collect the original radar signal data, and denoise and enhance the collected radar signal data based on wavelet transform and adaptive filtering algorithm; Then, based on the denoised and enhanced radar signal data, short-time Fourier transform and continuous wavelet transform are used to perform time-frequency domain analysis to obtain time-frequency domain feature data; Step 2: Use the obtained time-frequency domain feature data and the multimodal data fusion model based on convolutional neural network to analyze the time-frequency domain feature data to obtain the target detection information, including distance, speed, acceleration, azimuth and type; Step 3: Based on the target detection information and the target historical trajectory information, the LSTM model is used to predict the target trajectory to obtain the target's future position and trajectory. The Kalman filter is then used to correct and update the target position in combination with the target's future position and trajectory to obtain the updated target position and state estimate. Step 4: Use a deep reinforcement learning algorithm to observe the state of the target and the state of the gimbal based on the updated target position and state estimation and the current gimbal state, perform strategy learning, and optimize the gimbal control strategy; Step 5: The gimbal control system generates corresponding control instructions based on the optimized control strategy to adjust the gimbal movement; and updates the actual state of the gimbal and feeds it back to the deep reinforcement learning algorithm for further learning and optimization.

2. The radar-assisted single target tracking and pan / tilt linkage control method according to claim 1, characterized in that: The specific steps of denoising and enhancing the radar signal are as follows: Wavelet transform: Discrete wavelet transform is used to decompose the original radar signal data into low-frequency components and high-frequency components; Local noise estimation: A local noise estimation method is used to calculate the noise level of each component that is heavier than the high-frequency component; Soft threshold denoising: Soft threshold processing is performed on the high-frequency components after noise estimation: D k [n] denoised =sgn(D k [n])·max(|D k [n]|-λ k ,0); Where: represents the wavelet threshold of the kth layer; γ represents the constant factor; represents the noise estimate obtained by wavelet decomposition at the kth layer; D k [n] represents the high-frequency component of the k-th wavelet decomposition; D k [n] denoised Represents the high-frequency component after soft threshold denoising; Adaptive filter output: y[n]=w T [n]·D k [n] denoised ; Where: w[n] represents the weight vector of the adaptive filter, which is a parameter that changes with time and is used to adjust the weight vector according to the input signal D k [n] denoised To adjust the output; w T [N]·D k [n] denoised represents the dot product operation, which represents the linear combination of the weight vector and the input signal; Error calculation: e[n]=D k [n] denoised -y[n]; Where: e[n] represents the error of the adaptive filter; y[n] represents the output of the filter; Weight update: w[n+1]=w[n]+µ k ·e[n]·D k [n] denoised 4 Where: w[n+1] represents the updated weight vector; μ k represents the step size factor; w[n] represents the weight vector before updating; Updated high frequency section: D k [n] denoised,updated =D k [n] denoised -w T [n]·D k [n] denoised ; Where: D k [n] denoised,updated Represents the high-frequency component after adaptive filtering; The final signal synthesis: y[n]'=A k [n]+D k [n] denoised,updated ; Where: y[n]' represents the final denoised signal; A k [n] represents the low-frequency component obtained by wavelet transform.

3. The radar-assisted single target tracking and pan-tilt linkage control method according to claim 2, characterized in that: The compensation factor is dynamically adjusted based on the noise estimate obtained by wavelet decomposition of the kth layer: Where: α represents the adjustment factor; Represents the noise estimate obtained by the k-th layer wavelet decomposition.

4. The radar-assisted single target tracking and pan-tilt linkage control method according to claim 1, characterized in that: The specific steps for time-frequency domain analysis based on denoised and enhanced radar signal data are as follows: Short-time Fourier transform: The denoised and enhanced radar data is divided into several small blocks, and Fourier transform is performed in each time window through a Gaussian window to obtain a two-dimensional complex matrix containing the amplitude and phase information of the signal at different time points and frequency points; Instantaneous frequency extraction of short-time Fourier transform: Calculate the instantaneous frequency of the signal based on the obtained two-dimensional complex matrix: Where: arg(X(τ,f)) represents the signal phase; f inst (τ) represents the instantaneous frequency of the short-time Fourier transform; X(τ,f) represents the two-dimensional complex matrix output by the short-time Fourier transform; represents the derivative with respect to the time parameter τ; Instantaneous amplitude extraction: the loss amplitude is the amplitude value of the short-time Fourier transform amplitude spectrum; Wavelet transform: Continuous wavelet transform is used to perform time-frequency analysis on the denoised and enhanced radar data to obtain wavelet transform coefficients; Frequency feature extraction: Based on the wavelet transform coefficients obtained by continuous wavelet transform, the frequency response at different scales is calculated to obtain the frequency features; Instantaneous frequency of wavelet transform: Based on the time-frequency spectrum of wavelet transform, the instantaneous frequency is obtained by analyzing the wavelet transform phase transformation of the signal: Where: represents the instantaneous frequency of wavelet transform; arg(W y[n]' (a, b)) represents the phase part of wavelet transform; represents the derivative of the phase with respect to the translation parameter b; Energy spectrum extraction: The energy spectrum of the signal is obtained by squaring and summing the wavelet transform coefficients of different scales; Spectrum center frequency: Extract spectrum center frequency and spectrum bandwidth based on Fourier transform results; Integration of time-frequency domain features: The extracted features are integrated into a time-frequency domain feature vector, where the time-frequency domain feature vector includes: instantaneous frequency and instantaneous amplitude of short-time Fourier transform, instantaneous frequency, energy spectrum, spectrum center frequency and spectrum bandwidth of wavelet transform.

5. The radar-assisted single target tracking and pan-tilt linkage control method according to claim 1, characterized in that: The multimodal data fusion model includes an input layer, a time domain feature separation module, a convolutional neural network module, a self-attention mechanism, a multi-layer perceptron fusion module and an output layer; The input layer is used to input the extracted time-frequency domain features; The time domain feature separation module establishes an independent input channel for the features extracted by short-time Fourier transform and continuous wavelet transform, and each input channel is sent to a convolutional neural network module for feature extraction; The convolutional neural network module includes multiple convolutional layers and multiple pooling layers. The features of each input channel sent to the convolutional neural network module are processed to output the local feature representation of each modality; The self-attention mechanism performs weighted fusion based on the output of the convolutional neural network module to obtain a weighted feature representation of each convolutional neural network; The multi-layer perceptron fusion module fuses all feature representations weighted by the self-attention mechanism and outputs final target information, including target detection information, wherein the target detection information includes distance, speed, acceleration, azimuth and type.

6. The radar-assisted single target tracking and pan-tilt linkage control method according to claim 1, characterized in that: Step 3 obtains the updated target state estimate by fusing the prediction result of the LSTM model and the prediction state of the Kalman filter: Where: represents the predicted state of the Kalman filter; represents the predicted value of the LSTM model; H represents the measurement matrix; ω represents a weight factor between 0 and 1; represents the updated target position and state estimate.

7. The radar-assisted single target tracking and pan / tilt linkage control method according to claim 1, characterized in that: The deep reinforcement learning algorithm in step 4 is as follows: State definition: Define the state vector Where P t represents the updated target position estimate; represents the updated target state estimate; C t Indicates the position of the PTZ; Indicates the speed status of the gimbal; Action space definition: Use the gimbal angle and angular velocity as the control action A t =[Δθ,Δv θ ], where Δθ represents the change in the gimbal angle; Δv θ Indicates the change in the angular velocity of the gimbal; Reward function: Where: ||e P || represents the target position error; represents the target velocity error; |Δθ| and |Δv θ | represents the change in the angle and angular velocity of the gimbal; α, β, γ, δ represent weight coefficients; in: e P =P t -C t ; Where: k1 and k2 represent control gains, which adjust the weights of position error and speed error respectively; k3 and k4 represent control gains; express rate of change over time; Q value update: Update the Q value function Q(S) based on the Q-learning algorithm t ,A t ): Where: R t represents the reward at the current moment; η represents the learning rate; γ' represents the discount factor; A' represents the next action; The control action of the gimbal is selected based on the deep reinforcement learning algorithm, and the control action is sent to the gimbal control system to adjust the angle and angular velocity of the gimbal.

8. The radar-assisted single target tracking and pan / tilt linkage control method according to claim 7, characterized in that: The control action of the pan / tilt is selected by using the ε-greedy strategy to select the optimal action from the Q value according to S t And the corresponding Q value function, select the action corresponding to the maximum Q value.

Citation Information

Cited By

  • Intelligent holder adaptive tracking control method based on reinforcement learning

    CN121857317A