Spectrum analyzer parameter automatic adjustment method and system based on deep learning
By constructing a deep Q-learning network model, the parameters of the spectrum analyzer are automatically adjusted, solving the problem that the adjustment of spectrum analyzer parameters depends on human experience, and achieving more efficient and accurate spectrum measurement, especially in complex signal environments.
Patent Information
- Application Number
- CN202511196439.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-26
- Publication Date
- 2025-12-05
- Estimated Expiration
- 2045-08-26
AI Technical Summary
Spectrum analyzer parameter adjustments rely on human experience, leading to inconsistencies, poor adaptability, low efficiency, and high skill requirements for operators, especially in complex or small-signal environments where measurement results are unsatisfactory.
A deep Q-learning network model is constructed to automatically adjust key parameters of the spectrum analyzer, including center frequency, bandwidth, attenuation, and reference level, through a Markov decision process, thereby optimizing measurement performance using deep learning.
Significantly improves the efficiency and accuracy of spectrum measurement, reduces human intervention, adapts to various complex signal environments, quickly reaches the optimal measurement settings, and maintains stable performance.
Smart Images

Figure CN120750464B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of radio spectrum measurement technology, and more specifically to a method and system for automatic parameter adjustment of spectrum analyzers based on deep learning. Background Technology
[0002] Spectrum analyzers are indispensable instruments in fields such as wireless communication, electromagnetic compatibility testing, and signal processing. However, selecting the optimal combination of measurement parameters for a specific signal scenario is often a challenging task, requiring operators with specialized knowledge and extensive experience. Inappropriate parameter settings can lead to inaccurate measurement results or significantly increase measurement time.
[0003] Traditionally, testers manually adjust various parameters of the spectrum analyzer based on experience, such as center frequency, bandwidth, resolution bandwidth (RBW), video bandwidth (VBW), attenuation, and reference level. This process is not only time-consuming but also easily affected by human factors, especially in complex or small-signal environments. Therefore, traditional spectrum analyzer parameter adjustment has the following limitations:
[0004] 1. Inconsistencies caused by human factors;
[0005] 2. Poor adaptability to complex signal environments;
[0006] 3. Low measurement efficiency;
[0007] 4. High skill requirements for operators. Summary of the Invention
[0008] The purpose of this invention is to provide a method and system for automatic parameter adjustment of spectrum analyzers based on deep learning. By modeling the parameter adjustment process of the spectrum analyzer as a deep Q-learning network model of a Markov decision process (MDP), the key parameters of the spectrum analyzer are automatically adjusted to optimize measurement performance. This method not only reduces human intervention but also adapts to various complex signal environments, providing a new solution for the automation and intelligence of spectrum measurement.
[0009] To achieve the above objectives, this application provides the following solution:
[0010] On the one hand, this application provides a method for automatic parameter adjustment of a spectrum analyzer based on deep learning, specifically including the following steps:
[0011] S1. Receive the spectrum analyzer's various parameters and their ranges, as well as the signal parameters sent by the signal generator;
[0012] S2. Construct a deep Q-learning network model. Based on the various parameters and ranges of the spectrum analyzer and the signal parameters, set the training parameters and training conditions of the deep Q-learning network.
[0013] S3. Train a deep Q-learning network model based on training parameters and training conditions. Adaptively adjust the various parameters of the spectrum analyzer within the range of each parameter. Select the optimal parameter configuration action as the optimal combination of spectrum analyzer parameters by predicting the Q value.
[0014] In some specific implementations, the training parameters include the state space and action space of the deep Q-learning network. The state space includes the spectrum analyzer parameter range and signal parameters. The spectrum analyzer parameters include center frequency, sweep width, RBW, VBW, attenuation value, and reference level. The signal parameters include signal frequency and signal power. The action space is the parameter combination composed of each spectrum analyzer parameter within its respective range in the state space.
[0015] In some specific implementations, the training parameters include the total number of training rounds, the size of the experience pool for each training round, and the number of time steps. For each time step in each training round, the training process of the deep Q-learning network model in step S3 includes:
[0016] S31. Select an action space from the experience pool and decode the action space into a combination of spectrum analyzer parameters;
[0017] S32. Obtain trace data obtained by measuring signal parameters under the current spectrum analyzer parameter combination;
[0018] S33. Perform spectral analysis on the trace data to obtain spectral data;
[0019] S34. Analyze the spectrum data to obtain the signal-to-noise ratio, amplitude accuracy, frequency accuracy, signal power, scan time, and noise floor of the signal under test under the current spectrum analyzer parameters;
[0020] S35. Input the signal-to-noise ratio, amplitude accuracy, frequency accuracy, signal power, scan time, and noise floor into the reward function to calculate the composite reward value R;
[0021] S36. Store the spectrum analyzer parameter combinations and their corresponding composite reward values R in the experience pool;
[0022] S37. Repeat steps S31-S36 until training is complete.
[0023] In some specific implementations, the reward function is as follows:
[0024] R = wSNR × SNR + wA × A + wF × F + wSWT× SWT + wP × P + wN × N
[0025] Wherein, SNR represents signal-to-noise ratio, A represents amplitude accuracy, F represents frequency accuracy, P represents signal power, SWT represents scan time, and N represents noise floor; wSNR , wA , wF , wSWT , wP , wN This represents the weighting coefficient.
[0026] In some specific implementations, the specific process of calculating the signal-to-noise ratio in step S34 is as follows:
[0027] S341. According to the bandwidth and resolution bandwidth, the spectrum data is sampled to obtain several frequency points and the amplitude value corresponding to each frequency point;
[0028] S342. Use the sliding window technique to calculate the average and standard deviation of the amplitude values corresponding to all frequency points, and find the peak value of the amplitude value of the spectrum data based on the average and standard deviation.
[0029] S343. Divide several frequency points into signal points and noise points according to the peak value of the amplitude, and calculate the signal power and noise power;
[0030] S344. The signal-to-noise ratio is calculated based on the signal power and noise power.
[0031] In some specific embodiments, the specific process of step S342 is as follows:
[0032] The number and size of the sliding windows are set according to the bandwidth. The average value of the amplitude at each frequency point within each sliding window is calculated to obtain the average value of each sliding window.
[0033] Calculate the mean and standard deviation of all sliding windows within the bandwidth;
[0034] The dynamic threshold is calculated based on the mean and standard deviation;
[0035] The peak amplitude value of the spectrum data is obtained by using dynamic threshold detection.
[0036] In some specific implementations, the method for calculating the dynamic threshold based on the mean and standard deviation is as follows:
[0037] Dynamic threshold = mean + standard deviation * threshold factor.
[0038] In some specific implementations, when training a deep Q-learning network model, an adaptive parameter tuning strategy based on ε-greedy policy actions is adopted, including dynamic decay control, RBW / VBW linkage, and sweep width optimization.
[0039] In some specific implementations, the dynamic attenuation control process is as follows:
[0040] When the detected signal power is greater than the first preset threshold, the attenuation value is automatically increased according to the preset step value. If the detected power is less than the second preset threshold, the attenuation is turned off and the preamplifier is enabled.
[0041] In some specific implementations, the RBW / VBW linkage process is as follows:
[0042] The RBW / VBW ratio is dynamically set according to the signal-to-noise ratio. VBW = RBW × Δx, and the value of Δx ranges from 1 / 10 to 1 / 100.
[0043] Secondly, this application provides a deep learning-based automatic parameter adjustment system for a spectrum analyzer, comprising:
[0044] The signal receiving module is used to receive various parameters and their ranges from the spectrum analyzer and the signal parameters from the signal generator.
[0045] The system initialization module is used to build a deep Q-learning network model. Based on the various parameters and ranges of the spectrum analyzer and the signal parameters, it sets the training parameters and training conditions of the deep Q-learning network.
[0046] The adaptive adjustment module is used to train a deep Q-learning network model based on training parameters and training conditions. It adaptively adjusts various parameters of the spectrum analyzer within the range of various parameters of the spectrum analyzer, and selects the optimal parameter configuration action as the optimal combination of spectrum analyzer parameters through Q value prediction.
[0047] The beneficial effects of this invention are as follows:
[0048] This application constructs a deep Q-learning network model, using spectrum analyzer parameters as the state space of the deep Q-learning network. Through real-time interaction with the spectrum analyzer, the deep Q-learning network model learns how to maximize amplitude and frequency accuracy. By analyzing the spectral data output by each set of spectrum analyzer parameters, a reward value for each set of parameters is calculated. After a certain number of iterations, the set of spectrum analyzer parameters with the highest reward value is selected as the spectrum analyzer parameter settings for analyzing the signal under test. This significantly improves the efficiency and accuracy of spectrum measurement, especially in complex and small-signal environments. Compared with traditional manual parameter adjustment, this method not only reduces human intervention but also adapts to various complex signal environments, achieving optimal measurement settings faster and maintaining stable performance under various signal scenarios. Attached Figure Description
[0049] Figure 1 A flowchart of an automatic parameter adjustment method for a spectrum analyzer provided in an embodiment of this application;
[0050] Figure 2 A flowchart illustrating the training and application process of a deep Q-learning network model provided in this application embodiment. Detailed Implementation
[0051] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit the present invention or its application or use. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0052] Unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps described in these embodiments do not limit the scope of the invention.
[0053] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.
[0054] Furthermore, for clarity and brevity, descriptions of well-known structures, functions, and configurations may have been omitted. Those skilled in the art will recognize that various changes and modifications can be made to the examples described herein without departing from the spirit and scope of this disclosure.
[0055] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0056] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not as limitations. Therefore, other examples of exemplary embodiments may have different values.
[0057] Example 1
[0058] like Figure 1 As shown, this embodiment provides a method for automatic parameter adjustment of a spectrum analyzer based on deep learning, specifically including the following steps:
[0059] S1. Receive the spectrum analyzer's various parameters and their ranges, as well as the signal parameters sent by the signal generator;
[0060] S2. Construct a deep Q-learning network model. Based on the various parameters and ranges of the spectrum analyzer and the signal parameters, set the training parameters and training conditions of the deep Q-learning network.
[0061] The training parameters include the state space and action space of the deep Q-learning network. The state space includes the spectrum analyzer parameter range and signal parameters. The spectrum analyzer parameters include center frequency, sweep width, RBW, VBW, attenuation value, and reference level. The signal parameters include signal frequency and signal power. The action space is the combination of parameters formed by the various spectrum analyzer parameters within their respective ranges in the state space.
[0062] S3. Train a deep Q-learning network model based on training parameters and training conditions. Adaptively adjust the various parameters of the spectrum analyzer within the range of each parameter. Select the optimal parameter configuration action as the optimal combination of spectrum analyzer parameters by predicting the Q value.
[0063] Training parameters include the total number of training rounds, the size of the experience pool for each training round, and the number of time steps. For each time step in each training round, the training process of the deep Q-learning network model in step S3 includes:
[0064] S31. Select an action space from the experience pool and decode the action space into a combination of spectrum analyzer parameters;
[0065] S32. Obtain trace data obtained by measuring signal parameters under the current spectrum analyzer parameter combination;
[0066] S33. Perform spectral analysis on the trace data to obtain spectral data;
[0067] S34. Analyze the spectrum data to obtain the signal-to-noise ratio, amplitude accuracy, frequency accuracy, signal power, sweep time, and noise floor of the signal under test with the current spectrum analyzer parameters;
[0068] The specific process of calculating the signal-to-noise ratio in step S34 is as follows:
[0069] S341. Sample the spectrum data according to the bandwidth and resolution bandwidth to obtain a number of frequency points and the amplitude values corresponding to each frequency point;
[0070] S342. Use the sliding window technique to calculate the average value and standard deviation of the amplitude values corresponding to all frequency points, and find the peak value of the amplitude value of the spectrum data based on the average value and standard deviation;
[0071] The specific process of step S342 is as follows:
[0072] Set the number and size of the sliding windows according to the bandwidth, and calculate the average value of the amplitude values of each frequency point within each sliding window to obtain the average value of each sliding window;
[0073] Calculate the average value and standard deviation of all sliding windows within the bandwidth;
[0074] Calculate the dynamic threshold according to the average value and standard deviation;
[0075] The method of calculating the dynamic threshold according to the average value and standard deviation is as follows:
[0076] Dynamic threshold = average value + standard deviation * threshold factor.
[0077] Use the dynamic threshold to detect the peak value of the amplitude value of the spectrum data.
[0078] For example, sliding window dynamic threshold:
[0079] 1. Calculate the sliding mean μ(f) and standard deviation σ(f) (window size = 10 points) for the spectrum data x(f);
[0080] 2. Set the dynamic threshold to μ(f) + 2σ(f).
[0081] 3. Detect the peaks exceeding the threshold as signals, and the rest as noise.
[0082] Noise floor estimation:
[0083] The average power of the filtered noise region is Pnoise, and it is compared with the instrument's nominal noise floor (DANL). If Pnoise < DANL, then force Pnoise = DANL;
[0084] S343. Divide a number of frequency points into signal points and noise points according to the peak value of the amplitude value, and calculate the signal power and noise power;
[0085] S344. The signal-to-noise ratio is calculated based on the signal power and noise power.
[0086] S35. Input the signal-to-noise ratio, amplitude accuracy, frequency accuracy, signal power, scan time, and noise floor into the reward function to calculate the composite reward value R;
[0087] The reward function is:
[0088] R = wSNR × SNR + wA × A + wF × F + wSWT × SWT + wP × P + wN × N
[0089] Wherein, SNR represents signal-to-noise ratio, A represents amplitude accuracy, F represents frequency accuracy, P represents signal power, SWT represents scan time, and N represents noise floor; wSNR , wA , wF , wSWT , wP , wN Represents the weighting coefficient, for example wSNR =0.3; wA =0.2; wF =0.2; wSWT =-0.1; wP =0.1; wN =0.1
[0090] S36. Store the spectrum analyzer parameter combinations and their corresponding composite reward values R in the experience pool;
[0091] S37. Repeat steps S31-S36 until training is complete.
[0092] When training the deep Q-learning network model, an adaptive parameter tuning strategy based on ε-greedy policy actions is adopted, including dynamic decay control, RBW / VBW linkage, and sweep width optimization.
[0093] 1. The attenuation control process is as follows:
[0094] When the detected signal power is greater than the first preset threshold (e.g., greater than -10dBm), the attenuation value is automatically increased according to the preset step value (10dB step) to prevent overload; if the detected power is less than the second preset threshold (e.g., less than -50dBm), the attenuation is turned off and the preamplifier is enabled.
[0095] 2. The process of RBW / VBW linkage is as follows:
[0096] The RBW / VBW ratio is dynamically set based on the signal-to-noise ratio. VBW = RBW × Δx, where Δx ranges from 1 / 10 to 1 / 100, and is used to balance resolution and scanning time.
[0097] 3. Width Optimization
[0098] An initial wide scan (Span=3GHz) quickly locates the signal, followed by a narrow scan (Span=100kHz) to improve accuracy.
[0099] Understandably, this embodiment proposes a deep reinforcement learning-based method for automatically optimizing measurement parameter settings of a spectrum analyzer. A Deep Q-Learning Network (DQN) model is proposed, which can automatically adjust key parameters such as center frequency (CF), bandwidth (Span), resolution bandwidth (RBW), video bandwidth (VBW), attenuation (Att), and reference level (RefLevel) based on given signal characteristics. Through real-time interaction with the spectrum analyzer and signal source, the DQN agent learns how to maximize the signal-to-noise ratio (SNR), minimize amplitude and frequency errors, while maintaining a reasonable scan time. This method significantly improves the efficiency and accuracy of spectrum measurements, especially in complex and small-signal environments. Experimental results show that compared to traditional manual parameter adjustment, this method can reach optimal measurement settings faster and maintain stable performance under various signal scenarios.
[0100] In this application, the spectrum analyzer and signal generator are controlled by SCPI commands to acquire spectrum data in real time. The test parameters (RBW, VBW, attenuation value, etc.) and signal characteristics (SNR, frequency error, etc.) are encoded into a multi-dimensional state vector. A deep Q-network (DQN) agent is used to select the optimal parameter configuration action through Q-value prediction. A composite reward mechanism is equipped, which integrates signal-to-noise ratio (SNR), scan time (SWT), and frequency / power accuracy to design a multi-objective reward function. An adaptive parameter tuning strategy is adopted to dynamically adjust the exploration rate (ε-greedy) and learning rate according to environmental feedback.
[0101] The state space of a deep Q-learning network model is defined as follows: State vector S = [Signal frequency, Signal power, Center frequency, Span, Resolution bandwidth (RBW), Video bandwidth (VBW), Attenuation value (ATT), Reference level (Ref Level)].
[0102] Motion space design: Discrete motions: RBW (15 levels), VBW ratio (5 levels), attenuation (4 levels), reference level (8 levels), sweep width scaling factor (±10%). Total number of motions: 15×5×4×8×3=7200 combinations. The specific parameters and their ranges are shown in the table below:
[0103] Table 1 Discretization parameters and their possible values
[0104]
[0105] like Figure 2 As shown, the automatic parameter adjustment method in this embodiment is divided into a model training phase and a model application phase. The model training process is as follows:
[0106] 1. Initialize the environment and configure the interactive environment.
[0107] Hardware device integration:
[0108] Spectrum Analyzer: A wideband device supporting SCPI command control (such as the K©ysight N9030B), with a frequency coverage of 10Hz to 50GHz, a maximum input power of +30dBm, and a noise floor of -170dBm. Signal Generator: A programmable RF source (such as the R&S SMBV100B), with an output range of -145dBm to +25dBm and a frequency resolution of 0.01Hz.
[0109] Control interface: Implements LAN or GPIB communication based on the PyVISA library, supporting cross-platform operation (Windows / Linux).
[0110] Data acquisition protocol:
[0111] SCPI instruction set:
[0112] / *Example of spectrum analyzer parameter settings* /
[0113] FREQ:CENT 1GHz / / Set the center frequency
[0114] FREQ:SPAN 100MHz / Set sweep width
[0115] BAND:RES 10kHz / / Set RBW
[0116] BAND:VID 1kHz / / Set VBW
[0117] POW:ATT 20dB / 1 sets the input attenuation.
[0118] :DISP:WIND:TRAC:Y:RLEV-30dBm / / Set reference level
[0119] Asynchronous communication mechanism: Multi-threading technology is used to achieve parallel processing of command sending and data acquisition, avoiding blocking and waiting.
[0120] 2. Configure instrument parameters
[0121] State vector definition: The RF test environment is abstracted as an 8-dimensional state vector S = [fsignal, Psigual, fcenter, Span, RBW, VBW, ATT, RefLevel], where:
[0122] figmal represents the frequency (unit: Hz) of the signal to be measured set by the signal generator;
[0123] Psigual indicates the output power set by the signal generator (unit: dBm);
[0124] Fcenter represents the current center frequency of the spectrum analyzer (unit: Hz);
[0125] Span indicates the spectrum analyzer sweep width (unit: Hz);
[0126] RBW represents resolution bandwidth (unit: Hz);
[0127] VBW represents video bandwidth (unit: Hz);
[0128] ATT represents the input attenuation value (unit: dB);
[0129] Refevel indicates the reference level (unit: dBm);
[0130] Normalize the above states:
[0131] The state vector is preprocessed using Min-Max standardization, as shown in the formula:
[0132]
[0133] Where Smin and Smax are the allowable ranges for each parameter (e.g., RBW ∈ [1Hz, 1MHz]);
[0134] 3. Decision-making phase (DQN structure initialization)
[0135] Neural network architecture design:
[0136] Input layer: 8 nodes, corresponding to the dimensions of the state vector (each node corresponds to a parameter of a state space).
[0137] Hidden layer: 3 fully connected layers, 128 nodes per layer, with ReLU activation function;
[0138] Output layer: Action space Q-value, with dimension |A|=7200.
[0139] Experience replay mechanism:
[0140] Replay buffer: Capacity 2000 experience tuples (St, At, Rt, St+1, Done);
[0141] Priority sampling: Sampling weights are assigned based on the TD error δ = (Qtarget (target Q value) - Qpred (predicted Q value)) to accelerate convergence;
[0142] 4. Optimize the execution phase
[0143] Action mapping rules:
[0144] Decoding discrete actions into specific parameter combinations (decomposing a single action index into multiple parameters), the decoding code is as follows:
[0145] rbw_idx = action % len(RBW_VALUES) # Take the modulo operation to get the RBW index
[0146] vbw_idx = (action / / len(RBW_VALUES)) % len(VBW_RATIOS)
[0147] att_idx = (action / / (len(RBW_VALUES)*len(VBW_RATIOS))) % len(ATT_VALUES)
[0148] reflevel_idx = (action / / (len(RBW_VALUES)*len(VBW_RATIOS)*len(ATT_VALUES))) %len(REF_LEVEL_VALUES)
[0149] span_change =
[0150] (action / / (len(RBW_VALUES)*len(VBW_RATIOS)*len(ATT_VALUES)*len(REF_LEVEL_VALUES))) -1
[0151] Parameter constraint handling:
[0152] Dynamic range protection: If RefLevel > ATT-10, then automatically limit RefLevel to ATT-10.
[0153] Frequency boundary detection: When fcenter ± Span / 2 exceeds the instrument range, reset the sweep width to the maximum allowable value.
[0154] The DQN model is trained by replaying experience and updating the target network.
[0155] Specifically, the deep Q-learning network model mentioned above is trained by initializing the constants and state space parameters of the neural network, starting the training rounds, and the training process for each round includes:
[0156] 1. Data collection: Control the spectrum analyzer and signal generator through SCPI commands to build the test environment, set the signal generator parameters (including center frequency and power (reference level)), turn on the output, and make the signal generator output a random test signal;
[0157] 2. Select an action to be executed from the experience pool (the action includes spectrum analyzer parameters and signal parameters), encode the spectrum analyzer parameters and signal parameters into a state vector, and use a deep Q-network (DQN) agent to generate parameter configuration actions;
[0158] 3. Configure the spectrum analyzer according to its parameters and collect data from it. Collect trace data output by the spectrum analyzer for spectrum analysis. The signal processing flow is as follows: use the sliding window analysis method to standardize the trace data and process it with the sliding signal noise separation algorithm to obtain the signal-to-noise ratio, amplitude accuracy, and frequency accuracy.
[0159] Sliding window signal-noise separation algorithm: such as center frequency, bandwidth, resolution bandwidth (RBW), video bandwidth (VBW), attenuation and reference level, etc. When calculating the signal-to-noise ratio, the sliding window is used to calculate local statistics, and then the signal peak is detected based on dynamic threshold. The peak value and noise are separated from the amplitude data, and finally the SNR / accuracy is calculated.
[0160] 4. Calculate the composite reward based on parameters such as signal-to-noise ratio, scanning time, and accuracy. Store the experience (i.e., spectrum analyzer parameters and corresponding reward values) in the experience pool, update the state, and determine whether the current round has ended. If not, iterate through the data of all time steps starting from each time step. After a round ends, perform experience replay and determine whether the number of rounds is a multiple of 10. If yes, update the target model. If no, determine whether all rounds have been completed. If no, return to each round. After all rounds have been completed, save the trained model and use the trained model for testing.
[0161] When applying the pre-trained model (i.e., not in training mode), load the pre-trained model, input the test signal parameters, use the pre-trained model to predict the optimal parameter configuration, execute the test, and display the results.
[0162] Example 2
[0163] This embodiment 2 applies the automatic adjustment method in embodiment 1 to provide a deep learning-based automatic parameter adjustment system for a spectrum analyzer, including:
[0164] The signal receiving module is used to receive various parameters and their ranges from the spectrum analyzer and the signal parameters from the signal generator.
[0165] The system initialization module is used to build a deep Q-learning network model. Based on the various parameters and ranges of the spectrum analyzer and the signal parameters, it sets the training parameters and training conditions of the deep Q-learning network.
[0166] The adaptive adjustment module is used to train a deep Q-learning network model based on training parameters and training conditions. It adaptively adjusts various parameters of the spectrum analyzer within the range of various parameters of the spectrum analyzer, and selects the optimal parameter configuration action as the optimal combination of spectrum analyzer parameters through Q value prediction.
[0167] It also includes a signal generator module for generating the radio frequency signal to be tested, which is then input into the adaptive adjustment module. The adaptive adjustment module includes a spectrum analysis module for acquiring and processing spectrum data of the input radio frequency signal to be tested, and a deep reinforcement learning agent module for processing the spectrum data according to the trained deep Q-learning network model and outputting decision configuration actions.
[0168] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Based on the technical essence of the present invention, any simple modifications, equivalent substitutions, and improvements made to the above embodiments within the spirit and principles of the present invention shall still fall within the protection scope of the present invention.
Claims
1. A method for automatic adjustment of parameters of a spectrum analyzer based on deep learning, characterized in that, Specifically comprising the following steps: S1, receiving the spectrum analyzer parameters and their ranges sent by the spectrum analyzer and the signal parameters sent by the signal generator; S2, constructing a deep Q learning network model, setting the training parameters and training conditions of the deep Q learning network according to the spectrum analyzer parameters and their ranges and the signal parameters; S3, training the deep Q learning network model based on the training parameters and training conditions, adaptively adjusting the spectrum analyzer parameters within the spectrum analyzer parameter range, and selecting the optimal parameter configuration action as the optimal spectrum analyzer parameter combination through Q value prediction; The training parameters include the state space and action space of the deep Q learning network, wherein the state space includes the spectrum analyzer parameter range and the signal parameter, the spectrum analyzer parameters include the center frequency, the sweep width, the RBW, the VBW, the attenuation value and the reference level; the signal parameter includes the signal frequency and the signal power, and the action space is the parameter combination composed of the spectrum analyzer parameters within their respective ranges in the state space; The training parameters include the total training rounds, the experience pool size of each training, and the number of time steps, and for each time step in each training round, the training process of the deep Q learning network model in step S3 includes: S31, selecting the action space from the experience pool and decoding the action space into the spectrum analyzer parameter combination; S32, obtaining the trace data obtained by measuring the signal parameter under the current spectrum analyzer parameter combination; S33, performing spectrum analysis on the trace data to obtain the spectrum data; S34, analyzing the spectrum data to obtain the signal-to-noise ratio, amplitude accuracy, frequency accuracy, signal power, scanning time and noise floor of the signal to be measured under the current spectrum analyzer parameters; S35, inputting the signal-to-noise ratio, amplitude accuracy, frequency accuracy, signal power, scanning time and noise floor into the reward function formula to calculate the composite reward value R; S36, storing the spectrum analyzer parameter combination and its corresponding composite reward value R in the experience pool; S37, repeating steps S31-S36 until the training is completed.
2. The deep learning based spectrum analyzer parameter automatic adjustment method of claim 1, wherein, The reward function formula is: R = wSNR × SNR + wA × A + wF × F + wSWT × SWT + wP × P + wN × N wherein SNR represents a signal-to-noise ratio, A represents an amplitude accuracy, F represents a frequency accuracy, P represents a signal power, SWT represents a scan time, and N represents a noise floor; wSNR 、 wA 、 wF 、 wSWT 、 wP 、 wN represents a weight coefficient.
3. The deep learning based spectrum analyzer parameter auto-adjustment method of claim 1, wherein, The specific process of calculating the signal-to-noise ratio in step S34 is: S341, sampling the spectrum data according to the bandwidth and resolution bandwidth to obtain a plurality of frequency points and the amplitude values corresponding to each frequency point; S342, calculating the average value and standard deviation of the amplitude values corresponding to all frequency points using the sliding window technology, and finding the amplitude value peak of the spectrum data based on the average value and the standard deviation; S343, dividing a plurality of frequency points into signal points and noise points according to the amplitude value peak, and calculating the signal power and noise power; S344, calculating the signal-to-noise ratio according to the signal power and the noise power.
4. The deep learning based spectrum analyzer parameter auto-adjustment method of claim 1, wherein, The specific process of step S342 is: Set the number and size of the sliding window according to the bandwidth, calculate the average value of the amplitude values of each frequency point in each sliding window to obtain the average value of each sliding window; Calculate the average value and standard deviation of all sliding windows within the bandwidth; Calculate the dynamic threshold value according to the average value and the standard deviation; The amplitude value peak of the spectrum data is detected using a dynamic threshold.
5. The deep learning based spectrum analyzer parameter automatic adjustment method of claim 1, wherein, During training of the deep Q-learning network model, an adaptive parameter adjustment strategy based on an ε-greedy policy action is adopted, including dynamic attenuation control, RBW / VBW linkage, and sweep width optimization.
6. The deep learning based spectrum analyzer parameter auto-adjustment method of claim 5, wherein, The process of dynamic attenuation control is as follows: When the detected signal power is greater than a first preset threshold, the attenuation value is automatically increased by a preset step value, and if the detected power is less than a second preset threshold, the attenuation is turned off and the preamplifier is enabled.
7. The deep learning based spectrum analyzer parameter auto-adjustment method of claim 6, wherein, The process of RBW / VBW linkage is as follows: The RBW / VBW ratio is dynamically set according to the signal-to-noise ratio, and VBW=RBW×△x, with the value of△x ranging from 1 / 10 to 1 / 100.
8. A deep learning based spectrum analyzer parameter automatic adjustment system, characterized in that, The method comprises the following steps: a signal receiving module for receiving the parameters of the spectrum analyzer and their ranges sent by the spectrum analyzer and the signal parameters sent by the signal generator; a system initialization module for constructing a deep Q-learning network model, setting the training parameters and training conditions of the deep Q-learning network according to the parameters of the spectrum analyzer and their ranges and the signal parameters; the training parameters include the state space and action space of the deep Q-learning network, wherein the state space includes the parameter ranges of the spectrum analyzer and the signal parameters, the parameters of the spectrum analyzer include the center frequency, sweep width, RBW, VBW, attenuation value, and reference level; the signal parameters include the signal frequency and signal power, and the action space is a parameter combination composed of the parameters of the spectrum analyzer within their respective ranges; an adaptive adjustment module for training the deep Q-learning network model based on the training parameters and training conditions, adaptively adjusting the parameters of the spectrum analyzer within their respective ranges, and selecting the optimal parameter configuration action as the optimal parameter combination of the spectrum analyzer by Q value prediction; the training parameters include the total training rounds, the experience pool size for each training, and the number of time steps, and for each time step in each training round, the training process of the deep Q-learning network model includes: selecting the action space from the experience pool and decoding the action space into the parameter combination of the spectrum analyzer; obtaining the trace data obtained by measuring the signal parameters under the current parameter combination of the spectrum analyzer; performing spectrum analysis on the trace data to obtain the spectrum data; analyzing the spectrum data to obtain the signal-to-noise ratio, amplitude accuracy, frequency accuracy, signal power, scan time, and noise floor of the signal to be measured under the current parameter combination of the spectrum analyzer; inputting the signal-to-noise ratio, amplitude accuracy, frequency accuracy, signal power, scan time, and noise floor into the reward function formula to calculate the composite reward value R; storing the parameter combination of the spectrum analyzer and the corresponding composite reward value R in the experience pool; repeating the above steps until the training is completed.
Citation Information
Patent Citations
Cognitive radio network dynamic spectrum access method based on deep reinforcement learning
CN115190489A
Signal noise suppression and signal integrity guaranteeing method based on deep learning
CN119169986A