Digital sensitivity test method based on IQ signal sampling, storage and playback

By using IQ signal sampling, storage, and playback methods, combined with inverse reinforcement learning to construct test strategies, the problems of low testing efficiency and insufficient consistency of digital radio receivers are solved, achieving efficient and stable testing in complex environments.

CN121923743APending Publication Date: 2026-04-2463963 TROOP OF THE PLA
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
63963 TROOP OF THE PLA
Filing Date
2026-03-09
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing digital sensitivity testing methods for digital radio receivers are inefficient, lack repeatability and consistency, and are poorly adaptable to complex environments, making it difficult to fully utilize IQ data storage and playback for adaptive testing.

Method used

A method based on IQ signal sampling, storage, and playback is adopted, combined with inverse reinforcement learning to construct a test strategy. Expert trajectories are formed through IQ sample data, reward functions and state-action value functions are constructed, random test strategies are generated, and the test process is optimized to determine digital sensitivity.

Benefits of technology

It improves testing efficiency, reduces test points, enhances test consistency and repeatability, enables stable testing in complex environments, and improves the maintainability and traceability of the testing system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121923743A_ABST
    Figure CN121923743A_ABST
Patent Text Reader

Abstract

The invention discloses a digital sensitivity test method based on IQ signal sampling, storage and playback, and belongs to the technical field of radio measurement. The method comprises the following steps: controlling to input a test signal of power to a tested radio station, calculating a bit error rate to demodulate a valid data proportion, and forming and storing an expert track; performing inverse reinforcement learning on the expert trajectory to obtain reward function parameters; estimating a state transition probability, an action value function and a state value function based on the expert trajectory and / or a new trajectory obtained by a subsequent test; and acquiring a random test strategy: iteratively selecting an input power gear according to the random test strategy and carrying out test to generate a new track update state transition probability and / or a value function until a convergence condition is met, and determining the minimum input signal power meeting a set threshold condition as the receiver digital sensitivity of the tested radio station. According to the invention, the digital sensitivity of the radio station receiver to be tested is determined with fewer test points and higher efficiency, and the test consistency and repeatability are improved at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless communication testing and measurement technology, and in particular to a method and apparatus for testing the digital sensitivity of a digital radio receiver based on IQ signal sampling, storage and playback. Specifically, it relates to a technical solution that utilizes IQ sampling and storage, trajectory playback and inverse reinforcement learning to construct an adaptive testing strategy, and determines the receiver's digital sensitivity under the constraints of bit error rate and the proportion of effective demodulated data. Background Technology

[0002] Digital sensitivity of a digital radio receiver is typically used to characterize the minimum input signal power required for the receiver to maintain a certain demodulation quality under specified operating conditions. Current testing methods usually use bit error rate (BER), frame error rate (FER), or audio quality metrics as criteria. A known modulated signal is injected into the receiver under test through a signal source and a programmable attenuator. The input power is gradually reduced, and the BER or frame error rate is statistically analyzed at each power point. When the metric changes from meeting a threshold to not meeting a threshold, the power point near the threshold is taken as the digital sensitivity or its estimate.

[0003] Existing methods have at least the following shortcomings: 1) Low testing efficiency: Conventional tests often employ fixed step sizes or manual experience strategies (e.g., coarse scanning followed by fine scanning), requiring repeated measurements in the threshold neighborhood. This is especially problematic when the device under test (DUT) is experiencing unstable synchronization, re-acquisition, frame synchronization, or a large number of statistical bits, significantly increasing testing time. 2) Strategy reliance on manual experience, resulting in insufficient repeatability and consistency: Different testers may choose different parameters such as power step size, dwell time, and statistical length, leading to insufficient repeatability and comparability of test results. Furthermore, under different models and channel conditions, fixed step sizes often fail to balance efficiency and accuracy. 3) Weak adaptability to complex environments or state drift: When the DUT experiences AGC dynamics, frequency offset / phase noise changes, frequency hopping, or link state drift, the bit error rate and the proportion of effective demodulated data may exhibit lag or non-ideal monotonic characteristics with power changes. Fixed testing strategies are prone to generating invalid or duplicate measurement points. 4) Existing waveform storage and playback methods usually still rely on fixed test strategies: Even if waveform or IQ data can be stored and played back to reproduce real signal scenarios, the selection of power points in the test process often adopts preset scanning or manual adjustment, making it difficult to make full use of historical data to form an adaptive optimal / near-optimal test strategy.

[0004] Therefore, a digital sensitivity testing method is needed that can combine IQ sampling storage and playback capabilities and learn testing strategies from historical trajectories to improve testing efficiency and consistency, and stably determine digital sensitivity under the constraints of bit error rate and effective demodulated data ratio. Summary of the Invention

[0005] The purpose of this invention is to provide a digital sensitivity testing method and apparatus based on IQ signal sampling, storage, and playback. By forming a test trajectory from "input power level - demodulation quality index" and performing playback training, an inverse reinforcement learning is used to construct a reward function and obtain a random testing strategy accordingly. This enables the testing process to determine the digital sensitivity of the radio receiver under test with fewer test points and higher efficiency while meeting the bit error rate and demodulation effective data ratio thresholds, and at the same time improving the consistency and repeatability of the test.

[0006] To achieve the above objectives, this invention provides a digital sensitivity testing method based on IQ signal sampling, storage, and playback. S01: Control the input power to the radio station under test to be... The test signal is sampled from the radio station under test and stored as IQ sample data; the IQ sample data is processed to obtain the received bit sequence, and compared with the reference bit sequence to obtain the bit error rate. Simultaneously calculate the proportion of valid demodulated data. ;by As a state, with As an action, an expert trajectory is formed and stored:

[0007] Where T is the number of data sets;

[0008] S02: Regarding the expert trajectory The corresponding IQ sample data were replayed, and the parameters of the reward function were obtained using inverse reinforcement learning. And construct the state-action reward function:

[0009]

[0010] in, For the preset characteristic function;

[0011] S03: Based on the expert trajectory and / or the state transition probability estimated from a new trajectory obtained from subsequent tests. Action value function With state value function ;

[0012] S04: The random testing strategy is obtained according to the following formula:

[0013]

[0014] S05: According to the random testing strategy Iteratively select the input power level and perform tests to generate a new trajectory. Update the state transition probabilities and / or the value function until the convergence condition is met:

[0015] in, For the threshold, Number of data sets; A set of states;

[0016] and will satisfy and Minimum input signal power under threshold conditions The receiver digital sensitivity of the radio station under test is determined, among which... To preset the bit error rate threshold, This is a preset threshold for the percentage of valid demodulated data.

[0017] To achieve the aforementioned objective, the present invention also provides a digital sensitivity testing device, characterized in that it comprises: a programmable signal source, a programmable attenuator, an IQ sampling and storage module, a memory, and a processor; wherein the programmable signal source and the programmable attenuator are used to input test signals of different power levels to the radio under test, the IQ sampling and storage module is used to sample and store the I-channel and Q-channel signals of the radio under test, and the processor is used to execute the aforementioned digital sensitivity testing method based on IQ signal sampling, storage, and playback to output the receiver digital sensitivity of the radio under test.

[0018] Compared with existing technologies, the digital sensitivity testing method and apparatus for digital radio receivers based on IQ signal sampling, storage, and playback provided by this invention has at least the following beneficial effects: 1) Improved testing efficiency and reduced measurement points: By using a learned random testing strategy, more informative power levels are preferentially selected in the threshold neighborhood, reducing invalid power points and repeated measurements. 2) Improved testing consistency and repeatability: The testing process is abstracted into a state-action trajectory and decisions are made based on a unified learning strategy, reducing reliance on human experience and enhancing the comparability of test results from different personnel and batches. 3) Full utilization of IQ storage and playback capabilities to achieve offline learning and experiment reproduction: By storing and playing back IQ data and trajectories, the testing process can be reproduced and the strategy updated without occupying the RF environment or repeating field conditions, improving the maintainability and traceability of the testing system. 4) Strong scalability: characteristic function It can integrate information such as power change, logarithmic scale of bit error rate, and percentage of effective data; the action set can adopt a combination of coarse and fine-grained level design; the reward function can also introduce a threshold bias term to adapt to different systems and different criteria. Attached Figure Description

[0019] Figure 1This is a block diagram of the digital sensitivity testing device provided by the present invention;

[0020] Figure 2 This is a flowchart of the digital sensitivity testing method for a digital radio receiver based on IQ signal sampling, storage, and playback provided by the present invention. Detailed Implementation

[0021] The present invention will be further described below with reference to embodiments. It should be understood that the following embodiments are used to explain the present invention and not to limit the scope of protection of the present invention; in the absence of conflict, the embodiments and technical features disclosed in this specification can be combined with each other.

[0022] Figure 1 This is a block diagram of the digital sensitivity testing device provided by the present invention, as shown below. Figure 1 As shown, the digital sensitivity testing device includes at least: a programmable signal source: used to generate a test signal containing a known reference bit sequence, with modulation method, symbol rate, shaping filter, etc., matched to the system of the radio station under test; and a programmable attenuator: used to generate discrete power levels and control the input power. IQ sampling and storage module: including analog-to-digital converter and memory, used to sample, buffer and persistently store the I and Q output signals of the radio under test; Processor (and its memory): used to perform operations such as synchronization, demodulation, decoding, error statistics, valid data determination, inverse reinforcement learning, value iteration, policy generation and convergence determination, and output digital sensitivity results to the display.

[0023] In this invention, the IQ sampling and storage module adds a uniform timestamp to each segment of IQ data during sampling. Sampling rate identifier Local oscillator / intermediate frequency configuration identifiers and synchronization quality indicators The processor, during playback demodulation, is based on... Select a synchronization strategy and output the synchronization reliability. By using timestamps and synchronization quality metrics as conditional variables in demodulation statistics, the bit error rate statistics and valid data determination can distinguish between "bit errors caused by insufficient power" and "invalid frames caused by synchronization failure," thus improving the robustness of sensitivity determination. Valid data determination employs multi-criteria gating, including at least: decoding verification results (such as CRC / checksum), frame synchronization success flag, demodulation soft information confidence threshold, and effective payload percentage. Threshold conditions. By jointly gating multiple criteria, it is possible to stably distinguish between "usable data segments" and "failed data segments" even under low signal-to-noise ratio or multipath conditions, and to avoid mistakenly including synchronization failures or demodulation phase-locking anomalies in sensitivity statistics.

[0024] Optionally, the digital sensitivity testing device further includes a power calibration submodule (not shown in the figure), which is used to calibrate at each discrete power level. Next, the insertion loss of the programmable attenuator and the connection link is calibrated, and a range-actual input power mapping table is generated. By tracing and calibrating the power of each gear, subsequent adjustments can be made. The sensitivity results obtained by using cables, connectors and temperature drift as independent variables are reproducible and comparable across devices, and reduce systematic errors introduced by cables, connectors and temperature drift.

[0025] Figure 2 This is a flowchart of the digital sensitivity testing method for a digital radio receiver based on IQ signal sampling, storage, and playback provided by the present invention, as shown below. Figure 2 As shown, the digital sensitivity testing method based on IQ signal sampling, storage, and playback provided by this invention is applied to the digital sensitivity measurement of digital radio receivers. The method is executed by a testing device, which includes at least a programmable signal source, a programmable attenuator, an IQ sampling and storage module, and a processor. The method includes:

[0026] S01: Construct a set of discrete input power levels Each action Corresponding input power:

[0027] in, For minimum test power, For power stepping, K is the number of actions;

[0028] The programmable signal source and programmable attenuator are controlled to input power to the radio station under test. The test signal is obtained from the I-channel and Q-channel outputs of the radio under test. With Q-channel signal and through the IQ sampling and storage module and The data is sampled and stored as IQ sample data; the processor synchronizes, demodulates, and decodes the IQ sample data to obtain the received bit sequence, and compares it with the reference bit sequence to obtain the bit error rate. Simultaneously calculate the proportion of valid demodulated data. The bit error rate and the percentage of valid demodulated data are used together as the status: As a state, with As an action, an expert trajectory is formed and stored:

[0029] ,

[0030] Where T represents the index of the last time step of the expert trajectory.

[0031] S02: Regarding the expert trajectory The corresponding IQ sample data were replayed, and the parameters of the reward function were obtained using inverse reinforcement learning. And construct the state-action reward function:

[0032]

[0033] in For the preset characteristic function, express The transpose of ;

[0034] S03: Based on the expert trajectory and / or the state transition probability estimated from a new trajectory obtained from subsequent tests. The action value function is calculated using the following soft Bellman equation. With state value function :

[0035] ,

[0036] ,

[0037] in E[ ]: discount factor; E[ ]: expectation operator, for possible Find the average; : indicates that the next state will be sampled according to this distribution.

[0038] S04: The random testing strategy is obtained according to the following formula:

[0039] ,

[0040] S05: According to the aforementioned random testing strategy Iteratively select the input power level and perform tests to generate a new trajectory. Update the state transition probabilities and / or the value function until the convergence condition is met:

[0041] ,

[0042] and will satisfy and Minimum input signal power under threshold conditions The receiver digital sensitivity of the radio station under test is determined, among which... To preset the bit error rate threshold, The threshold for the percentage of valid demodulated data is set, and M is the length of the new trajectory; The convergence threshold (a positive number) is used to measure whether the difference in the value function between two iterations is small enough.

[0043] This invention uses input power levels Compared with actual input power By binding and sampling, storing, and replaying the IQ data of the radio under test, the testing process is upgraded from "real-time one-time measurement" to "repeatable digital testing process", thereby achieving the goal that the same test data can be verified, replayed, and compared.

[0044] This invention calculates the bit error rate by demodulating the IQ samples to obtain the received bit sequence and comparing it with a reference bit sequence. Simultaneously calculate the proportion of valid demodulated data. This allows the sensitivity assessment to simultaneously reflect both the "bit error rate" and the "effective demodulation level," thereby reducing the risk of misjudging synchronization failures / invalid frames caused by relying solely on BER and improving the reliability of the judgment.

[0045] This invention is achieved by using Build state, with Construct actions and form expert trajectories This allows the traditional process of "searching power levels based on human experience" to be explicitly digitized into a learnable decision sequence, thereby enabling expert experience to be transferred, reused, and used for subsequent strategy generation.

[0046] This invention replays expert trajectories and their corresponding IQ samples, and uses inverse reinforcement learning to obtain reward parameters. , build This allows implicit preferences such as "compliance tendency / test efficiency / power cost" to be quantified into the reward model, thereby reducing the need for manually setting heuristic rules and improving the adaptability and generalization ability of strategy generation.

[0047] This invention estimates state transition probabilities based on expert trajectories and / or novel trajectories. And iteratively update the action value function. With state value function This allows the testing strategy to be continuously revised using historical statistical patterns, thereby achieving stable convergence and reducing invalid test rounds even when the channel fluctuates or the equipment status changes.

[0048] This invention is achieved by using and Computing random testing strategies This allows the strategy to automatically balance between "exploring undertested power levels" and "utilizing known better power levels," thereby avoiding getting stuck in local experience and improving the efficiency of finding the minimum usable input power.

[0049] This invention generates new trajectories iteratively according to a random testing strategy. And in satisfying The test process terminates at a certain time, thus providing a clear and calculable convergence stopping criterion, thereby reducing test time and computational overhead and increasing automation while ensuring accuracy.

[0050] This invention achieves by satisfying and Under the threshold condition, the minimum input power will be By defining the sensitivity as digital, the output results are strictly aligned with the preset quality threshold and are verifiable, thereby achieving a clear definition of the sensitivity index, acceptable results, and consistent comparison across different devices / batches.

[0051] In this invention, the test signal is a standard frame signal or a pseudo-random sequence signal containing a known reference bit sequence, and the bit error rate is... It is obtained by comparing the received bit sequence with the reference bit sequence bit by bit.

[0052] This invention introduces a known reference bit sequence, thereby providing a unique benchmark for bit error statistics, which in turn ensures that the bit error rate calculation results are objective and consistent, and allows for direct comparison between different test batches and different devices.

[0053] This invention employs standard frame signals or pseudo-random sequence signals, thereby making the test signals representative and controllable in terms of frame structure and statistical characteristics. This enables the test signals to cover typical modulation / coding / shaping filtering conditions and improve the versatility and reproducibility of test scenarios.

[0054] This invention calculates by bit-by-bit comparison. This allows the bit error rate to be determined directly by bit-level differences, thereby avoiding deviations introduced by frame-level judgment, soft decision thresholds, or intermediate indicators, and improving the accuracy and reliability of bit error statistics.

[0055] This invention obtains by bit-by-bit comparison As input to the state variables, subsequent learning and policy iteration enable policy updates to be based on fine-grained, differentiable trend error information, thereby achieving more stable power level search convergence and reducing invalid test rounds.

[0056] In this invention, the percentage of effective demodulated data... Calculate as follows:

[0057] ,

[0058] in, For the first The number of data bits or data frames in the group of data that are determined to be valid through frame synchronization, CRC check, or decoding check. For the first The total number of bits or frames in the data set.

[0059] This invention uses frame synchronization, CRC check, or decoding check as validity criteria to statistically analyze... This ensures that the "availability" judgment aligns with the actual reception effectiveness of the communication system, thereby eliminating interference from invalid data segments such as synchronization failures and decoding failures on performance statistics and improving the reliability of sensitivity assessment. The proportion is obtained by normalizing the effective data volume to the total data volume. This allows for direct comparison of results from different test rounds under varying sampling lengths and frame counts, thereby enhancing the scale consistency and cross-round comparability of the metrics. By allowing... and The method uses two granularities, "bit count" or "frame count," to make the percentage metric adaptable to different systems (continuous bit stream or frame structure services) and different statistical needs, thereby improving the method's versatility and engineering feasibility.

[0060] In this invention, the discrete input power level set From the input power range Power Step It is obtained by quantization, or by combining the coarse measurement range set and the fine measurement range set, wherein the fine measurement range set covers the threshold power neighborhood obtained by coarse measurement.

[0061] This invention utilizes the continuous input power range In steps Quantization constructs discrete gear sets This discretizes the test action space and gives it clear boundaries, thereby facilitating policy generation and value iteration calculation, reducing search complexity, and improving feasibility. This is achieved by setting power steps. This allows for configurable sensitivity measurement resolution, achieving a trade-off between test accuracy and test time, and adapting to engineering needs of different standards or devices under test. By employing a hierarchical combination structure of "coarse measurement range set + fine measurement range set," the test first rapidly locates the threshold power with a larger step size, then performs a fine search within the threshold power neighborhood with a smaller step size. This significantly reduces the number of test rounds and time consumed by full-range fine-step scanning, while maintaining the judgment accuracy near the threshold. By having the fine measurement range set cover the threshold power neighborhood obtained from the coarse measurement, the determination of the minimum usable input power is concentrated in the high-sensitivity region, thereby reducing the probability of misjudgment near the boundary power and improving the stability and reproducibility of the final digital sensitivity results.

[0062] In this invention, the characteristic function It contains at least one or more of the following features: , , , Power change and the intersection of the above features, where To prevent positive numbers with odd logarithms.

[0063] This invention introduces and As a feature, reward learning and policy updates can simultaneously characterize the error rate and the degree of effective demodulation, thereby avoiding misjudgments caused by evaluating only a single indicator and improving the reliability and robustness of sensitivity judgment.

[0064] By introducing The logarithmic characteristics amplify the differences in the low BER region and make the numerical changes smoother, thereby enhancing the identification ability near the threshold (low BER region) and improving the convergence accuracy of the strategy in the critical power neighborhood; at the same time, by setting Thus avoid The logarithm diverges over time, thereby ensuring the stability of feature values ​​and avoiding numerical instability during training and iteration.

[0065] By introducing input power characteristics This makes the learning objective explicitly include "power cost," thereby making it more likely to choose a lower input power level when the threshold condition is met, thus more effectively approaching the minimum available input power and shortening the testing process.

[0066] By introducing power change characteristics This enables the strategy learning to perceive the power adjustment range in adjacent rounds, thereby suppressing frequent and large jumps in power levels, reducing test oscillations, and improving the stability and reproducibility of the search process.

[0067] By introducing the intersection terms of the above features (e.g.) , , (etc.), so that the reward model can express the coupling relationship between "power-bit error-validity", thereby more accurately depicting the impact mechanism of different power levels on bit error and valid data, and improving the expressive power and generalization ability of the reward function obtained by inverse reinforcement learning.

[0068] In this invention, the inverse reinforcement learning employs maximum entropy inverse reinforcement learning, ensuring that the trajectory probabilities satisfy:

[0069] in The partition function is used; the parameters are obtained by maximizing the log-likelihood function of the demonstrated trajectory. :

[0070]

[0071] in, The dataset consists of the expert trajectories. This is the regularization coefficient.

[0072] This invention models the trajectory distribution using the maximum entropy form, thereby preserving the maximum uncertainty for unobserved factors while satisfying expert demonstration constraints. This allows for the existence of multiple approximately equivalent expert strategies and avoids over-deterministic assumptions, thus improving the model's adaptability to random factors such as channel fluctuations and synchronization contingencies.

[0073] By combining trajectory probability with cumulative reward Exponential correlation and partition function Normalization assigns a higher probability to high-return trajectories and forms a comparable probability scale, thereby quantifying "expert preferences" with a unified probability framework, which facilitates subsequent strategy evaluation and generation.

[0074] Parameters are learned by maximizing the log-likelihood function of the demonstration trajectory. This allows the reward function to fit the implicit test preferences within the expert trajectory in a data-driven manner (such as approaching the threshold power faster and reducing invalid measurements), thereby reducing the reliance on manually set heuristic rules and improving the objectivity and transferability of the reward function.

[0075] By introducing a regularization term This constrains the magnitude of the reward parameters and suppresses overfitting, thereby maintaining learning stability even when the number of expert trajectories is limited or the noise is high, and improving the generalization ability to new trajectories / new devices.

[0076] By using the dataset composed of expert trajectories As a unified training sample set, it supports the cumulative learning of data from multiple tests and various operating conditions, thereby continuously improving the accuracy of reward estimation and enhancing the robust convergence of the testing strategy as the data grows.

[0077] In this invention, the state transition probability The trajectory is obtained by statistical estimation from the expert trajectory and / or the new trajectory, and satisfies the following:

[0078] ,

[0079] in, To transition from state in the trajectory In action Transition to state The count.

[0080] This invention uses statistical estimation based on expert trajectories and / or new trajectories to directly derive the state transition model from actual test data rather than subjective assumptions. This enables data-driven modeling without establishing an accurate channel / receiver analytical model, improving the method's versatility and engineering feasibility.

[0081] By counting transfers Normalized to conditional probability Thus, for a given The sum of all next state probabilities is 1, thereby ensuring the normality of the transition probabilities and their computability for value iteration, and improving the numerical stability of policy evaluation and updating.

[0082] By simultaneously utilizing expert trajectories and new trajectories for updates, the transition probability can be adaptively corrected as the state of the tested radio station changes, environmental fluctuations occur, or the testing strategy changes. This leads to a continuous improvement in the accuracy of transition estimation and enhances the robust convergence of the strategy under dynamic conditions.

[0083] By using a ternary count of "state-action-next state" as the estimation basis, the transition model can explicitly characterize the input power level action pair. , By understanding the influence of state evolution, we can strengthen the data coupling between actions and performance statistics, reduce blind scanning, and improve the efficiency of search threshold power.

[0084] By employing a counting form of transition estimation, it is easier to perform smooth expansion when the sample size is insufficient (e.g., optional addition of prior / Laplacian smoothing), thereby reducing the risk of probability collapse caused by zero counts and improving reliability under small sample conditions.

[0085] In this invention, a temperature parameter is introduced. Then, the state value function and the random testing strategy respectively satisfy:

[0086]

[0087]

[0088] Among them, when When taking smaller values, the random testing strategy tends to be closer to selection. The strategy with the greatest certainty.

[0089] This invention uses temperature parameters Scaling is applied to the log-sum-exp aggregation to adjust the state value function. The system allows for controlled switching between "maximum approximation" and "smooth averaging," thereby preventing the value function from becoming overly sensitive to fluctuations in individual samples near the power threshold and improving the numerical stability of value iteration.

[0090] Through strategy Introduced in China By controlling the entropy of the control strategy, the testing strategy can adaptively balance between "exploring different power levels" and "utilizing the current optimal power level", thereby reducing blind full-range scanning, reducing invalid test rounds, and improving the efficiency of locating the threshold power.

[0091] By when Taking a smaller value allows the strategy to approach a deterministic optimal choice, thus enabling the strategy to focus on selecting high-power options when approaching convergence or having already located the threshold power neighborhood. The power level is adjusted to accelerate convergence and more accurately approximate the minimum available input power (digital sensitivity).

[0092] By when By taking a larger value while maintaining a certain degree of randomness (maximum entropy characteristic), the strategy will not fall into local optima too early or be misled by accidental low-error samples, thereby improving robustness and generalization ability under conditions of channel fluctuations, synchronization randomness, or insufficient samples.

[0093] By As a configurable parameter, it allows for the selection of different exploration intensities under different systems and testing duration constraints, thereby enabling the same method to be flexibly adapted between "high-precision slow testing" and "fast coarse testing" scenarios.

[0094] In this invention, a reward bias term related to the bit error rate threshold is introduced into the state-action reward function, such that:

[0095]

[0096] in, ,

[0097] and , .

[0098] in, State-action feature vector; : Target achievement reward constant; : Penalty constant for not meeting the target.

[0099] This invention explicitly encodes "whether the bit error rate threshold is met" into the reward function as a segmented bias term, thereby enabling policy updates to directly revolve around... The compliance / non-compliance boundary is optimized to enhance the search guidance of threshold drive and reduce invalid test rounds for irrelevant power levels.

[0100] Through Positive rewards will be given in a timely manner. ,exist Punishment at the time This allows the strategy to develop a stable preference for "compliant states" and suppress "non-compliant states," thereby accelerating the approach to the minimum input power that meets the threshold condition, shortening the convergence time, and improving testing efficiency.

[0101] By combining the bias term with the feature weighting term By superimposing these features, while retaining fine-grained feature optimizations such as power cost and effective data ratio, a strong constraint-based threshold guidance is added, thereby achieving a balance between multiple objectives such as "meeting the target first" and "minimum power / stability", improving the ability to express rewards and the controllability of the strategy.

[0102] By adopting constant type and As a threshold bias, this simplifies reward calculation, makes it fast, and insensitive to noise, thereby improving real-time performance during online testing and reducing the impact of instantaneous changes. The risk of volatility causing sharp fluctuations in rewards.

[0103] By It can be set to a positive number and is configurable, so that the intensity of the penalty for "meeting the standard" and "not meeting the standard" can be adjusted according to different standards, thereby adapting to the sensitivity threshold requirements under different systems / test specifications and improving the versatility of the method.

[0104] Obtain a random testing strategy Then, the following process is executed:

[0105] In each round of testing, based on the current state Depend on Sampling to obtain action Control signal source and attenuator to set input power Collect I / Q data and demodulate statistics to obtain new data. and Record transfer A new trajectory is formed:

[0106] Model Update and Convergence Determination: Updating Counts Using New Trajectories With transition probability and update When the convergence condition is met:

[0107] At that time, it was assumed that the strategy and value function were stable.

[0108] After convergence, the digital sensitivity output rule iterates through or explores the power level set obtained by the strategy, and selects power points that meet the following threshold conditions:

[0109]

[0110] The minimum input power that satisfies the condition:

[0111]

[0112] The digital sensitivity of the radio receiver under test was determined.

[0113] In this invention, and It can be set according to the standards or test specifications of the system being tested, for example Statistical length The confidence level can be set to a fixed number of bits or a fixed number of frames, depending on the required confidence level.

[0114] In some systems, the percentage of valid demodulated data can be indirectly reflected by the bit error rate or frame check results. In this case, the state can be simplified to... And accordingly, the threshold condition is simplified to .

[0115] This invention, by adjusting the current state in each round of testing... according to Sampling to obtain action This allows for a probabilistic balance between exploration and utilization in power level selection, thereby avoiding redundant testing caused by fixed-step full scan and improving threshold power positioning efficiency.

[0116] The input power is set by controlling the signal source and attenuator. After collecting I / Q data, demodulation statistics are obtained to obtain new data. and This closes the data link of "action-observation-performance statistics," thereby establishing a calculable causal relationship between power adjustment and changes in error / effectiveness indicators, and improving the interpretability and reproducibility of strategy iteration.

[0117] Transfer by record And form a new trajectory This allows online test data to be continuously accumulated into learnable samples, thereby supporting learning while testing and adaptively updating the model according to changes in operating conditions, and improving the ability to adapt to channel fluctuations and equipment state drift.

[0118] Update the count by utilizing the new trajectory With transition probability and iteratively update This allows strategy evaluation and improvement to be continuously revised based on real observations and statistics, thereby reducing reliance on prior analytical models and improving the versatility and engineering feasibility under different systems and equipment.

[0119] By adopting convergence conditions The decision strategy and value function are stable, thus providing the testing process with clear and calculable stopping criteria. This allows for a reduction in the number of iterations and testing time while ensuring the stability of the results, and avoids overtesting or undertesting.

[0120] After convergence, the set of power levels explored by the strategy is filtered, and based on... and As a quality threshold, the sensitivity judgment is made consistent with the system standards / test specifications, thereby achieving clearly defined, acceptable, and cross-batch comparable output indicators.

[0121] By satisfying the minimum input power of the threshold condition By defining it as digital sensitivity, the final output directly corresponds to the engineering definition of "minimum usable input power," thereby improving the interpretability of the conclusions and reducing the subjectivity of human selection of the power point.

[0122] By allowing and Set according to institutional standards, and with statistical length By setting the confidence level to a fixed number of bits or a fixed number of frames, both the threshold and the statistical confidence level can be configured, thereby achieving flexible adaptation under different specifications and confidence level requirements and improving the statistical reliability of the results.

[0123] By simplifying the state in some systems to The threshold conditions are simplified accordingly, allowing the method to reduce state dimensions and computational complexity in scenarios where the proportion of effective data can be indirectly reflected by bit errors or frame verification. This reduces implementation costs and expands applicability without compromising decision consistency.

[0124] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely illustrative of the principles of the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of this invention is defined by the appended claims and their equivalents.

Claims

1. A digital sensitivity testing method based on IQ signal sampling, storage, and playback, comprising: S01: Control the input power to the radio station under test. The test signal is sampled from the radio station under test and stored as IQ sample data; The received bit sequence is obtained by processing the IQ sample data and then compared with the reference bit sequence to obtain the bit error rate. Simultaneously calculate the proportion of valid demodulated data. ;by As a state, with As an action, an expert trajectory is formed and stored: Where T is the number of data sets; S02: Regarding the expert trajectory The corresponding IQ sample data were replayed, and the parameters of the reward function were obtained using inverse reinforcement learning. And construct the state-action reward function: in, Preset characteristic function; S03: Based on the expert trajectory and / or the state transition probability estimated from a new trajectory obtained from subsequent tests. Action value function With state value function ; S04: The random testing strategy is obtained according to the following formula: S05: According to the random testing strategy Iteratively select the input power level and perform tests to generate a new trajectory. Update the state transition probabilities and / or the value function until the convergence condition is met: in, For the threshold, Number of data sets; A set of states; and will satisfy and Minimum input signal power under threshold conditions The receiver digital sensitivity of the radio station under test is determined, among which... To preset the bit error rate threshold, This is a preset threshold for the percentage of valid demodulated data.

2. The digital sensitivity testing method based on IQ signal sampling, storage, and playback according to claim 1, wherein the test signal is a standard frame signal or a pseudo-random sequence signal containing a known reference bit sequence, and the bit error rate... It is obtained by comparing the received bit sequence with the reference bit sequence bit by bit.

3. The digital sensitivity testing method based on IQ signal sampling, storage, and playback according to claim 1, wherein the effective demodulation data ratio... Calculate as follows: , in, For the first The number of data bits or data frames in the group of data that are determined to be valid through frame synchronization, CRC check, or decoding check. For the first The total number of bits or frames in the data set.

4. The digital sensitivity testing method based on IQ signal sampling, storage, and playback according to claim 1, wherein the discrete input power level set... From the input power range Power Step It is obtained by quantization, or by combining the coarse measurement range set and the fine measurement range set, wherein the fine measurement range set covers the threshold power neighborhood obtained by coarse measurement.

5. The digital sensitivity testing method based on IQ signal sampling, storage, and playback according to claim 1, wherein the characteristic function It contains at least one or more of the following features: , , , Power change and the intersection of the above features, where To prevent positive numbers with odd logarithms.

6. The digital sensitivity testing method based on IQ signal sampling, storage, and playback according to claim 1, wherein the inverse reinforcement learning adopts maximum entropy inverse reinforcement learning, and the trajectory probability satisfies: in The partition function is used; the parameters are obtained by maximizing the log-likelihood function of the demonstrated trajectory. : in, The dataset consists of the expert trajectories. is the regularization coefficient.

7. The method according to claim 1, characterized in that, The state transition probability The trajectory is obtained by statistical estimation from the expert trajectory and / or the new trajectory, and satisfies the following: ,in, To transition from state in the trajectory In action Transition to state The count.

8. The method according to claim 1, characterized in that, Introducing temperature parameters Then, the state value function and the random testing strategy respectively satisfy: ,, , Among them, when When taking smaller values, the random testing strategy tends to be closer to selection. The strategy with the greatest certainty.

9. The method according to claim 1, characterized in that, In the state-action reward function, a reward bias term related to the bit error rate threshold is introduced, such that... in, , and , .

10. A digital sensitivity testing device, characterized in that, include: The system comprises a programmable signal source, a programmable attenuator, an IQ sampling and storage module, a memory, and a processor; wherein the programmable signal source and the programmable attenuator are used to input test signals of different power levels to the radio under test, the IQ sampling and storage module is used to sample and store the I and Q signals of the radio under test, and the processor is used to execute the method described in any one of claims 1 to 9 to output the receiver digital sensitivity of the radio under test.