Beidou passive room dynamic closed-loop calibration method based on reinforcement learning
By applying time-symmetric micro-perturbations to the BeiDou passive indoor distribution system and utilizing gradient confidence discrimination and reinforcement learning, the problems of blind adjustment and uncertainty in complex controlled systems are solved, enabling real-time perception and stable adjustment of local performance changes and providing fault early warning.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-20
- Publication Date
- 2026-03-17
AI Technical Summary
Existing technologies struggle to effectively detect local performance changes in complex, distributed controlled systems, leading to delayed or blind adjustment responses. Furthermore, they lack clear gradient information to guide parameter adjustments, making it impossible to achieve a balance between adjustment efficiency and system stability. In particular, they are unable to cope with disturbances and system drift in dynamic environments.
A reinforcement learning-based dynamic closed-loop calibration method for BeiDou passive indoor distribution systems is adopted. By actively detecting and applying time-symmetric micro-perturbations, the response gradient is monitored. Gradient confidence is used to identify and reinforce learning agents to obtain the local instantaneous causal relationship between control parameters and system performance in real time, thereby achieving directional optimization adjustment.
It enables real-time perception and effective adjustment of local performance changes in dynamic environments, avoids gradient illusion caused by external disturbances, ensures the stability of the control system and the effectiveness of adjustment, and provides early warning and fault prediction of system health status.
Smart Images

Figure CN121165697B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a BeiDou passive indoor distribution dynamic closed-loop calibration method based on reinforcement learning, belonging to the general field of control or regulation system technology. Background Technology
[0002] Currently, dynamic closed-loop calibration for applications such as large-scale indoor wireless signal distribution systems typically involves online adjustment of multiple adjustable control parameters, such as the gain or delay of each signal, to continuously optimize and maintain the stability of one or more system performance indicators. This approach relies on feedback control to address performance degradation caused by environmental changes or system drift. However, as these controlled systems become increasingly widespread in physical space, their structures more complex, and their operating environments more dynamic and uncertain, the inherent limitations of the aforementioned feedback adjustment method, which relies on overall performance indicators, become apparent. When a control... When a system is dealing with a controlled object that has a complex internal state, strong coupling of multiple variables, and is difficult to model accurately, if it relies solely on the system's final output performance indicators, which have been averaged spatially or temporally, such as the average error of regional positioning, for passive monitoring, the feedback information obtained will suffer from fundamental information loss when guiding control decisions. For example, performance improvement in a local area may numerically cancel out performance degradation in another area caused by different factors, resulting in no significant change in the overall average performance indicators of the system. Consequently, the control system cannot effectively perceive the local problems that are occurring, and misses the opportunity for adjustment.
[0003] Furthermore, even if the control system detects a decline in overall performance, its subsequent adjustments lack direct basis in terms of direction and magnitude. Because the impact of adjustments to various control parameters on the final performance index is often highly nonlinear and interconnected, without a clear quantification of the local, instantaneous causal relationship between each control parameter and system performance, any adjustment based on a global optimization algorithm is essentially an inefficient search in a multidimensional space. This adjustment method not only struggles to guarantee convergence efficiency, but its coupling effect also increases the risk of system oscillations. It may even cause performance degradation in other related parts while attempting to fix one problem. Simply improving the computational power of the optimization algorithm cannot fundamentally compensate for the lack of decision-making basis due to insufficient input information dimensions. Moreover, existing technologies in control... At the methodological level, there are also inherent defects in the inability to effectively cope with dynamic environmental interference and changes in the system's own state. For example, Chinese invention patent application CN102012498A discloses a Beidou passive positioning receiver. Although it achieves passive positioning, its core technology lies in receiving signals from three satellites and combining them with external elevation data to perform positioning calculations. Its positioning process is essentially an open-loop calculation based on a static model. This method lacks an effective online perception and closed-loop calibration mechanism for dynamic performance degradation caused by environmental changes or system component drift. When there is instantaneous strong interference in the external environment, its positioning accuracy will be significantly affected, and it cannot guarantee service stability in complex dynamic environments. This further highlights the importance of establishing a method that can identify interference and perform real-time dynamic closed-loop adjustment of the system.
[0004] Existing technologies have several limitations: 1. The overall average performance index relied upon by the control system can mask local and asynchronous performance changes within the controlled object, leading to delayed or missing adjustment responses; 2. Under conditions where the controlled object model is unknown and multiple variables are strongly coupled, the control system lacks explicit gradient information to guide parameter adjustment, resulting in blind and uncertain adjustment actions; 3. Any attempt to overcome the lack of information through optimization algorithms cannot change the black-box nature of the control process, making it difficult to achieve an effective balance between adjustment efficiency and system stability. Therefore, the technical problem to be solved by this invention is how to establish an online mechanism in a complex, distributed, and model-unknown controlled system to obtain the local instantaneous causal relationship between the adjustment of each control parameter and the overall system performance in real time, thereby transforming global search-based optimization into directed gradient-guided optimization adjustment. Summary of the Invention
[0005] This invention provides a BeiDou passive indoor distribution dynamic closed-loop calibration method based on reinforcement learning. Its main purpose is to solve the problem of how to obtain the local and instantaneous causal relationship between control parameters and system performance online in complex controlled systems with unknown models, so as to overcome the problem of blind adjustment and uncertainty caused by information lack in traditional feedback control.
[0006] To achieve the above objectives, this invention provides a reinforcement learning-based dynamic closed-loop calibration method for BeiDou passive indoor distribution systems, comprising the following steps executed by a processor:
[0007] Step a: Before performing active detection on at least one adjustable system parameter in the BeiDou passive indoor distribution system, obtain a preliminary benchmark performance index.
[0008] Step b, perform active detection, which includes: based on the current parameter values of the adjustable system parameters, sequentially apply a positive perturbation and an equal negative perturbation to form a time-symmetric perturbation action;
[0009] Step c: During the perturbation action, monitor the positioning performance metrics of the system to calculate the response gradient;
[0010] Step d: After the perturbation action ends and the current parameter values are restored, obtain a post-baseline performance index.
[0011] Step e, perform gradient confidence screening, which includes: comparing the preceding benchmark performance index with the following benchmark performance index, and determining the response gradient as a valid gradient only when the difference between the preceding benchmark performance index and the following benchmark performance index is less than a preset threshold.
[0012] Step f: Based on the effective gradient, construct a state vector that characterizes the current local response of the system.
[0013] Step g: The state vector is used as the input to the reinforcement learning agent, which outputs a calibration action to adjust one or more adjustable system parameters.
[0014] Step h involves performing the calibration action and then looping back to step a.
[0015] Preferably, the method further includes: storing a series of state vectors generated at different time points by step f to form a state vector time series; performing trend analysis on the state vector time series to identify a long-term unidirectional drift trend of at least one component in the state vector; and generating a prognostic diagnostic information based on the long-term unidirectional drift trend, the prognostic diagnostic information being used to indicate that there is a gradual performance degradation of the hardware unit corresponding to at least one component.
[0016] Preferably, the method further includes: evaluating the gradient signal-to-noise ratio of the detection process based on one or more effective gradients determined in step e. And based on the gradient signal-to-noise ratio The magnitudes of the positive and negative perturbations applied during subsequent execution step b are dynamically adjusted; where the gradient signal-to-noise ratio... The calculation follows The rules, among which, This is the average magnitude of multiple valid gradients acquired within a preset time window. represents the standard deviation of the magnitudes of multiple effective gradients.
[0017] Preferably, the reward function of the reinforcement learning agent is calculated based on the product of the adjustment amount corresponding to the calibration action and the effective gradient. When the product is positive, a positive reward is given, and when the product is negative, a negative reward is given.
[0018] Preferably, between step f and step g, the method further includes: calculating the information entropy of the state vector consisting of one or more effective gradients; determining the calibration action as a preset gradient descent action when the information entropy is less than a preset information entropy threshold; and outputting the calibration action by the reinforcement learning agent only when the information entropy is not less than the preset information entropy threshold.
[0019] Preferably, the cyclic execution of steps a to h is triggered when the system's positioning performance index is detected to be lower than a preset performance threshold.
[0020] Preferably, trend analysis of the state vector time series further includes: analyzing the gradient components of multiple parameters that are geographically adjacent or functionally related, and whether their amplitudes show a trend of increase, in order to identify regional instability symptoms.
[0021] Preferably, the response gradient in step c is calculated based on the difference between the performance index monitored during the application of a positive perturbation and the performance index monitored during the application of a negative perturbation.
[0022] Preferably, the state vector also includes effective gradient information determined at historical moments to characterize the dynamic change trend of the system response characteristics.
[0023] Preferably, in step e, if the difference between the current benchmark performance index and the subsequent benchmark performance index is not less than a preset threshold, the currently calculated response gradient is discarded, and steps a to e are immediately re-executed.
[0024] Compared with the prior art, the beneficial effects of the present invention are:
[0025] 1. By applying a symmetrical perturbation to a single adjustable parameter in the system and simultaneously monitoring changes in system performance indicators during this action, a response gradient characterizing the instantaneous correlation between parameter adjustment and performance change is calculated. Subsequently, one or more such response gradients are constructed into a state vector, which is then input into a reinforcement learning agent, which outputs a calibration action. This approach transforms the information basis for control system decision-making from the traditional macroscopic average performance indicators characterizing the past overall operating state of the system to an information field that can reflect the local instantaneous causal relationship between each control input and system output in real time. This makes the control adjustment process no longer a passive response based on historical statistical results, but a direct optimization based on the current internal response characteristics of the system, avoiding the problems of blind adjustment and uncertainty caused by the mutual coupling of multiple variables or the performance averaging effect.
[0026] 2. Before applying the symmetrical perturbation action, a pre-reference performance index is obtained, and after the action ends and the parameters are restored, a post-reference performance index is obtained. Only when the two reference indices are consistent within a preset range is the response gradient calculated during the period determined as a valid gradient for state characterization. This mechanism compares the temporal symmetry of the active probing action with the asymmetry of external random disturbances, introducing a built-in criterion for the control system to distinguish the reliability of information sources. In this way, the system has the ability to distinguish the performance changes caused by its own probing behavior from the performance fluctuations caused by irrelevant external events, avoiding the gradient illusion caused by instantaneous environmental changes from misleading the decision-making process, and ensuring the stability of the entire closed-loop control system in an open dynamic environment.
[0027] 3. The method of this invention stores the effective state vectors generated at different time points to form a time series, and performs trend analysis on the series to identify the long-term change trends of specific components, thereby generating prognostic diagnostic information characterizing the long-term health status of the system. This reuses the gradient information originally used for real-time tactical calibration as a long-term observation reflecting the state evolution of the controlled physical system itself. When a hardware unit in the system experiences progressive aging or performance degradation, the adjustment effort made by the control system to maintain apparent performance will be recorded in the time series in the form of a continuous unidirectional drift of the corresponding gradient component. This allows the method to provide early warning of potential failure risks of the physical system while completing the dynamic calibration task. This approach extends the capabilities from functional maintenance to health status diagnosis. Simultaneously, based on one or more calculated response gradients, the signal-to-noise ratio (SNR) of the detection process is evaluated, and the amplitude of the perturbation applied during the next active detection is dynamically adjusted according to the evaluation results. This establishes a regulatory closed loop concerning the detection behavior itself, outside of the control closed loop, enabling the system's observation methods to be adaptive. When the system is in a stable state, the perturbation amplitude is automatically reduced to minimize the impact on normal service; conversely, when system performance deteriorates and response becomes sluggish, the amplitude is automatically increased to obtain gradient information with a higher SNR. This achieves a dynamic balance between the effectiveness of the detection behavior and its impact on the system, improving the method's universality under different operating conditions. Attached Figure Description
[0028] Figure 1 This is a flowchart illustrating the self-verifying dynamic closed-loop calibration method of the present invention.
[0029] Figure 2 This is an image showing the offline calibration results of the key parameters actively detected by this invention;
[0030] Figure 3 This is a deployment architecture diagram of the system of the present invention in a multi-regional physical environment. Detailed Implementation
[0031] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0032] This invention discloses a reinforcement learning-based dynamic closed-loop calibration method for BeiDou passive indoor distribution systems. The method is implemented as a processor-executed, iterative closed-loop adjustment process. This process mainly consists of steps such as active detection and response gradient acquisition, gradient confidence discrimination, state representation, and reinforcement learning decision-making and execution. It is used to identify the local and instantaneous causal relationships between control inputs and system states in a multivariable coupled controlled system, transforming global search-based optimization into a gradient-guided adjustment method. In the operation of large indoor wireless signal distribution systems, environmental changes or system component drift often lead to local positioning performance degradation, while existing adjustment methods rely on monitoring the overall average performance indicators of the system. Measurements often fail to effectively detect and respond to such localized, asynchronous performance changes. To address this issue, the present invention employs an active detection step to acquire the system's local response characteristics. This step targets an adjustable system parameter in the BeiDou passive indoor distribution system, such as the gain of a specific signal channel, and performs a time-symmetric perturbation. Specifically, this action involves sequentially applying a positive perturbation and an equal negative perturbation based on the current value of the adjustable system parameter. During the perturbation period, changes in the system's positioning performance indicators are monitored. The response gradient corresponding to the adjustable system parameter is then calculated based on the difference between the performance indicators monitored during the positive and negative perturbation periods. For example, if a gain parameter... The current value is dB, setting the disturbance amplitude for dB, then the active detection action is to sequentially set this parameter value to dB and maintain for a time window T, then set to dB and maintain the same time window T, finally recovering to dB, if the average positioning error of the system monitored within these two time windows T are respectively m and m can be determined according to the formula The response gradient of this parameter is calculated. The calculation of this response gradient, m / dB, provides a quantitative and causal basis for subsequent parameter adjustment decisions.
[0033] In dynamic operating environments, external random events unrelated to control actions may interfere with performance indicators, leading to distorted calculated response gradients, i.e., gradient illusions, which can mislead control actions. To address this, the present invention establishes a gradient confidence discrimination procedure. This procedure compares the time symmetry of active detection actions with the asymmetry of external random disturbances, introducing a criterion for discriminating the reliability of information sources for the control system. Specifically, before performing active detection, a preliminary benchmark performance indicator is obtained. And after the perturbation action ends and the parameters are restored to their current values, a post-baseline performance index is obtained. Subsequently, the preceding and following benchmark performance metrics are compared, and only when the absolute value of the difference between them is... Less than a preset threshold Only when the time is right will the response gradient calculated during that period be determined as a valid gradient; a preset threshold is used. The value can be determined through a standardized offline calibration process. This involves continuously monitoring the natural fluctuations of system performance indicators over a period of time (e.g., 30 minutes) during the initial system deployment phase, without applying any perturbations, and calculating their standard deviation. and set a preset threshold Set it to a multiple of that standard deviation, for example ,when If a significant external disturbance occurs during the detection process, the currently calculated response gradient is discarded, and the system re-executes the detection steps. This discrimination procedure enables the system to proactively distinguish between performance changes caused by its own detection behavior and fluctuations caused by external events, ensuring the reliability of the decision-making basis.
[0034] After acquiring one or more effective gradients, in order to achieve a comprehensive characterization of the system state and make adjustment decisions, the method of this invention constructs one or more effective gradients into a state vector characterizing the current local response characteristics of the system. The state vector This is used as input to a reinforcement learning agent; the reinforcement learning agent, such as an actor-critic model based on a deep neural network, is trained through offline simulation or online learning to master a mapping policy from the state vector space to the calibrated action space. The agent receives the state vector... Then, it outputs a calibration action for adjusting one or more adjustable system parameters. ,in This refers to the specific adjustment amount for the i-th parameter; during the training process of the agent, its reward function is set as the adjustment amount corresponding to the calibration action. With effective gradient product When the value is positive, a positive reward is given; otherwise, a negative reward is given. This reward mechanism prompts the agent to learn how to make effective optimization adjustments based on gradient information. Finally, the system executes the calibration action output by the agent and returns to the active exploration step in a loop to continuously adapt to the dynamic changes in the system state.
[0035] Example 1: This example demonstrates the application of a technical solution in a specific industrial control scenario. In a large automated logistics warehouse consisting of multiple zones, a Beidou passive indoor distribution system provides navigation signals for an autonomous mobile robot (AGV) cluster. The signal amplifier in zone A experiences slow gain thermal drift due to continuous operation, resulting in a slight and persistent degradation of signal coverage quality in that area. Simultaneously, an AGV fully loaded with metal containers rapidly passes through the adjacent zone B, causing a momentary asymmetric disturbance to the signal environment in zone B. Under these conditions, if a control system relies solely on the warehouse's overall average positioning error as feedback, its monitored overall performance indicators show no significant change because the performance degradation in zone A is numerically offset by random fluctuations in zone B or other areas. This results in the control system failing to respond to the progressive problem occurring in area A. The method of this invention is then triggered. First, an active detection step is performed on the adjustable gain parameter in area A where performance degradation occurs. Before and after the application of a time-symmetric perturbation, the pre-reference performance index and post-reference performance index obtained by the controller are basically consistent. Based on this, the gradient confidence discrimination step determines a weak non-zero response gradient calculated during the detection period as a valid gradient. When the method performs active detection on the gain parameter in area B as planned, the passage of the AGV causes a jump in the positioning performance index during the perturbation, resulting in a significant deviation of the post-reference performance index from the pre-reference performance index after the detection ends. Based on this, the gradient confidence discrimination step determines that the detection was contaminated by an external event and discards the response gradient calculated during the period.
[0036] The coordinated operation of the aforementioned active detection step and gradient confidence screening step ensures that the information used to construct the state vector and input it into the reinforcement learning agent only includes data that has passed reliability verification. This state vector represents that the gain parameter in region A has a definite optimization direction, while the state of the parameters in region B is not included in the decision-making basis due to unreliable information. Based on this state vector, the reinforcement learning agent outputs a calibration action that only slightly improves the gain parameter in region A, while not performing any operation on region B. After this calibration action is performed, the signal quality in region A is compensated, and the performance is restored to a normal level. At the same time, the system avoids making an erroneous adjustment to the parameters in region B that could cause signal oscillations by refusing to use contaminated gradient information. When facing distributed, multivariable, and strongly disturbed controlled objects, this closed-loop adjustment method does not rely on building a complete system model for its control behavior. Instead, it continuously confirms the local effectiveness of the control action through an online self-verifying information acquisition and decision-making process.
[0037] Example 2: To objectively verify the improvement effect of the method of the present invention on the stability of the closed-loop control system under the condition of transient external disturbance, a numerical simulation test platform was built in this example. The platform simulates a Beidou passive indoor distribution system with a single adjustable system gain parameter. Its function is to continuously compensate for a preset, slow system parameter drift through closed-loop adjustment based on a quantifiable positioning performance index. The platform can inject an asymmetric transient disturbance simulating a sudden change in the external environment into the positioning performance index at a specified time point. The test is used to compare and analyze the operating performance of two control methods with and without gradient confidence discrimination steps. The test sets up a control group and the sample group of the present invention. The control method of the control group includes active detection, response gradient calculation and reinforcement learning decision, but all calculated response gradients are considered valid. The sample group of the present invention adopts the complete technical method disclosed, including the gradient confidence discrimination step. The test is carried out for 10 adjustment cycles. The true value of the system gain parameter is set from the initial value. Starting at dB, it decreases linearly with each cycle. dB is used to simulate component performance degradation. During the active detection window in the 5th adjustment cycle, a transient disturbance is superimposed on the positioning performance index. The fluctuation amplitude of the performance index caused by this disturbance is greater than the normal fluctuation caused by parameter perturbation. For the sample group of this invention, the preset threshold for gradient confidence discrimination is... The standard deviation of the system's performance index under interference-free conditions was set to twice the standard deviation of the performance index. During the fifth adjustment period of the experiment, the control group calculated a response gradient that was abnormally induced by interference during the detection period and used it for decision-making, causing its reinforcement learning agent to output a large error gain compensation action. The sample group of this invention conducted detection at the same time, and its gradient confidence discrimination step detected that the difference between the pre-benchmark performance index and the post-benchmark performance index exceeded a preset threshold. Therefore, the currently calculated response gradient is deemed invalid and discarded, and no adjustment action is performed; the operating status and performance data records of the two test groups within 10 adjustment cycles are shown in Table 1.
[0038] Table 1: Performance comparison of the two control methods under drift and transient disturbances in the simulated system.
[0039]
[0040] As shown in Table 1, the parameter values of the control group deviated from the true values in the 5th cycle due to incorrect adjustment, resulting in a persistently high performance error in the following cycles, indicating control instability. In contrast, the parameter values of the sample group of this invention closely followed the changes in the true parameter values throughout the entire experiment, and its performance error remained at a low and stable level, unaffected by significant transient disturbances. The experimental results demonstrate that the introduction of the gradient confidence screening step enables the closed-loop control system to effectively identify and reject feedback information contaminated by external transient events, avoiding decision-making misguidance caused by erroneous gradients.
[0041] To further clarify from a mechanistic perspective how the absence of the gradient confidence screening step described in this invention directly leads to the failure of conventional online gradient detection methods in dynamic environments, the following comparative example 1 is provided.
[0042] Comparative Example 1: This comparative example aims to simulate a conventional technical path that a person skilled in the art might use when attempting to solve the background technical problem. That is, based on introducing active detection to obtain the response gradient, the gradient is directly used for feedback adjustment. However, it does not include the key gradient confidence discrimination step of this invention. The experimental conditions of this comparative example are completely consistent with those of Example 2: the same numerical simulation test platform is used to simulate a Beidou passive indoor distribution system with a single adjustable system gain parameter; the true value of the system gain parameter is also set to start from an initial value of 10.0dB and decrease linearly by 0.1dB in each adjustment cycle to simulate the same component performance degradation; and in the active detection window of the 5th adjustment cycle, a transient disturbance simulating a sudden change in the external environment, exactly the same as in Example 2, is injected into the positioning performance index. The only difference between this and Example 2 is that the control method used in this comparative example does not perform any confidence discrimination after performing active detection and calculating the response gradient. Instead, it directly regards the response gradient as a valid gradient and uses it as the input for subsequent reinforcement learning agent decision-making.
[0043] In the first to fourth adjustment cycles, due to the absence of external interference, the comparative method was able to calculate a basically accurate response gradient, and its parameter adjustment results were close to those of Embodiment 2 of the present invention, basically following the slow drift of the true parameter values. In the fifth adjustment cycle, when an external transient disturbance was injected, the comparative control method, during active detection, incorrectly attributed the drastic fluctuations in performance caused by the disturbance to its own perturbation action, thus calculating a response gradient with an abnormal amplitude and incorrect direction. Due to the lack of a verification step, this illusory gradient was directly input to the reinforcement learning agent, causing the controller to output a large... The erroneous gain compensation action caused the system parameter value to be incorrectly adjusted from nearly 9.6dB to 7.25dB. In the 6th to 10th adjustment cycles, although the external disturbance had disappeared, the system parameter had deviated significantly from the true operating point due to the erroneous adjustment in the 5th cycle. Although the control system subsequently attempted to correct back based on the gradient calculated under disturbance-free conditions, it required multiple adjustment cycles to slowly and gradually correct the previous huge deviation. During this period, the system's positioning performance error remained at an unacceptably high level, exhibiting significant control instability and adjustment oscillations. Detailed cycle-by-cycle performance data are recorded in Table 2.
[0044] Table 2: Performance record of the conventional control method in Comparative Example 1 under drift and transient disturbances in the simulated system.
[0045]
[0046] Experimental results show that, in the absence of a gradient confidence discrimination mechanism, the conventional technical approach of simply using active detection to obtain gradients cannot withstand transient disturbances from the external environment in its design principle. The system will make catastrophic decision-making errors due to gradient illusion, causing the closed-loop control system to lose stability in a dynamic open environment.
[0047] Example 3: This example combines Figures 1 to 3 This section explains the reinforcement learning-based dynamic closed-loop calibration method for BeiDou passive indoor distribution systems, such as... Figure 1As shown, this process is triggered by system performance monitoring. When the positioning performance index is detected to be lower than a preset threshold, a calibration loop is initiated. First, a pre-test performance index is acquired to record the system's stable state before detection. Then, active detection is performed, which involves applying a time-symmetric perturbation to an adjustable parameter and calculating the response gradient that quantifies the relationship between parameter adjustment and performance change during this period. After detection, a post-test performance index is acquired to record the system's stable state after detection. Next, the core gradient confidence screening stage is entered. By judging whether the difference between the pre-test and post-test performance indices is less than a preset threshold, it is confirmed whether the detection process has been affected by external interference. If the difference is not less than the threshold, then... If a gradient is deemed invalid, it is discarded and the probing is repeated. If the difference is less than a threshold, the gradient is confirmed as valid. Based on one or more valid gradients, a state vector characterizing the current local response of the system is constructed. This state vector is input into the reinforcement learning agent, which decides and outputs a calibration action. The system then executes this action to adjust one or more system parameters. On the other hand, valid state vectors that have passed confidence level identification can be stored to form a state vector time series. By performing trend analysis and diagnosis on this series, progressive hardware performance degradation can be identified, and prognostic diagnostic information can be generated, thereby outputting a warning of potential fault risks.
[0048] like Figure 2 As shown in the figure, the horizontal axis represents the perturbation duration T in milliseconds (ms), and the vertical axis represents the effective gradient recognition rate in percentages (%). The figure displays two curves, representing the trend of the effective gradient recognition rate with perturbation duration T under the conditions of a selected perturbation amplitude (solid line) and a contrasting perturbation amplitude (dashed line), respectively. It can be seen from the figure that under both perturbation amplitudes, the recognition rate increases with time T, and the recognition rate under the selected perturbation amplitude is better than that under the contrasting perturbation amplitude at most time points. A vertical dashed line marks the final selected initial operating point, with a corresponding perturbation duration of 200 ms. This operating point was chosen to minimize the impact on normal system service while ensuring a high recognition rate. Figure 3As shown, the entire system is logically divided into an upper-layer cloud control and operation and maintenance center and an lower-layer physical deployment environment. The cloud center houses a central controller, which integrates a reinforcement learning agent and a gradient confidence discrimination module, and is connected to an effective gradient time series database for storing historical data. The data in this database is used for long-term trend analysis, and the analysis results are sent to the operation and maintenance management platform to be presented to the operation and maintenance personnel in the form of prognostic diagnosis or alarm information. The physical deployment environment is divided into several areas, such as area A, area B, and area C, which has persistent strong interference. Each area is equipped with adjustable hardware units such as signal amplifiers, as well as devices such as autonomous mobile robots or mobile terminals. The central controller implements control by sending detection and calibration actions to the adjustable hardware units in the physical environment, while the autonomous mobile robots in the deployment environment feed back performance indicators to the central controller in real time, thus forming a complete dynamic closed-loop control and diagnostic system.
[0049] Example 4: This example illustrates the calibration method for key operating parameters and the switching procedure for decision logic in the aforementioned technical solution, thereby addressing the technical problem of parameter setting during system deployment. Before system deployment, the perturbation amplitude in its active detection mechanism needs to be determined. The initial operating point is determined by the duration T of the perturbation. Its selection involves a technical trade-off between ensuring effective detection of the probe signal and minimizing the disturbance to normal system service caused by the probe action. To determine these parameters, this embodiment employs a standardized offline calibration procedure performed in a controlled test environment. First, the gain is adjusted to a selected adjustable system parameter. Inject a known weak linear drift rate, i.e. dB / min; subsequently, different perturbation amplitudes were used. (exist dB to Within dB range, dB is the step size) and the perturbation duration T (in ms to Within the ms range, Combinations of steps (ms being the step size) are used to repeatedly perform the active probing step, and the calculated response gradient for each combination is recorded. ;Calculate gradient by comparison Determine a minimum theoretical gradient corresponding to the known drift that can stably reflect the true response of the system. Combinations with the shortest T, such as dB and T ms, as an initial parameter for the online operation of the system, this procedure provides a reproducible method for determining the initial value of the parameter.
[0050] To adapt to changes in signal-to-noise ratio during system online operation, the method of this invention also includes a perturbation amplitude... The dynamic adaptive adjustment closed loop operates on a fixed time scale, such as an evaluation window of 100 adjustment cycles. The steps are as follows: First, based on multiple effective gradients determined by the gradient confidence discrimination step within this period, the gradient signal-to-noise ratio (GSNR) of the detection process is evaluated. Its calculation follows... The rule, where g is the mean magnitude of multiple valid gradients acquired within a preset time window, The magnitude standard deviations of multiple effective gradients are used; subsequently, the gradient signal-to-noise ratio (GSNR) is calculated based on its current value and a preset target value. ( The comparison results (set to 5.0) dynamically adjust the perturbation amplitude applied during the next active detection. If the current If it falls below the target value, the increase will be proportional. Conversely, the value decreases, and this adjustment loop is used to dynamically adjust the parameters of the detection behavior. To coordinate the computational overhead of complex decision-making by the reinforcement learning agent with the need for rapid response under simple conditions, the present invention introduces a decision-making pattern arbitration mechanism based on state vector information entropy before the reinforcement learning decision-making step. Specifically, the state vector is composed of one or more effective gradients. Then, the information entropy H of the vector is calculated first. If the gradient field has a sharp shape, that is, a few gradient components have large absolute values while others are small, the information entropy H is low, indicating that the main factors affecting the system performance degradation are clear. If H is less than a preset information entropy threshold, the system will bypass the reinforcement learning agent and directly execute a preset gradient descent action, that is, select the parameter with the largest gradient and adjust it by a fixed step size along the gradient direction. Only when the gradient field has a flat shape, that is, all gradient components have similar absolute values and random directions, and the information entropy H is high and not less than the information entropy threshold, will the reinforcement learning agent be activated to make a decision. This arbitration mechanism switches between two decision modes based on the information entropy of the state vector.
[0051] Example 5: This example aims to illustrate the function of using the effective gradient information generated by the method during long-term operation to perform trend analysis in order to achieve the prognostic function of system health status; In a Beidou passive indoor distribution system that has been deployed and continuously operated for more than six months, the controller is configured to store the effective gradients corresponding to all adjustable system parameters determined by the gradient confidence screening step at different time points, forming an effective gradient time series database with timestamps. A background diagnostic module performs trend analysis on the data in this database with a 24-hour cycle.
[0052] At a certain point, the diagnostic module's analysis algorithm detected that the 30-day moving average of the effective gradient of the gain adjustment parameter corresponding to amplifier unit A-07 in the system showed a continuous unidirectional positive drift over the past three consecutive analysis periods, and its slope exceeded a preset statistical threshold. Simultaneously, the algorithm analyzed the effective gradient time series of the corresponding parameters of units A-06 and A-08, which are geographically adjacent to A-07, and unit B-02, which is associated with it on the signal link, and found no synchronous trend of increase. Based on this, the system determined that unit A-07 itself had a high probability of progressive performance degradation, rather than regional instability, and generated a prognostic diagnostic message stating that the control compensation of hardware unit A-07 showed a continuous unidirectional increase, indicating that the unit might have progressive performance degradation and recommending preventative maintenance. This information was pushed to the operation and maintenance management platform. Maintenance personnel conducted on-site inspections and confirmed that the amplifier unit's gain reduction was due to component aging, and replaced it.
[0053] Example 6: This example aims to illustrate the parameter scanning strategy adopted by the method in a large-scale, high-dynamic interference environment with a large number of adjustable parameters, in order to balance detection efficiency and system response speed, as well as the fault-tolerant adjustment mechanism for dealing with persistent strong interference environments. In an indoor distributed system of a transportation hub with hundreds of adjustable gain and delay parameters, if all parameters are subjected to complete sequential polling detection, the entire scanning cycle is too long and cannot respond to rapidly changing environments in a timely manner. To solve this problem, the method of this invention adopts a scanning strategy of grouped polling and dynamic focusing. First, all adjustable system parameters are divided into multiple parameter groups according to their physical location or functional correlation. Under normal operating conditions, the system performs parameter scanning in each adjustable group. Instead of probing all parameters within a cycle, a polling approach is used to perform active probing and gradient calculation on different parameter groups sequentially, maintaining state monitoring of the entire system at a low time cost. When the system detects that the effective gradient magnitude of one or more parameters exceeds the preset attention threshold during the probing of a certain parameter group, the control system will automatically enter dynamic focus mode. In this mode, the system will temporarily suspend polling of other parameter groups and concentrate the probing resources of the following adjustment cycles on continuously probing and calibrating the abnormal parameter group and its adjacent parameter groups at a higher frequency until the overall magnitude of the state vector in that region recovers to a stable level. Then, the system will exit focus mode and resume regular global polling scan.
[0054] Furthermore, when a certain area in the system, such as a security checkpoint, experiences long-term and intense radio frequency interference due to continuous equipment operation, the gradient confidence assessment step for the parameters in that area may fail multiple times consecutively, resulting in the parameter status of that area not being effectively updated for a long period. To address this boundary condition, the method of this invention also includes a fault-tolerant upgrade procedure. When the system detects that the number of consecutive failures in the gradient confidence assessment of a specific parameter exceeds a preset retry limit, such as 5 times, the system will activate adaptive detection enhancement logic. That is, in the next detection of that parameter, the applied perturbation amplitude will be automatically and temporarily reduced. Increase by a fixed multiple, such as increasing to To enhance the discernibility of the detection signal in a strong noise background; if the detection with enhanced perturbation amplitude still fails to obtain an effective gradient multiple times, the system will mark the parameter as a persistent strong interference state and temporarily remove it from the target set of dynamic closed-loop calibration. At the same time, an alarm message will be generated to the operation and maintenance management platform to indicate that there are abnormal environmental conditions in the area that require manual intervention. This procedure enables the system to avoid getting stuck in an ineffective detection loop when facing localized persistent severe operating conditions and provides a fault handling closed loop in collaboration with manual operation and maintenance.
[0055] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0056] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A dynamic closed-loop calibration method for a Beidou passive room based on reinforcement learning, characterized in that, The method comprises the following steps executed by the processor: Step a, before performing active probing on at least one adjustable system parameter in the Beidou passive room subsystem, obtaining a pre-reference performance index; Step b, performing active probing, which comprises: on the basis of the current parameter value of the adjustable system parameter, sequentially applying a positive disturbance and an equal negative disturbance to form a time-symmetric perturbation action; Step c, during the perturbation action, monitoring the positioning performance index of the system to calculate the response gradient; Step d, after the perturbation action ends and returns to the current parameter value, obtaining a post-reference performance index; Step e, performing gradient confidence discrimination, which comprises: comparing the pre-reference performance index with the post-reference performance index, and only when the difference between the pre-reference performance index and the post-reference performance index is less than a preset threshold, the response gradient is determined as a valid gradient; Step f, based on the valid gradient, a state vector representing the current local response characteristics of the system is constructed; Step g, taking the state vector as the input of the reinforcement learning agent, and outputting a calibration action for adjusting one or more adjustable system parameters by the reinforcement learning agent; Step h, performing the calibration action, and returning to step a; In step c, the response gradient is calculated according to the difference between the performance index monitored during the application of the positive disturbance and the performance index monitored during the application of the negative disturbance; In step e, when the difference between the pre-reference performance index and the post-reference performance index is not less than the preset threshold, the currently calculated response gradient is discarded, and steps a to e are immediately re-executed.
2. The BeiDou passive room dynamic closed-loop calibration method based on reinforcement learning according to claim 1, characterized in that, Further comprising: storing a series of state vectors generated by step f at different time points to form a state vector time sequence; performing trend analysis on the state vector time sequence to identify the long-term one-way drift trend of at least one component in the state vector; and generating a prognosis diagnosis information based on the long-term one-way drift trend, the prognosis diagnosis information being used to indicate that there is progressive performance degradation in the hardware unit corresponding to the at least one component.
3. The BeiDou passive room dynamic closed-loop calibration method based on reinforcement learning according to claim 1, characterized in that, Further comprising: evaluating a gradient signal-to-noise ratio of the probing procedure based on the one or more valid gradients determined in step e ; and the gradient signal-to-noise ratio dynamically adjusting the amplitudes of the positive and negative perturbations applied in the subsequent execution of step b; wherein the gradient signal-to-noise ratio is calculated following the rule wherein is the mean of the amplitudes of the valid gradients acquired in a pre-set time window, is the standard deviation of the amplitudes of the valid gradients.
4. The dynamic closed-loop calibration method based on reinforcement learning for Beidou passive room according to claim 1, wherein, The reward function of the reinforcement learning agent is calculated based on the product of the adjustment amount corresponding to the calibration action and the valid gradient, when the product is positive, a positive reward is given, and when the product is negative, a negative reward is given.
5. The dynamic closed-loop calibration method based on reinforcement learning for Beidou passive room according to claim 1, wherein, Between step f and step g, further comprising: calculating the information entropy of the state vector composed of one or more valid gradients; when the information entropy is less than a preset information entropy threshold, determining the calibration action as a preset gradient descent action; when the information entropy is not less than the preset information entropy threshold, the calibration action is output by the reinforcement learning agent.
6. The dynamic closed-loop calibration method based on reinforcement learning for Beidou passive room according to claim 1, wherein, The cycle execution of steps a to h is triggered when the positioning performance index of the system is lower than a preset performance threshold.
7. The dynamic closed-loop calibration method based on reinforcement learning for Beidou passive room according to claim 2, characterized in that, The trend analysis on the state vector time sequence further comprises: analyzing whether the amplitude of the gradient component of multiple parameters adjacent in geographical position or associated in function has a trend of increasing.
8. The BeiDou passive room dynamic closed-loop calibration method based on reinforcement learning according to claim 1, characterized in that, The state vector also includes valid gradient information determined at the historical time, to represent the dynamic change trend of the system response characteristic.
Citation Information
Patent Citations
Beidou passive positioning receiver
CN102012498A
Indoor distribution system monitoring method and device and electronic equipment
CN120151908A
Wireless communication battery management system with signal adaptive adjustment function
CN120414798A