Beidou passive indoor distribution dynamic closed-loop calibration method based on reinforcement learning

By adopting a reinforcement learning-based dynamic closed-loop calibration method for BeiDou passive indoor distribution systems, the problem of obtaining the causal relationship between control parameters and performance in complex controlled systems is solved, and the system can achieve stable adjustment and fault early warning in dynamic environments.

CN121165697AActive Publication Date: 2025-12-19HUNAN XIANGYINHE SENSOR TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202511705544.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-20
Publication Date
2025-12-19
Estimated Expiration
2045-11-20

AI Technical Summary

Technical Problem

Existing technologies struggle to obtain real-time local and instantaneous causal relationships between control parameters and system performance in complex, distributed controlled systems, leading to blind and uncertain adjustments and an inability to effectively cope with dynamic environmental disturbances and system state changes.

Method used

A BeiDou passive indoor distribution dynamic closed-loop calibration method based on reinforcement learning is adopted. By actively detecting and applying time-symmetric micro-perturbations, the system performance index changes are monitored, the response gradient is calculated, and the calibration action is output through gradient confidence discrimination and reinforcement learning agent, thus establishing an information field of local instantaneous causal relationships.

Benefits of technology

It enables direct optimization based on the current internal response characteristics of the system, avoids blind adjustment caused by multivariable coupling and environmental interference, ensures the stability and effectiveness of the control system in dynamic environments, and provides fault early warning capability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121165697A_ABST
    Figure CN121165697A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of general control or adjustment systems, and discloses a Beidou passive indoor distribution dynamic closed-loop calibration method based on reinforcement learning, which comprises the following steps: respectively acquiring front and rear reference performance indexes before and after applying a time-symmetric perturbation action to an adjustable parameter of a system; according to the method, the time symmetry of the active detection action is compared with the asymmetry of the external random interference, the time symmetry of the active detection action is compared with the asymmetry of the external random interference, and the accuracy of the calibration action is improved. A built-in information source reliability discrimination criterion is introduced for a control system, so that the system can actively distinguish performance changes caused by self detection behaviors and fluctuations caused by irrelevant external events, decision misguidance caused by gradient illusion is avoided, and the operation stability of a closed-loop regulation system in a dynamic open environment is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to a Beidou passive room dynamic closed-loop calibration method based on reinforcement learning and belongs to the technical field of general control or regulation systems. BACKGROUND

[0002] Currently, for dynamic closed-loop calibration of large indoor wireless signal distribution systems and the like, a commonly used method is to adjust multiple adjustable control parameters such as signal gains or time delays in the system online, so as to continuously optimize and maintain the stability of one or more system performance indicators. The basis of this method is to use feedback control to deal with the performance degradation caused by environmental changes or system drift. However, as such controlled systems are increasingly widely distributed in physical space, their structures are becoming increasingly complex, and the dynamicity and uncertainty of their operating environments are increasing, the inherent limitations of the aforementioned feedback regulation method that relies on overall performance indicators begin to appear. When a control system is faced with a controlled object whose internal state is complex, multivariables are strongly coupled, and an accurate model is difficult to establish, if only the performance indicators of the system output, which are averaged in space or time, such as the average error of regional positioning, are used for passive monitoring, the feedback information obtained will have fundamental information loss when guiding control decisions. For example, the performance improvement of a local area may be offset by the performance deterioration of another area caused by different factors in numerical value, resulting in no significant change in the average performance indicators of the system as a whole, and thus the control system cannot effectively perceive the local problem that is occurring and misses the adjustment opportunity.

[0003] Further, even if the control system monitors the decline in overall performance, the subsequent adjustment action lacks direct basis in direction and amplitude, because the adjustment of each control parameter in the system often has a highly nonlinear and mutually coupled effect on the final performance indicator, in the absence of clear quantification of the local, instantaneous causal relationship between each control parameter and system performance, any adjustment based on global optimization algorithm is essentially an inefficient search in multidimensional space, which not only makes it difficult to guarantee convergence efficiency, but also increases the risk of system oscillation, and even causes the deterioration of the performance of other related parts while trying to repair a problem. Simply improving the computing power of the optimization algorithm cannot fundamentally make up for the lack of decision basis caused by insufficient input information dimensions; Furthermore, the existing technology also has the inherent defect of being unable to effectively respond to dynamic environmental disturbances and system state changes in the control method level, for example, the Chinese patent application with the publication number CN102012498A discloses a Beidou passive positioning receiver, which realizes passive positioning, but its technical core lies in positioning solution by receiving signals from three satellites and combining with external elevation data, the positioning process is essentially an open-loop calculation based on a static model, this method lacks effective online perception and closed-loop calibration mechanism for dynamic performance degradation caused by environmental changes or system component drift, and when there is instantaneous strong interference in the external environment, the positioning accuracy will be significantly affected, which cannot guarantee the service stability in complex dynamic environment, which further highlights the importance of establishing a method that can distinguish interference and dynamically adjust the system in real time.

[0004] The existing technology mainly has the following limitations: 1. The overall average performance indicator relied on by the control system can mask the local, asynchronous performance changes in the controlled object, resulting in lagging or missing adjustment response; 2. Under the condition that the model of the controlled object is unknown and the multivariables are strongly coupled, the control system lacks clear gradient information to guide parameter adjustment, making the adjustment action blind and uncertain; 3. Any attempt to overcome information deficiency through optimization algorithm cannot change the black box property of the control process, making it difficult to achieve effective balance between adjustment efficiency and system stability. Therefore, how to establish an online mechanism in a complex, distributed, model-unknown controlled system to obtain the local instantaneous causal relationship between the adjustment of each control parameter and the overall performance of the system in real time, so as to change the global search optimization into directional gradient-based optimization adjustment, has become the technical problem to be solved by the present application. SUMMARY

[0005] The application provides a Beidou passive room sub-division dynamic closed-loop calibration method based on reinforcement learning, which mainly aims to solve how to obtain a local and instantaneous causal relationship between control parameters and system performance in a complex and model-unknown controlled system online, so as to overcome the regulation blindness and uncertainty caused by information loss of traditional feedback control.

[0006] To achieve the above-mentioned purpose, the application provides a Beidou passive room sub-division dynamic closed-loop calibration method based on reinforcement learning, which comprises the following steps executed by a processor: Step a: before performing active detection on at least one adjustable system parameter in the Beidou passive room sub-division system, obtaining a pre-positioned reference performance index; Step b: performing active detection, which comprises: on the basis of the current parameter value of the adjustable system parameter, sequentially applying a positive disturbance and an equal negative disturbance to form a time-symmetrical micro-disturbance action; Step c: during the micro-disturbance action, monitoring the positioning performance index of the system to calculate a response gradient; Step d: after the micro-disturbance action ends and returns to the current parameter value, obtaining a post-positioned reference performance index; Step e: performing gradient confidence discrimination, which comprises: comparing the pre-positioned reference performance index with the post-positioned reference performance index, and only when the difference between the pre-positioned reference performance index and the post-positioned reference performance index is less than a preset threshold, the response gradient is determined as a valid gradient; Step f: based on the valid gradient, forming a state vector representing the current local response characteristics of the system; Step g: taking the state vector as the input of a reinforcement learning agent, and outputting a calibration action for adjusting one or more adjustable system parameters from the reinforcement learning agent; Step h: performing the calibration action, and returning to step a cyclically.

[0007] Preferably, it further comprises: storing a series of state vectors generated by step f at different time points to form a state vector time sequence; performing trend analysis on the state vector time sequence to identify a long-term one-way drift trend of at least one component in the state vector; and based on the long-term one-way drift trend, generating a prognosis diagnosis information, which is used to indicate that there is a gradual performance degradation in a hardware unit corresponding to the at least one component.

[0008] Preferably, it further comprises: based on one or more valid gradients determined in step e, evaluating the gradient signal-to-noise ratio of the detection process ; and according to the gradient signal-to-noise ratio , dynamically adjusting the amplitudes of the positive and negative disturbances applied when subsequently performing step b; wherein the gradient signal-to-noise ratio the calculation follows rules, wherein, is the mean value of the amplitudes of the plurality of effective gradients obtained within a preset time window, is the standard deviation of the amplitudes of the plurality of effective gradients.

[0009] Preferably, the reward function of the reinforcement learning agent is calculated based on the product of the adjustment amount corresponding to the calibration action and the effective gradient, and when the product is positive, a positive reward is given, and when the product is negative, a negative reward is given.

[0010] Preferably, between step f and step g, further comprising: calculating the information entropy of the state vector composed of one or more effective gradients; when the information entropy is less than a preset information entropy threshold, determining the calibration action as a preset gradient descent action; and when the information entropy is not less than the preset information entropy threshold, outputting the calibration action by the reinforcement learning agent.

[0011] Preferably, the loop of steps a to h is triggered when it is monitored that the positioning performance indicator of the system is lower than a preset performance threshold.

[0012] Preferably, the trend analysis on the state vector time sequence further comprises: analyzing whether there is a trend of increasing in the amplitude of the gradient component of the plurality of parameters that are adjacent in geographical position or are functionally associated, to identify regional instability signs.

[0013] Preferably, the response gradient in step c is calculated according to the difference between the performance indicator monitored during the application of the positive disturbance and the performance indicator monitored during the application of the negative disturbance.

[0014] Preferably, the state vector further comprises effective gradient information determined at a historical time, to represent the dynamic change trend of the system response characteristics.

[0015] Preferably, in step e, when the difference between the pre-position reference performance indicator and the post-position reference performance indicator is not less than a preset threshold, the currently calculated response gradient is discarded, and steps a to e are immediately re-executed.

[0016] Compared with the prior art, the beneficial effects of the present application are: 1、By applying a symmetric perturbation action to a single adjustable parameter in the system, and synchronously monitoring the change of system performance index during the action, the response gradient representing the instantaneous correlation between the parameter adjustment and performance change is calculated; then, one or more such response gradients are constructed into a state vector, which is input to the reinforcement learning agent and outputs the calibration action; in this way, the information base on which the control system makes decisions is changed from the traditional macroscopic average performance index representing the overall past running state of the system, to an information field that can reflect the local instantaneous causal relationship between each control input and system output in real time, so that the control adjustment process is no longer a passive response based on historical statistics, but a direct optimization based on the current system internal response characteristics, avoiding the problems of adjustment blindness and uncertainty caused by multi-variable coupling or performance averaging effect.

[0017] 2、Before applying the symmetric perturbation action, the pre-reference performance index is obtained, and after the parameter is restored at the end of the action, the post-reference performance index is obtained; only when the two reference indexes are consistent within the preset range, the response gradient calculated during the action is determined as the valid gradient for state representation; this mechanism compares the time symmetry of the active detection action with the asymmetry of external random interference, and introduces a built-in criterion for distinguishing the reliability of information sources for the control system; in this way, the system has the ability to distinguish the performance changes caused by its own detection behavior from the performance fluctuations caused by irrelevant external events, avoiding the misleading of the decision-making process caused by the gradient illusion due to environmental instantaneous changes, ensuring the stability of the entire closed-loop adjustment system running in an open dynamic environment.

[0018] 3、The method of the application stores the effective state vectors generated at different time points to form a time sequence, and performs trend analysis on the sequence to identify the long-term change trend of a specific component, and further generates prognosis diagnosis information representing the long-term health status of the system; this reuses the gradient information originally used for real-time tactical calibration as a long-term observation reflecting the state evolution of the controlled physical system itself; when a hardware unit in the system gradually ages or performance degrades, the control system's efforts to maintain apparent performance will leave a record in the time sequence in the form of a corresponding gradient component that continuously drifts in one direction, enabling the method to complete dynamic calibration tasks while also providing early warning of potential fault risks in the physical system, achieving an expansion of capabilities from function maintenance to health status diagnosis; at the same time, based on one or more response gradients, the signal-to-noise ratio of the detection process is evaluated, and the amplitude of the perturbation action applied during the next active detection is dynamically adjusted based on the evaluation results; this establishes a regulation closed loop for the detection behavior itself outside the control closed loop, enabling the system's observation means to be adaptive; when the system is in a stable state, the perturbation amplitude is automatically reduced to reduce the impact on normal service, while when the system performance is deteriorating and the response is sluggish, the amplitude is automatically increased to obtain gradient information with a higher signal-to-noise ratio, enabling a dynamic balance between the effectiveness of the detection behavior and the impact on the system, improving the method's versatility under different working conditions. BRIEF DESCRIPTION OF DRAWINGS

[0019] Fig. 1 A flowchart of the self-verification dynamic closed-loop calibration method of the application; Fig. 2 A result graph of offline calibration of the active detection key parameters of the application; Fig. 3 A deployment architecture diagram of the system of the application in a multi-region physical environment. DETAILED DESCRIPTION

[0020] To make the objectives, technical solutions, and advantages of the application clearer, the technical solutions of the application will be described in detail below with reference to the drawings. Obviously, the described embodiments are part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of the application.

[0021] The application discloses a Beidou passive indoor sub-division dynamic closed-loop calibration method based on reinforcement learning, which is realized as a closed-loop adjustment process iterated by a processor, and mainly comprises steps of active detection and response gradient acquisition, gradient confidence discrimination, state representation, and reinforcement learning decision and execution, so that the local and instantaneous causal relationship between the control input and the system state is identified online in a multivariate coupled controlled system, and the global search type optimization is changed into the adjustment mode based on the gradient guidance; in the operation process of a large indoor wireless signal distribution system, the environment change or system component drift often leads to the deterioration of the local positioning performance, and the existing adjustment mode depends on the monitoring of the overall average performance index of the system, so that the local and asynchronous performance change cannot be effectively perceived and responded; in order to solve the problem, the method adopts the active detection step to acquire the local response characteristics of the system, the step is performed on a adjustable system parameter in the Beidou passive indoor sub-division system, for example, the gain of a specific signal channel, a time-symmetric micro-perturbation action is performed, specifically, on the basis of the current parameter value of the adjustable system parameter, a positive perturbation and an equal negative perturbation are sequentially applied, and the change of the system positioning performance index is monitored during the application of the perturbation, and then the response gradient of the adjustable system parameter is calculated according to the difference between the performance indexes monitored during the application of the positive perturbation and the negative perturbation; for example, if the current value of a gain parameter is 20 dB, the perturbation amplitude is 1 dB, the active detection action is to sequentially set the parameter value to 21 dB and maintain a time window T, then set the parameter value to 19 dB and maintain the same time window T, and finally restore the parameter value to 20 dB, if the monitored average positioning errors of the system in the two time windows T are 0.5 m and 0.3 m respectively, the response gradient of the parameter can be calculated according to the formula 0.2 m / dB, and the calculation of the response gradient provides a quantitative decision basis with causal relationship for subsequent parameter adjustment.

[0022] In a dynamic running environment, external random events irrelevant to the control action may interfere with the performance index, so that the calculated response gradient is distorted, that is, the gradient illusion is generated, so that the adjustment action is misled; therefore, the method establishes a gradient confidence discrimination procedure, the procedure compares the time symmetry of the active detection action with the asymmetry of the external random interference, introduces the discrimination criterion of the reliability of the information source for the control system, and the specific implementation is that, before the active detection is performed, a pre-set reference performance index is acquired. ​​​​​​​​​​​And after the perturbation action ends and the parameters are restored to their current values, a post-baseline performance index is obtained. Subsequently, the preceding and following benchmark performance metrics are compared, and only when the absolute value of the difference between the two is... Less than a preset threshold Only when the time is right will the response gradient calculated during that period be determined as a valid gradient; a preset threshold is used. The value can be determined through a standardized offline calibration process. This involves continuously monitoring the natural fluctuations of system performance indicators over a period of time (e.g., 30 minutes) during the initial system deployment phase, without applying any perturbations, and calculating their standard deviation. and set a preset threshold Set it to a multiple of that standard deviation, for example ,when If a significant external disturbance occurs during the detection process, the currently calculated response gradient is discarded, and the system re-executes the detection steps. This discrimination procedure enables the system to proactively distinguish between performance changes caused by its own detection behavior and fluctuations caused by external events, ensuring the reliability of the decision-making basis.

[0023] After acquiring one or more effective gradients, in order to achieve a comprehensive characterization of the system state and make adjustment decisions, the method of this invention constructs one or more effective gradients into a state vector characterizing the current local response characteristics of the system. The state vector This is used as input to a reinforcement learning agent; the reinforcement learning agent, such as an actor-critic model based on a deep neural network, is trained through offline simulation or online learning to master a mapping policy from the state vector space to the calibrated action space. The agent receives the state vector... Then, it outputs a calibration action for adjusting one or more adjustable system parameters. ,in This refers to the specific adjustment amount for the i-th parameter; during the training of the agent, its reward function is set as the adjustment amount corresponding to the calibration action. With effective gradient product When the value is positive, a positive reward is given; otherwise, a negative reward is given. This reward mechanism prompts the agent to learn how to make effective optimization adjustments based on gradient information. Finally, the system executes the calibration action output by the agent and returns to the active exploration step in a loop to continuously adapt to the dynamic changes in the system state.

[0024] Embodiment 1: This embodiment is an application example of the technical solution in a specific industrial control scenario; in a large automated logistics warehouse composed of multiple sub-zones, the Beidou passive room system provides navigation signals for a cluster of autonomous mobile robots AGVs, wherein the signal amplifier in sub-zone A has a slow gain thermal drift due to continuous operation, resulting in a weak and continuous degradation of signal coverage quality in this area; at the same time, an AGV loaded with a metal container is passing through sub-zone B adjacent to it at a high speed, causing a transient asymmetric disturbance to the signal environment of sub-zone B; under this working condition, if a control system only relies on the overall average positioning error of the warehouse as feedback, the overall performance indicator monitored by it does not change significantly, because the performance degradation in sub-zone A is numerically offset by the random fluctuations in sub-zone B or other areas, resulting in the control system not responding to the progressive problem occurring in sub-zone A; the method of the present application is triggered, first performing an active detection step on the adjustable gain parameter in sub-zone A where performance degradation occurs, and the pre-reference performance indicator and the post-reference performance indicator obtained by the controller before and after the action of applying a time-symmetric perturbation are substantially the same, and the gradient confidence discrimination step determines a weak non-zero response gradient calculated during the detection period as an effective gradient; when the method performs an active detection on the gain parameter in sub-zone B as planned, the passing of the AGV causes the positioning performance indicator during the micro-perturbation action to jump, so that the post-reference performance indicator after the detection ends deviates significantly from the pre-reference performance indicator, and the gradient confidence discrimination step determines that this detection is contaminated by external events, and discards the response gradient calculated during the detection.

[0025] The cooperative operation of the aforementioned active detection step and the gradient confidence discrimination step makes the information finally used to construct the state vector and input to the reinforcement learning agent contain only data that has passed the reliability verification, and the state vector represents that there is a certain optimization direction for the gain parameter in sub-zone A, while the state of the parameter in sub-zone B is not included in the decision basis due to unreliable information; the reinforcement learning agent outputs a calibration action that only slightly improves the gain parameter in sub-zone A according to this state vector, and does not perform any operation on sub-zone B; after the calibration action is executed, the signal quality in sub-zone A is compensated and the performance is restored to the normal level, at the same time, the system avoids adjusting the parameter in sub-zone B which may cause signal oscillation due to a false adjustment; this closed-loop adjustment method does not rely on establishing a complete system model when facing distributed, multivariable and strongly disturbed controlled objects, but continuously confirms the local effectiveness of the control action through an online self-verification information acquisition and decision-making process.

[0026] Example 2: To objectively verify the improvement effect of the method of the present invention on the stability of the closed-loop control system under the condition of transient external disturbance, a numerical simulation test platform was built in this example. The platform simulates a Beidou passive indoor distribution system with a single adjustable system gain parameter. Its function is to continuously compensate for a preset, slow system parameter drift through closed-loop adjustment based on a quantifiable positioning performance index. The platform can inject an asymmetric transient disturbance simulating a sudden change in the external environment into the positioning performance index at a specified time point. The test is used to compare and analyze the operating performance of two control methods with and without gradient confidence discrimination steps. The test sets up a control group and the sample group of the present invention. The control method of the control group includes active detection, response gradient calculation and reinforcement learning decision, but all calculated response gradients are considered valid. The sample group of the present invention adopts the complete technical method disclosed, including the gradient confidence discrimination step. The test is carried out for 10 adjustment cycles. The true value of the system gain parameter is set from the initial value. Starting at dB, it decreases linearly with each cycle. dB is used to simulate component performance degradation. During the active detection window in the 5th adjustment cycle, a transient disturbance is superimposed on the positioning performance index. The fluctuation amplitude of the performance index caused by this disturbance is greater than the normal fluctuation caused by parameter perturbation. For the sample group of this invention, the preset threshold for gradient confidence discrimination is... The standard deviation of the system's performance index under interference-free conditions was set to twice the standard deviation of the performance index. During the fifth adjustment period of the experiment, the control group calculated a response gradient that was abnormally induced by interference during the detection period and used it for decision-making, causing its reinforcement learning agent to output a large error gain compensation action. The sample group of this invention conducted detection at the same time, and its gradient confidence discrimination step detected that the difference between the pre-benchmark performance index and the post-benchmark performance index exceeded a preset threshold. Therefore, the currently calculated response gradient is deemed invalid and discarded, and no adjustment action is performed; the operating status and performance data records of the two test groups within 10 adjustment cycles are shown in Table 1.

[0027] Table 1: Performance comparison of the two control methods under drift and transient disturbances in the simulated system.

[0028]

[0029] From the data of Table 1, it can be seen that the parameter values of the control group deviate from the true values in the 5th period due to the error adjustment, resulting in the performance error continuously at a high level in the subsequent periods, showing control instability; the parameter values of the sample group of the application closely follow the changes of the true values of the parameters throughout the test process, and the performance error always maintains a stable low level, which is not significantly affected by transient disturbances; the test results show that the introduction of the gradient confidence discrimination step enables the closed-loop control system to effectively identify and reject feedback information contaminated by external transient events, avoiding the decision misleading caused by the wrong gradient.

[0030] To further illustrate the mechanism, how the lack of the gradient confidence discrimination step described in the application will directly lead to the failure of the conventional online gradient detection method in a dynamic environment, the following Comparative Example 1 is set.

[0031] Comparative Example 1: This comparative example aims to simulate a conventional technical path that a person skilled in the art may adopt when trying to solve the problems in the background art, i.e., directly using the gradient for feedback adjustment on the basis of introducing active detection to obtain the response gradient, but it does not contain the key gradient confidence discrimination step of the application. The test conditions of this comparative example are completely consistent with Example 2: the same numerical simulation test platform is used to simulate a Beidou passive indoor system with a single adjustable system gain parameter; the true value of the system gain parameter is also set to start from the initial value of 10.0 dB and linearly decrease by 0.1 dB in each adjustment period, to simulate the same component performance degradation; and in the active detection window period of the 5th adjustment period, a transient disturbance simulating an external environmental mutation is injected into the positioning performance index, which is exactly the same as in Example 2. The only difference between this comparative example and Example 2 is that the control method used in this comparative example, after performing active detection and calculating the response gradient, does not perform any confidence discrimination, but directly regards the response gradient as an effective gradient, and uses it as the input for the subsequent reinforcement learning agent to make decisions.

[0032] In the first to fourth adjustment periods, since there is no external disturbance, the method of the comparative example can calculate a substantially accurate response gradient, and the parameter adjustment result is close to that of the embodiment 2 of the present application, and can substantially follow the slow drift of the parameter true value; in the fifth adjustment period, when the external transient disturbance is injected, the control method of the comparative example incorrectly attributes the performance index fluctuation caused by the disturbance to the self-disturbance action during the active detection, thereby calculating an abnormal and directionally incorrect response gradient. Since there is no verification link, the illusory gradient is directly input to the reinforcement learning agent, resulting in a large and incorrect gain compensation action of the controller, which incorrectly adjusts the system parameter value from a level close to 9.6 dB to 7.25 dB; in the sixth to tenth adjustment periods, although the external disturbance has disappeared, the system parameter has deviated from the true working point due to the incorrect adjustment in the fifth period, and the control system attempts to adjust back according to the gradient calculated under the disturbance-free condition, but it takes multiple adjustment periods to slowly and gradually correct the previous large deviation; during this period, the positioning performance error of the system is continuously at an unacceptably high level, showing significant control instability and adjustment oscillation. The detailed performance data record of each period is recorded in Table 2.

[0033] Table 2: Performance record table of the conventional control method in the comparative example 1 under the system drift and transient disturbance.

[0034]

[0035] The test results show that, under the condition of lacking the gradient confidence discrimination mechanism, the conventional technical path of using only active detection to obtain the gradient cannot resist the transient disturbance of the external environment in its design principle, and the system will make disastrous decision errors due to the gradient illusion, resulting in the loss of stability of the closed-loop control system in the dynamic open environment.

[0036] Embodiment 3: The embodiment combines Figs. 1 to 3 the dynamic closed-loop calibration method for the Beidou passive room based on reinforcement learning, as follows: Fig. 1As shown, the flow is triggered by system performance monitoring, when the positioning performance indicators are monitored to be lower than the preset threshold, the calibration cycle is started, first, a pre-reference performance indicator is obtained to record the system stable state before detection, then active detection is performed, that is, a time-symmetric perturbation is applied to a certain adjustable parameter, and the response gradient quantifying the relationship between parameter adjustment and performance change is calculated during this period, after the detection is completed, a post-reference performance indicator is obtained to record the system stable state after detection, then the core gradient confidence discrimination link is entered, whether the difference between the pre-reference performance indicator and the post-reference performance indicator is less than the preset threshold is judged to confirm whether the detection process is disturbed by external interference, if the difference is not less than the threshold, it is determined that the gradient is invalid, discarded and re-executed, if the difference is less than the threshold, it is confirmed that the gradient is a valid gradient, and one or more valid gradients are further constructed into a state vector representing the current local response characteristics of the system, the state vector is input to the reinforcement learning agent on one hand, and the calibration action is output by the decision of the agent, and the action is executed by the system to adjust one or more system parameters, on the other hand, the valid state vector that passes the confidence recognition can also be stored to form a state vector time sequence, through trend analysis and diagnosis of the sequence, the performance degradation of the hardware can be identified, and the prognosis diagnosis information can be generated, so as to output potential fault risk warning.

[0037] As shown in the figure, Fig. 2 The horizontal coordinate of the figure is the perturbation duration T, the unit is millisecond ms, and the vertical coordinate is the effective gradient recognition rate, the unit is percentage %, two curves are shown in the figure, which respectively represent the trend of the effective gradient recognition rate changing with the perturbation duration T under the condition of using the selected perturbation amplitude solid line and the comparison perturbation amplitude dot-dash line, it can be seen from the figure that under the two perturbation amplitudes, the recognition rate increases with the increase of time T, and the recognition rate under the selected perturbation amplitude is better than that under the comparison perturbation amplitude at most time points, a vertical dotted line is marked in the figure to indicate the finally selected initial working point, which corresponds to a perturbation duration of 200 ms, the selection of this working point is to shorten the detection time as much as possible to reduce the influence on the normal service of the system under the premise of ensuring high recognition rate. Fig. 3As shown, the whole system is logically divided into the upper cloud control and operation center and the lower physical deployment environment. The cloud center is internally deployed with a central controller, which integrates the reinforcement learning agent and the gradient confidence discrimination module, and is connected with an effective gradient time series database for storing historical data. The data of the database are used for long-term trend analysis, and the analysis results are sent to the operation management platform in the form of prognosis diagnosis or alarm information to present to the operation personnel. The physical deployment environment is divided into several areas, such as A area, B area and C area with persistent strong interference. The adjustable hardware units such as signal amplifiers and devices such as autonomous mobile robots or mobile terminals are distributed in each area. The central controller implements control by sending detection and calibration actions to the adjustable hardware units in the physical environment, and the autonomous mobile robots in the deployment environment feedback the performance indicators to the central controller in real time, thereby forming a complete dynamic closed-loop control and diagnosis system.

[0038] Embodiment 4: This embodiment is used to illustrate the calibration method of the key operating parameters and the switching procedure of the decision logic in the foregoing technical solutions, to solve the technical problem of parameter setting in the specific deployment of the system. Before the deployment of the system, the perturbation amplitude and the perturbation duration T in the active detection mechanism of the system need to be determined. The selection of the initial working point involves a technical trade-off between ensuring that the detection signal can be effectively detected and reducing the disturbance of the detection action on the normal service of the system. To determine the above parameters, this embodiment adopts a standardized offline calibration procedure, which is performed in a controlled test environment. First, a known weak linear drift rate, i.e. dB / min, is injected into a selected adjustable system parameter, gain . Then, different combinations of perturbation amplitude (in the range of dB to dB, with a step of dB) and perturbation duration T (in the range of ms to ms, with a step of ms) are used to repeatedly perform the active detection step, and the calculated response gradient under each combination is recorded. By comparing the calculated gradient with the theoretical gradient corresponding to the known drift, a minimum and the shortest T combination that can stably reflect the true response of the system are determined, such as dB and T ms, as the initial parameters for online operation of the system. The procedure provides a reproducible determination method for the initial values of the parameters.

[0039] To adapt to changes in signal-to-noise ratio during system online operation, the method of this invention also includes a perturbation amplitude... The dynamic adaptive adjustment closed loop operates on a fixed time scale, such as an evaluation window of 100 adjustment cycles. The steps are as follows: First, based on multiple effective gradients determined by the gradient confidence discrimination step within this period, the gradient signal-to-noise ratio (GSNR) of the detection process is evaluated. Its calculation follows... The rule, where g is the mean magnitude of multiple valid gradients acquired within a preset time window, The magnitude standard deviations of multiple effective gradients are used; subsequently, the gradient signal-to-noise ratio (GSNR) is calculated based on its current value and a preset target value. ( The comparison results (set to 5.0) dynamically adjust the perturbation amplitude applied during the next active detection. If currently If it falls below the target value, the increase will be proportional. Conversely, the value decreases, and this adjustment loop is used to dynamically adjust the parameters of the detection behavior. To coordinate the computational overhead of complex decision-making by the reinforcement learning agent with the need for rapid response under simple conditions, the present invention introduces a decision-making pattern arbitration mechanism based on state vector information entropy before the reinforcement learning decision-making step. Specifically, the state vector is composed of one or more effective gradients. Then, the information entropy H of the vector is calculated first. If the gradient field has a sharp shape, that is, a few gradient components have large absolute values ​​while others are small, the information entropy H is low, indicating that the main factors affecting the system performance degradation are clear. If H is less than a preset information entropy threshold, the system will bypass the reinforcement learning agent and directly execute a preset gradient descent action, that is, select the parameter with the largest gradient and adjust it by a fixed step size along the gradient direction. Only when the gradient field has a flat shape, that is, all gradient components have similar absolute values ​​and random directions, and the information entropy H is high and not less than the information entropy threshold, will the reinforcement learning agent be activated to make a decision. This arbitration mechanism switches between two decision modes based on the information entropy of the state vector.

[0040] Example 5: This example aims to illustrate the function of using the effective gradient information generated by the method during long-term operation to perform trend analysis in order to achieve the prognostic function of system health status; In a Beidou passive indoor distribution system that has been deployed and continuously operated for more than six months, the controller is configured to store the effective gradients corresponding to all adjustable system parameters determined by the gradient confidence screening step at different time points, forming an effective gradient time series database with timestamps. A background diagnostic module performs trend analysis on the data in this database with a 24-hour cycle.

[0041] At a certain moment, the analysis algorithm of the diagnosis module detects that the 30-day moving average of the effective gradient of the gain adjustment parameter corresponding to the amplifier unit marked as A-07 in the system presents a continuous one-way positive drift in the past three consecutive analysis periods, and the slope exceeds the preset statistical threshold, while the algorithm analyzes the time series of the effective gradient of the parameters corresponding to the units A-06 and A-08 adjacent to A-07 in the geographical position and the unit B-02 associated with A-07 in the signal link, and no synchronous trend growth is found; based on this, the system determines that there is a high possibility of progressive performance degradation of the A-07 unit itself rather than regional instability, and generates a prognosis diagnosis information that the control compensation amount of the hardware unit A-07 presents a continuous one-way growth, prompting that the unit may have progressive performance degradation, and preventive maintenance is suggested, which is pushed to the operation and maintenance platform; the operation and maintenance personnel confirm the gain decrease of the amplifier unit caused by component aging and replace it based on this.

[0042] Embodiment 6: This embodiment aims to explain the parameter scanning strategy adopted by the method in a large-scale, high-dynamic interference environment with a large number of adjustable parameters, in order to balance the detection efficiency and system response speed, and the fault tolerance adjustment mechanism in the case of persistent strong interference environment; in an indoor distributed system of a traffic hub with hundreds of adjustable gain and delay parameters, if all parameters are completely sequentially polled for detection, the entire scanning period is too long to respond to the rapid changes in the environment in time; to solve this problem, the method of the present application adopts a scanning strategy of grouped polling and dynamic focusing. First, all adjustable system parameters are divided into multiple parameter groups according to their physical positions or functional correlations. In the normal running state, the system does not detect all parameters in each adjustment period, but uses polling to sequentially perform active detection and gradient calculation on different parameter groups to maintain the state monitoring of the entire system at a lower time cost. When the system finds that the effective gradient amplitude of one or more parameters in a parameter group exceeds the preset attention threshold in the detection, the control system will automatically enter the dynamic focusing mode. In this mode, the system will temporarily interrupt the polling of other parameter groups, and concentrate the detection resources of the subsequent several adjustment periods on the abnormal parameter group and its adjacent parameter groups for higher frequency continuous detection and calibration until the overall amplitude of the state vector of the region returns to the stable level, and then exit the focusing mode and restore the normal global polling scanning.

[0043] Further, when a certain area in the system, such as a security checkpoint, has long-term and strong RF interference due to continuous operation of the equipment, the gradient confidence discrimination step for the parameters of this area may fail continuously for multiple times, resulting in that the parameter state of this area cannot be effectively updated for a long time; to deal with this boundary condition, the method of the present application also includes a set of fault-tolerant upgrade procedures, when the system monitors that the gradient confidence discrimination of a certain parameter fails continuously for more than a preset retry upper limit, such as 5 times, the system will start the adaptive detection enhancement logic, that is, in the next detection for this parameter, the applied perturbation amplitude is increased by a fixed multiple, such as to to enhance the distinguishability of the detection signal in a strong noise background; if the detection with enhanced perturbation amplitude still fails to obtain an effective gradient for multiple times, the system will mark this parameter as a persistent strong interference state, and temporarily remove it from the target set of dynamic closed-loop calibration, while generating an alarm information to the operation and maintenance management platform, prompting that there is an abnormal environmental condition in this area that needs manual intervention for troubleshooting, this procedure enables the system to avoid falling into an invalid detection cycle when facing a local persistent harsh working condition, and provides a fault handling closed loop that cooperates with manual operation and maintenance.

[0044] It is obvious to those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application.

[0045] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not limiting, although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or replaced equivalently without departing from the spirit and scope of the technical solutions of the present application.

Claims

1. A dynamic closed-loop calibration method for a Beidou passive room based on reinforcement learning, characterized in that, The method comprises the following steps executed by the processor: Step a, before performing active probing on at least one adjustable system parameter in the Beidou passive room subsystem, obtaining a pre-reference performance index; Step b, performing active probing, which comprises: on the basis of the current parameter value of the adjustable system parameter, sequentially applying a positive disturbance and an equal negative disturbance to form a time-symmetric micro-disturbance action; Step c, during the micro-disturbance action, monitoring the positioning performance index of the system to calculate the response gradient; Step d, after the micro-disturbance action ends and returns to the current parameter value, obtaining a post-reference performance index; Step e, performing gradient confidence discrimination, which comprises: comparing the pre-reference performance index with the post-reference performance index, and only when the difference between the pre-reference performance index and the post-reference performance index is less than a preset threshold, the response gradient is determined as a valid gradient; Step f, based on the valid gradient, a state vector representing the current local response characteristics of the system is constructed; Step g, taking the state vector as the input of the reinforcement learning agent, and outputting a calibration action for adjusting one or more adjustable system parameters from the reinforcement learning agent; Step h, performing the calibration action, and returning to step a.

2. The BeiDou passive room dynamic closed-loop calibration method based on reinforcement learning according to claim 1, characterized in that, Further comprising: storing a series of state vectors generated by step f at different time points to form a state vector time sequence; performing trend analysis on the state vector time sequence to identify the long-term one-way drift trend of at least one component in the state vector; and based on the long-term one-way drift trend, generating a prognosis diagnosis information, the prognosis diagnosis information is used to indicate that there is progressive performance degradation in the hardware unit corresponding to the at least one component.

3. The BeiDou passive room dynamic closed-loop calibration method based on reinforcement learning according to claim 1, characterized in that, Further comprising: evaluating a gradient signal-to-noise ratio of the probing procedure based on the one or more valid gradients determined in step e ; And based on the gradient signal-to-noise ratio The magnitudes of the positive and negative perturbations applied during subsequent execution step b are dynamically adjusted; where the gradient signal-to-noise ratio... The calculation follows The rules, among which, This is the average magnitude of multiple valid gradients acquired within a preset time window. The standard deviation of the magnitude of multiple effective gradients.

4. The dynamic closed-loop calibration method based on reinforcement learning for Beidou passive room according to claim 1, wherein, The reward function of the reinforcement learning agent is calculated based on the product of the adjustment amount corresponding to the calibration action and the valid gradient, when the product is positive, a positive reward is given, and when the product is negative, a negative reward is given.

5. The dynamic closed-loop calibration method based on reinforcement learning for Beidou passive room according to claim 1, wherein, Between step f and step g, further comprising: calculating the information entropy of the state vector composed of one or more valid gradients; when the information entropy is less than a preset information entropy threshold, the calibration action is determined as a preset gradient descent action; when the information entropy is not less than the preset information entropy threshold, the calibration action is output by the reinforcement learning agent.

6. The dynamic closed-loop calibration method based on reinforcement learning for Beidou passive room according to claim 1, wherein, The cycle execution of steps a to h is triggered when the positioning performance index of the system is lower than a preset performance threshold.

7. The dynamic closed-loop calibration method based on reinforcement learning for Beidou passive room according to claim 2, characterized in that, The trend analysis on the state vector time sequence further comprises: analyzing whether the amplitude of the gradient component of multiple parameters that are adjacent in geographical position or functionally related has a trend of increasing.

8. The dynamic closed-loop calibration method based on reinforcement learning for Beidou passive room according to claim 1, wherein, The calculation of the response gradient in step c is based on the difference between the performance index monitored during the application of the positive disturbance and the performance index monitored during the application of the negative disturbance.

9. The dynamic closed-loop calibration method based on reinforcement learning for Beidou passive room according to claim 1, wherein, The state vector further comprises valid gradient information determined at historical time points, which is used to represent the dynamic change trend of the system response characteristics.

10. The dynamic closed-loop calibration method based on reinforcement learning for Beidou passive room according to claim 1, wherein, In step e, when the difference between the pre-reference performance index and the post-reference performance index is not less than the preset threshold, the currently calculated response gradient is discarded, and steps a to e are immediately re-executed.

Citation Information

Patent Citations

  • Beidou passive positioning receiver

    CN102012498A

  • Power regulation method, device and equipment of indoor distribution system and medium

    CN119110381A

  • Doppler very high frequency omnidirectional beacon transmission channel closed-loop calibration system and method

    CN119254349A

  • Indoor distribution system monitoring method and device and electronic equipment

    CN120151908A

  • Wireless communication battery management system with signal adaptive adjustment function

    CN120414798A