A voltage sampling circuit management system for high voltage buck topology output

By introducing a risk quantification unit, an intelligent adjustment unit, and a hierarchical execution unit into the high-voltage BUCK topology system, the problem of reduced voltage sampling accuracy caused by component thermal drift and grid harmonic interference was solved, and the long-term stability and reliability management of the system was achieved.

CN121689754BActive Publication Date: 2026-04-17米博电源(厦门)有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
米博电源(厦门)有限公司
Filing Date
2026-02-09
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

In existing high-voltage BUCK topology systems, the voltage sampling circuit suffers from long-term accuracy degradation due to component thermal drift, response delay, and grid harmonic interference, which may lead to system instability or catastrophic failures.

Method used

By introducing risk quantification units, intelligent adjustment units, and hierarchical execution units, and through trust entropy analysis, deep reinforcement learning, and hierarchical control, we can achieve forward-looking adjustment and risk management of the sampling system, ensuring that the system adopts the most appropriate control strategies under different risk levels.

Benefits of technology

It significantly improves the system's early warning capability and long-term reliability, and can maintain the long-term accuracy and reliability of voltage sampling under complex operating conditions, avoiding catastrophic consequences caused by slow drift.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121689754B_ABST
    Figure CN121689754B_ABST
Patent Text Reader

Abstract

This invention relates to the field of power electronic converter control technology, specifically a voltage sampling circuit management system for high-voltage BUCK topology output, comprising a risk quantification unit, an intelligent adjustment unit, and a hierarchical execution unit. The risk quantification unit performs risk quantification analysis on real-time acquired system physical quantities to generate a trust entropy. The intelligent adjustment unit constructs a state space based on the trust entropy, normalized voltage ripple, and normalized dynamic response recovery time, and performs deep reinforcement learning analysis based on the state space to output adjustment action commands. The hierarchical execution unit compares the trust entropy with a preset threshold to generate an operating status signal, and corrects or suspends the adjustment action commands in response to the operating status signal to execute the final control strategy. This invention realizes the transformation from passive response after a fault occurs to proactive prediction beforehand, ensuring the long-term accuracy of voltage sampling under complex operating conditions such as high voltage and strong interference.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of power electronic converter control technology, specifically a voltage sampling circuit management system for high-voltage BUCK topology output. Background Technology

[0002] In high-voltage BUCK topology systems, voltage sampling circuits are typically deployed to achieve accurate monitoring of the output voltage, thereby ensuring stable system operation. The core purpose is to ensure the long-term accuracy and reliability of output voltage sampling under complex operating conditions.

[0003] Existing voltage sampling technologies face severe reliability challenges during long-term operation. The core components in the circuit will experience thermal drift in performance parameters due to temperature changes. In addition, the inherent response delay of the circuit and harmonic interference from the power grid will cause a slow and continuous decline in the accuracy of the voltage sampling signal. If this gradual sampling drift problem cannot be predicted and managed in advance, the control system will make decisions based on inaccurate feedback information, which may eventually lead to system instability or catastrophic failures, thus posing a major threat to the long-term safety of the system. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention provides a voltage sampling circuit management system for high-voltage BUCK topology output. Specifically, the technical solution of this invention includes:

[0005] Risk quantification unit, intelligent adjustment unit, and hierarchical execution unit;

[0006] The risk quantification unit is used to perform risk quantification analysis on the real-time acquired system physical quantities in order to generate trust entropy;

[0007] The intelligent regulation unit is used to construct a state space based on trust entropy, normalized voltage ripple and normalized dynamic response recovery time, and to perform deep reinforcement learning analysis based on the state space in order to output regulation action commands.

[0008] The hierarchical execution unit is used to compare trust entropy with a preset threshold to generate an operating status signal, and to modify or suspend the adjustment action command in response to the operating status signal in order to execute the final control strategy.

[0009] Preferably, the system physical quantities include: the output current of the optocoupler isolation unit, the real-time resistance value of the temperature compensation resistor, the output voltage of the BUCK circuit, the local ambient temperature of the sampling system, and the normalized power grid harmonic interference intensity.

[0010] Preferably, the risk quantification analysis process is as follows:

[0011] Based on the local ambient temperature of the sampling system, the real-time drift of the optocoupler current transfer ratio is determined through a preset drift model function.

[0012] Based on the real-time resistance value of the temperature compensation resistor, the amount of drift to be compensated is determined by using a preset conversion function and drift model function, and taking into account the response delay.

[0013] Trust entropy is generated by integrating the real-time drift amount with the drift amount used for compensation, and in combination with the intensity of power grid harmonic interference.

[0014] The preferred deep reinforcement learning analysis process is as follows:

[0015] Based on the state space and the preset reward function, the switching frequency fine-tuning amount and the output voltage reference offset are determined.

[0016] The adjustment command is generated by combining the switching frequency fine-tuning amount and the output voltage reference offset.

[0017] Preferably, the reward function includes a linear performance penalty term and a nonlinear risk penalty term;

[0018] The linear performance penalty term is used for calculation based on normalized voltage ripple and normalized dynamic response recovery time;

[0019] The nonlinear risk penalty term is used to obtain the risk penalty value by exponentially processing the difference between the trust entropy and the preset security threshold.

[0020] Preferably, the operating status signal includes a safe operating signal, a resilience recovery signal, or an emergency setting signal;

[0021] The comparison and processing procedure is as follows:

[0022] If the trust entropy is less than or equal to the safety threshold in the preset threshold, a safe operation signal is generated;

[0023] If the trust entropy is greater than the safety threshold and less than or equal to the critical threshold in the preset threshold, a resilience recovery signal is generated.

[0024] If the trust entropy is greater than the critical threshold, an emergency tuning signal is generated.

[0025] Preferably, in response to a safety operation signal, the graded execution unit fully executes the adjustment action command.

[0026] Preferably, in response to the resilience recovery signal, the hierarchical execution unit executes the adjustment action instruction; wherein, the adjustment action instruction is an instruction generated by the intelligent adjustment unit under the dominance of the nonlinear risk penalty term, prioritizing the reduction of trust entropy.

[0027] Preferably, in response to an emergency setting signal, the hierarchical execution unit suspends the execution of adjustment action commands and forces the system to switch to preset deterministic safety mode parameters.

[0028] Compared with the prior art, the present invention has the following beneficial effects:

[0029] 1. This system proactively assesses and quantifies the long-term failure risk of the sampling circuit by quantifying the mismatch between the thermal drift of the core components and their compensation delay, thus realizing the transformation from passive response after a fault occurs to proactive prediction in advance, significantly improving the system's early warning capability.

[0030] 2. This system incorporates deep reinforcement learning and constructs a reward function that integrates nonlinear risk penalties, enabling the system to intelligently balance short-term performance metrics with long-term survivability. When potential risks increase, the system can automatically and smoothly shift its control focus from performance optimization to risk avoidance.

[0031] 3. Based on the real-time quantified risk level, this system has constructed three graded execution modes: safe operation, resilient recovery, and emergency tuning. This mechanism ensures that the system pursues optimal performance when the risk is low, prioritizes the recovery of system health when the risk is medium, and switches to deterministic safety parameters when the risk is high, thus ensuring the safety and appropriateness of the control strategy.

[0032] 4. This system incorporates the intensity of power grid harmonic interference into the risk quantification model, which significantly enhances the system's adaptability to harsh external operating conditions. When the power grid environment is complex, the system can more accurately assess and manage accumulated risks, ensuring the long-term accuracy of voltage sampling under complex operating conditions such as high voltage and strong interference. Attached Figure Description

[0033] The present invention will be further explained below with reference to the accompanying drawings and embodiments:

[0034] Figure 1 This is a structural diagram of the system of the present invention. Detailed Implementation

[0035] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to specific embodiments.

[0036] Example 1:

[0037] Please see Figure 1 A voltage sampling circuit management system for high-voltage BUCK topology output includes a risk quantization unit, an intelligent adjustment unit, and a hierarchical execution unit.

[0038] The risk quantification unit is used to perform risk quantification analysis on the real-time acquired system physical quantities in order to generate trust entropy;

[0039] The intelligent regulation unit is used to construct a state space based on trust entropy, normalized voltage ripple and normalized dynamic response recovery time, and to perform deep reinforcement learning analysis based on the state space in order to output regulation action commands.

[0040] The hierarchical execution unit is used to compare trust entropy with a preset threshold to generate an operating status signal, and to modify or suspend the adjustment action command in response to the operating status signal in order to execute the final control strategy.

[0041] This embodiment provides a voltage sampling circuit management system for high-voltage BUCK topology output, aiming to solve the problem of long-term reliability degradation of voltage sampling circuits in the prior art due to factors such as component thermal drift, response delay and power grid harmonic interference. By introducing a risk-quantitative forward-looking adjustment mechanism, a dynamic balance between short-term performance and long-term survivability of the system is achieved.

[0042] A voltage sampling circuit management system for high-voltage BUCK topology output aims to ensure the long-term accuracy and reliability of output voltage sampling under complex high-voltage conditions by real-time quantification and intelligent adjustment of the sampling system's own risks. The system includes a risk quantification unit, an intelligent adjustment unit, and a hierarchical execution unit.

[0043] The purpose of the risk quantification unit is to comprehensively analyze the key physical quantities of the system, proactively assess the long-term drift risk of the voltage sampling signal, and quantify this risk into a measurable index. In this embodiment, the unit collects a set of key physical quantities during system operation in real time, performs in-depth correlation and risk quantification analysis on these physical quantities, and finally generates a core risk index, namely trust entropy. It refers to a dimensionless index used to characterize the cumulative failure risk of a sampling system. Its function is to provide a decision-making basis for subsequent intelligent adjustment and hierarchical control. It is obtained by integrating the degree of mismatch of the physical drift model of the core components.

[0044] The purpose of the intelligent adjustment unit is to generate optimal control and adjustment commands based on the current system risk and performance, so as to achieve a dynamic balance between system performance and reliability. In this embodiment, the unit constructs a multi-dimensional state space based on three key indicators: trust entropy, normalized voltage ripple, and normalized dynamic response recovery time. It refers to the set of system environment states perceived by the deep reinforcement learning agent. Its role is to provide the agent with comprehensive information needed for decision-making. It is constructed by normalizing the trust entropy output by the risk quantification unit and the real-time performance indicators of the system.

[0045] The intelligent adjustment unit performs deep reinforcement learning analysis based on this state space. By maximizing a reward function that incorporates nonlinear risk penalties, it autonomously learns and outputs the optimal adjustment action command. The adjustment action command refers to the specific command generated by the intelligent adjustment unit for fine-tuning the core control parameters of the BUCK circuit. Its function is to actively intervene in the system operation to optimize its long-term or short-term performance. In this embodiment, it is specifically manifested as the fine-tuning amount of the switching frequency and the offset of the output voltage reference.

[0046] The purpose of the hierarchical execution unit is to perform risk review and correction on the instructions output by the intelligent adjustment unit, ensuring that the system can execute the most appropriate control strategy under any risk level. In this embodiment, the unit compares the real-time value of trust entropy with a set of preset thresholds to generate an operating status signal. The operating status signal refers to a discrete signal that represents the current risk level of the system. Its function is to trigger different control execution logics. Its source is the comparison result of trust entropy with safety threshold and critical threshold. In response to different operating status signals, the hierarchical execution unit dynamically corrects or suspends the adjustment action instructions generated by the intelligent adjustment unit, and finally executes a risk-adapted control strategy.

[0047] Through the collaborative work of the aforementioned units, a complete closed-loop control system is constructed, encompassing risk perception, intelligent decision-making, and hierarchical execution. This system not only optimizes the system's immediate performance but also quantifies and manages potential long-term failure risks using the trust entropy metric. It finds the optimal balance between performance and risk through deep reinforcement learning and ensures the safest control strategies are adopted at different risk levels through hierarchical execution logic. This significantly enhances the long-term operational reliability and antifragility of the high-voltage BUCK topology under complex operating conditions, avoiding catastrophic consequences caused by slow drift of the sampling system.

[0048] Example 2:

[0049] The system physical quantities include: the output current of the optocoupler isolation unit, the real-time resistance value of the temperature compensation resistor, the output voltage of the BUCK circuit, the local ambient temperature of the sampling system, and the normalized power grid harmonic interference intensity.

[0050] Based on the embodiment described in Example 1, the system physical quantities analyzed by the risk quantification unit are further defined. The purpose is to clarify the information input source for risk quantification and ensure the comprehensiveness and relevance of the model input. In this embodiment, the system physical quantities specifically include: the output current of the optocoupler isolation unit. Real-time resistance value of temperature compensation resistor Output voltage of the BUCK circuit Local ambient temperature of the sampling system and the normalized power grid harmonic interference intensity ;

[0051] The output current of the optocoupler and the local ambient temperature are the core observables for evaluating the specific thermal drift of the optocoupler current transfer; the resistance value of the temperature compensation resistor is the key state quantity for compensating for this drift; the output voltage of the BUCK circuit is the ultimate control target of the system; and the intensity of grid harmonic interference is the main external interference source affecting the stability of the system. By clearly collecting the above set of highly correlated physical quantities, a solid data foundation is provided for the accuracy of the risk quantification model, enabling the system to gain a deeper understanding of the physical state changes and mismatch relationships of the core components inside the sampling system, thereby achieving earlier and more accurate prediction of long-term drift risks.

[0052] Example 3:

[0053] The risk quantification analysis process is as follows:

[0054] Based on the local ambient temperature of the sampling system, the real-time drift of the optocoupler current transfer ratio is determined through a preset drift model function.

[0055] Based on the real-time resistance value of the temperature compensation resistor, the amount of drift to be compensated is determined by using a preset conversion function and drift model function, and taking into account the response delay.

[0056] Trust entropy is generated by integrating the real-time drift amount with the drift amount used for compensation, and in combination with the intensity of power grid harmonic interference.

[0057] Based on the embodiments described in Example 1, the risk quantification analysis process performed by the risk quantification unit is specifically described to disclose the generation mechanism of trust entropy. The process is designed as a temperature-driven trust entropy model, the core of which is to quantify the cumulative risk caused by the mismatch between the thermal drift of the optocoupler current transfer ratio and the dynamic compensation delay of the NTC thermistor.

[0058] To dynamically assess this risk, this embodiment introduces the instantaneous rate of change of trust entropy, which is calculated as follows:

[0059] ;

[0060] In this formula, Let be the instantaneous rate of change of trust entropy, with dimensions . Through the Integrating over time generates the dimensionless trust entropy. ; It is a drift model function preset according to the optocoupler datasheet or experimental calibration. It is used to describe the dimensionless drift of the current transfer ratio (CTR) as a function of temperature. For example, its form can be: ,in For reference temperature, These are calibration coefficients; The sampling system uses the local ambient temperature, which is collected in real time by a temperature sensor, as input to calculate the real-time drift. ; It is a temperature conversion function preset according to the NTC component datasheet or experimental calibration, used to convert the NTC resistance value into an estimated temperature. For example, it can be based on the Steinhart-Hart equation, and can be in the form of... , where A, B, and C are component characteristic coefficients; It is a dimensionless interference factor, which is determined by fitting through controlled harmonic injection experiments;

[0061] The real-time resistance value of the temperature-compensated resistor, acquired by the resistance measurement circuit and taking into account the response delay, is the response delay. The drift amount used for compensation can be determined by measuring the step temperature response of the NTC element. ; The initial nominal CTR value of the optocoupler, provided in the optocoupler datasheet, is used for normalization. In this model, it is required that... The value must be a positive real number to ensure the validity of the normalization calculation. This condition is satisfied for all physically available optocouplers. It is the dimensionless power grid harmonic interference intensity obtained by collecting and normalizing data from a harmonic sensor.

[0062] To ensure the feasibility of the model, the sources of the calibration parameters in the model are clarified; the basic cumulative coefficients are also explained. The unit is It was obtained through accelerated aging experiments; to clearly distinguish the variables in the calibration process from those in the model runtime, a set of constant ambient temperatures was defined in the calibration dataset. The mean time to failure of the system measured experimentally at this temperature is: By acquiring multiple different sets Data points were analyzed using regression methods such as least squares to evaluate the parameters. The goal of fitting the model is to make the failure time predicted by the model integral consistent with the experimental statistical results. Minimize the error between them; define a set of injected harmonics of known intensity in the calibration dataset as... The stable growth rate of trust entropy, measured experimentally at this harmonic intensity, is: By acquiring multiple sets The data points are fitted to determine the parameters. ;

[0063] The calculation logic of this formula lies in quantifying the mismatch between the real-time drift and the delay-compensated drift, and amplifying its impact through a squared term; simultaneously, it combines the intensity of power grid harmonic interference, and through... This further amplifies the rate of risk accumulation under harsh power grid conditions. The linear superposition model used here is to capture the main contradiction of the positive correlation between harmonic intensity and risk accumulation rate. Although the actual physical effects may be more complex, this model is sufficient to play an effective risk warning role in engineering applications. This process provides a clear and reproducible physical model and mathematical definition for trust entropy, profoundly revealing the inherent contradictions caused by the differences in dynamic characteristics of internal components, and transforming system risk assessment from post-judgment to pre-prediction. It should be noted that the temperature-driven trust entropy model proposed in this embodiment is mainly aimed at the cumulative risk caused by the thermal drift of the core optocoupler element and its compensation mismatch. In other applications, this model framework can also be extended by introducing other physical quantities that characterize system aging, such as the equivalent series resistance of the output filter capacitor, to construct a more multi-dimensional risk assessment model to more comprehensively depict the long-term survivability of the system.

[0064] Example 4:

[0065] The deep reinforcement learning analysis process is as follows:

[0066] Based on the state space and the preset reward function, the switching frequency fine-tuning amount and the output voltage reference offset are determined.

[0067] The adjustment command is generated by combining the switching frequency fine-tuning amount and the output voltage reference offset.

[0068] The reward function includes a linear performance penalty term and a non-linear risk penalty term;

[0069] The linear performance penalty term is used for calculation based on normalized voltage ripple and normalized dynamic response recovery time;

[0070] The nonlinear risk penalty term is used to obtain the risk penalty value by exponentially processing the difference between the trust entropy and the preset security threshold.

[0071] Based on the embodiments described in Example 1, the deep reinforcement learning analysis process of the intelligent adjustment unit and its core reward function are explained in detail to reveal how the intelligent agent makes decisions to balance short-term performance and long-term risk.

[0072] The intelligent adjustment unit is based on state space. With the preset reward function Decisions are made using a deep Q-network algorithm; Trust entropy Normalized voltage ripple and normalized dynamic response recovery time Composition; among which, normalized voltage ripple With normalized dynamic response recovery time By adjusting the output voltage of the BUCK circuit Real-time sampling and analysis are performed to evaluate key indicators of the system's short-term dynamic and steady-state performance, including the agent's action space. This includes the switching frequency of the BUCK circuit. fine-tuning amount and the output voltage reference value offset At each decision point, the agent selects an action that maximizes the cumulative reward in the future based on the current state. That is, based on the state space and the preset reward function, it determines the switching frequency fine-tuning amount and the output voltage reference offset, and combines the two to generate the adjustment action command.

[0073] In this process, the reward function The construction of this mechanism is key to guiding an agent to achieve a dynamic balance between short-term performance and long-term survivability. Its calculation method is as follows:

[0074] ;

[0075] The reward function is a dimensionless scalar, comprising a linear performance penalty term and a nonlinear risk penalty term; the linear performance penalty term... For use based on normalized voltage ripple With normalized dynamic response recovery time Calculations are performed; among which, the dimensionless performance weighting coefficients are used. and These are hyperparameters determined based on the steady-state and dynamic performance requirements of specific application scenarios; this term is a negative reward, and its value is proportional to the degree of performance degradation, incentivizing the agent to optimize conventional performance.

[0076] Non-linear risk penalty term Used for trust entropy-based With preset safety threshold The difference is indexed to obtain the risk penalty value; where the dimensionless risk weight coefficient is... Optimization is achieved through multiple rounds of iterative training to find the optimal balance between performance and risk; a dimensionless safety threshold. With scale factor The determination of this depends on the prior analysis of the trust entropy model: through simulation calculations, different... The value is mapped to the probability of future sampling failures in the system, and a value corresponding to an acceptablely low failure rate is selected, for example... of Value as And according to the failure rate To set the growth gradient Value; when much smaller At that time, this punishment can be ignored; but when Approaching or exceeding At this point, the negative reward will increase exponentially, creating a huge negative reward that forces the agent to take any action that can reduce the negative reward most quickly. The action; by designing this composite reward function, the intelligent adjustment unit is endowed with an antifragile decision-making ability, enabling it to automatically switch the strategy focus to risk avoidance when the risk increases, which is not available in existing linear control strategies.

[0077] Example 5:

[0078] Operating status signals include safe operation signals, resilience recovery signals, or emergency setting signals;

[0079] The comparison and processing procedure is as follows:

[0080] If the trust entropy is less than or equal to the safety threshold in the preset threshold, a safe operation signal is generated;

[0081] If the trust entropy is greater than the safety threshold and less than or equal to the critical threshold in the preset threshold, a resilience recovery signal is generated.

[0082] If the trust entropy is greater than the critical threshold, an emergency tuning signal is generated.

[0083] In response to the safety operation signal, the graded execution unit fully executes the adjustment action command;

[0084] In response to the resilience recovery signal, the hierarchical execution unit executes adjustment action instructions; wherein, the adjustment action instructions are instructions generated by the intelligent adjustment unit under the dominance of the nonlinear risk penalty term, with the priority of reducing trust entropy;

[0085] In response to an emergency setting signal, the hierarchical execution unit suspends the execution of adjustment action commands and forces the system to switch to preset deterministic safety mode parameters.

[0086] Based on the embodiments described in Example 1, the specific logic of the hierarchical execution unit is explained to disclose how the adjustment action instructions are modified or suspended according to real-time risks.

[0087] The core of the hierarchical execution unit lies in the comparison processing process, which involves trust entropy. The continuous values ​​are mapped to three discrete operating state regions, thereby generating a safe operation signal, a resilience recovery signal, or an emergency setting signal; this process relies on two preset thresholds: a safety threshold and a resilience recovery signal. and critical threshold ;in, Numerically equal to the reward function used ,and It is the trust entropy value corresponding to the risk point where the system state may become irreversible, determined based on simulation analysis, and ;

[0088] If trust entropy Less than or equal to the safety threshold If this signal is received, a safe operation signal will be generated. In response to this signal, the hierarchical execution unit will fully execute the adjustment action instructions output by the intelligent adjustment unit. This ensures that the performance optimization capabilities of the DRL agent can be fully utilized when the system is safe.

[0089] If trust entropy Greater than the safety threshold And less than or equal to the critical threshold This generates a resilience recovery signal; in response to this signal, the adjustment action instructions executed by the hierarchical execution unit are instructions generated under the dominance of the nonlinear risk penalty term in the reward function, prioritizing the reduction of trust entropy; because A higher value means that the exponential penalty term in the reward function forces the DRL agent to develop new optimal policies, such as choosing one that most effectively reduces... The mechanism enables actions that may temporarily increase voltage ripple or response time, even if these actions temporarily increase voltage ripple or response time; this mechanism achieves a smooth transition from performance optimization to risk avoidance, demonstrating a high degree of intelligence.

[0090] If trust entropy Greater than the critical threshold If this occurs, an emergency setting signal is generated; in response to this signal, the tiered execution unit will suspend the execution of the adjustment action command and force the system to switch to preset deterministic safety mode parameters; for example, adjusting the switching frequency. The frequency is forcibly set to 80% of the rated operating frequency. This value has been experimentally verified as the lowest frequency at which the output voltage reference can operate stably under any load. The original value is reduced by 2% to increase the system's stability margin, ensuring that hardware damage caused by overvoltage or oscillation can be avoided under any circumstances. This mechanism provides the system with a last line of defense. When the model predicts that the system is on the verge of failure, switching to fully validated deterministic safety parameters can guarantee the basic functions and hardware safety of the system with the highest probability, reflecting the design principle of fail-safety.

[0091] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A voltage sampling circuit management system for high-voltage BUCK topology output, characterized in that, It includes a risk quantification unit, an intelligent adjustment unit, and a tiered execution unit; The risk quantification unit is used to perform risk quantification analysis on the real-time acquired system physical quantities in order to generate trust entropy; The intelligent regulation unit is used to construct a state space based on trust entropy, normalized voltage ripple and normalized dynamic response recovery time, and to perform deep reinforcement learning analysis based on the state space in order to output regulation action commands. The hierarchical execution unit is used to compare trust entropy with a preset threshold to generate an operating status signal, and to modify or suspend the adjustment action command in response to the operating status signal in order to execute the final control strategy. The risk quantification analysis process is as follows: Based on the local ambient temperature of the sampling system, the real-time drift of the optocoupler current transfer ratio is determined through a preset drift model function. Based on the real-time resistance value of the temperature compensation resistor, the amount of drift to be compensated is determined by using a preset conversion function and drift model function, and taking into account the response delay. Based on the mismatch between the real-time drift and the drift used for compensation, and combined with the intensity of power grid harmonic interference, the trust entropy is generated by integrating over time using a mathematical model of the instantaneous rate of change of trust entropy. The formula for calculating the instantaneous rate of change of trust entropy is: ; in, Let the instantaneous rate of change of trust entropy be denoted as . Based on the cumulative coefficient, For the preset drift model function, The sampling system collects local ambient temperature data in real time. For the preset temperature conversion function, The real-time resistance value of the temperature compensation resistor is taken into account for response delay. To respond to the delay, This represents the initial nominal current transfer ratio of the optocoupler. The interference factor is dimensionless. The normalized power grid harmonic interference intensity; The deep reinforcement learning analysis process is as follows: Based on the state space and the preset reward function, the switching frequency fine-tuning amount and the output voltage reference offset are determined. By combining the switching frequency fine-tuning amount with the output voltage reference offset, an adjustment action command is generated; The calculation formula for the reward function is as follows: ; In the formula, For the reward function, and For performance weighting coefficients, For risk weighting coefficients, For normalized voltage ripple, To normalize the dynamic response recovery time, For trust entropy, As a safety threshold, is the scale factor.

2. A voltage sampling circuit management system for high-voltage BUCK topology output according to claim 1, characterized in that, The system physical quantities include: the output current of the optocoupler isolation unit, the real-time resistance value of the temperature compensation resistor, the output voltage of the BUCK circuit, the local ambient temperature of the sampling system, and the normalized power grid harmonic interference intensity.

3. A voltage sampling circuit management system for high-voltage BUCK topology output according to claim 1, characterized in that, The reward function includes a linear performance penalty term and a non-linear risk penalty term; Linear performance penalty term used for normalized voltage ripple With normalized dynamic response recovery time Perform the calculation; the formula for calculating the linear performance penalty term is: ; The nonlinear risk penalty term is used to obtain the risk penalty value by exponentially processing the difference between the trust entropy and the preset security threshold; the calculation formula for the nonlinear risk penalty term is: .

4. A voltage sampling circuit management system for high-voltage BUCK topology output according to claim 1, characterized in that, Operating status signals include safe operation signals, resilience recovery signals, or emergency setting signals; The comparison and processing procedure is as follows: If the trust entropy is less than or equal to the safety threshold in the preset threshold, a safe operation signal is generated; If the trust entropy is greater than the safety threshold and less than or equal to the critical threshold in the preset threshold, a resilience recovery signal is generated. If the trust entropy is greater than the critical threshold, an emergency tuning signal is generated.

5. A voltage sampling circuit management system for high-voltage BUCK topology output according to claim 4, characterized in that, In response to the safety operation signal, the hierarchical execution unit fully executes the adjustment action command.

6. A voltage sampling circuit management system for high-voltage BUCK topology output according to claim 4, characterized in that, In response to the resilience recovery signal, the hierarchical execution unit executes adjustment action instructions; among them, the adjustment action instructions are instructions generated by the intelligent adjustment unit under the dominance of the nonlinear risk penalty term, with the priority of reducing trust entropy.

7. A voltage sampling circuit management system for high-voltage BUCK topology output according to claim 4, characterized in that, In response to an emergency setting signal, the hierarchical execution unit suspends the execution of adjustment action commands and forces the system to switch to preset deterministic safety mode parameters.

Citation Information

Patent Citations

  • Transformer voltage regulation cooperative control system based on data acquisition of Internet of Things

    CN121417246A

  • Intelligent electric equipment monitoring and optimizing method

    CN121485294A