Non-contact high-voltage circuit breaker speed detection device

By combining non-contact sensors with meta-learning and reinforcement learning methods, the installation difficulties and safety issues of traditional high-voltage circuit breaker speed detection have been solved, achieving high-precision and real-time speed detection.

CN121856581APending Publication Date: 2026-04-14HEBEI HUAWAN ELECTRONIC TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-09-26
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Traditional methods for detecting the speed of high-voltage circuit breakers rely on contact sensors, which have problems such as demanding installation conditions, susceptibility to bumps or impacts, inability to achieve embedded online design, and long measurement time, and cannot meet the high requirements of modern power systems for real-time performance and accuracy.

Method used

By employing non-contact sensors combined with meta-learning and reinforcement learning methods, and through a process of acquisition-processing-meta-learning + reinforcement learning-computation-output, real-time and high-precision detection of the speed of high-voltage circuit breakers can be achieved.

Benefits of technology

It achieves real-time, high-precision, non-contact detection of high-voltage circuit breaker speed, avoiding the safety risks to equipment posed by traditional contact-type speed measurement, and improving the accuracy and efficiency of speed measurement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121856581A_ABST
    Figure CN121856581A_ABST
Patent Text Reader

Abstract

The invention discloses a non-contact high-voltage circuit breaker speed detection device. The device mainly solves the problems of safety risk and efficiency possibly brought by an existing contact type sensor in the speed measurement process. The device is combined with the content of a current'communication-perception-calculation 'integrated research subject for the first time, and comprises a non-contact sensor used for detecting the motion state of the high-voltage circuit breaker; the communication module is used for receiving data of the sensor and sending the data to the reinforcement learning module; the meta learning and reinforcement learning module is used for learning and optimizing a speed measurement algorithm; and the calculation module is used for processing the output of the meta learning and reinforcement learning module and calculating the speed of the high-voltage circuit breaker. According to the invention, non-contact high-precision speed measurement can be realized, the speed measurement efficiency is improved, and the safety risk is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of power system detection and control, and particularly relates to a non-contact high-voltage circuit breaker speed detection device. Background Technology

[0002] In the field of power system monitoring and control, the speed detection of high-voltage circuit breakers is a crucial issue. The opening and closing speed of a high-voltage circuit breaker is one of the key factors affecting its breaking performance. If the opening and closing speed is too fast, the operating mechanism and transmission components of the circuit breaker may exceed their mechanical strength tolerance, causing damage to the buffers and shortening the circuit breaker's service life. If the opening and closing speed is too slow, it may not be able to quickly interrupt normal operating current or fault current, potentially leading to prolonged arcing time between contacts, contact burn-out, or even explosion of the arc-extinguishing chamber. Therefore, accurate and real-time detection of the speed of high-voltage circuit breakers is of great significance for ensuring the safe operation of the power system.

[0003] However, traditional speed measurement methods primarily rely on contact sensors, which present several problems. First, contact sensors require installation on the crank arm of the high-voltage circuit breaker mechanism box or the crank arm of the transmission linkage, making installation conditions quite demanding; otherwise, inaccurate speed measurements or the absence of speed curves may occur. Second, contact sensors are easily bumped or impacted during detection, and any impact to the switch body could pose a safety risk. Furthermore, contact sensors cannot be embedded in an online design and cannot be integrated into the intelligent module of a smart switch. Finally, contact sensors have long measurement times and low efficiency, failing to meet the high real-time requirements of modern power systems. Summary of the Invention

[0004] To address the aforementioned issues, this invention proposes a non-contact high-voltage circuit breaker speed detection device, specifically a non-contact high-voltage circuit breaker speed detection device based on meta-learning and reinforcement learning. This device acquires data through a non-contact sensor, processes and learns the data through a meta-learning and reinforcement learning module, calculates the speed through a speed calculation module, and finally outputs the results through a display device or communication module. In this process, the meta-learning and reinforcement learning module learns how to extract the most useful information from the data through continuous trial and error and feedback, thereby improving the accuracy and efficiency of speed measurement. Simultaneously, the non-contact sensor design avoids the safety risks that traditional contact-based speed measurement methods may pose to high-voltage circuit breakers. Therefore, this device has broad application prospects and can be applied to scenarios such as power system operation monitoring, equipment maintenance, and fault diagnosis.

[0005] In power systems, high-voltage circuit breakers are critical equipment. Their main function is to quickly interrupt fault currents when a power system fault occurs, preventing the fault from escalating, while ensuring the continuous operation of the power system under normal conditions. The operating speed of a high-voltage circuit breaker, especially its opening and closing speed, directly affects its breaking performance. However, traditional speed measurement methods mainly rely on contact sensors, which have many problems, such as demanding installation conditions, susceptibility to impacts, inability to achieve embedded online design, and long measurement times. These problems limit their application in modern power systems. To solve these problems, researchers have begun to search for new speed measurement methods. Among them, non-contact sensors have received widespread attention because they do not require direct contact with the high-voltage circuit breaker, thus avoiding many of the problems of contact sensors. However, how to extract useful information from the large amount of data collected by non-contact sensors, and how to use this information for accurate speed calculation, are problems that need to be solved.

[0006] Furthermore, with the development of artificial intelligence technology, meta-learning and reinforcement learning, as methods that can learn how to extract the most useful information from data through continuous trial and error and feedback, have also been introduced into the speed detection of high-voltage circuit breakers. However, designing and implementing a non-contact high-voltage circuit breaker speed detection device based on meta-learning and reinforcement learning presents a completely new challenge.

[0007] Therefore, this invention proposes a non-contact high-voltage circuit breaker speed detection device, which aims to solve the above-mentioned problems by achieving real-time, high-precision, and non-contact detection of high-voltage circuit breaker speed through a process of acquisition-processing-meta-learning + reinforcement learning-calculation-output, and provides a new high-voltage circuit breaker speed detection device.

[0008] The embodiments of the present invention provide the following technical solutions:

[0009] A non-contact high-voltage circuit breaker speed detection device, applicable to high-voltage circuit breakers, combines relevant content from the "communication-sensing-computing" integrated research project. Through a process of data acquisition, processing, meta-learning + reinforcement learning, computation, and output, it achieves real-time, high-precision, non-contact detection of the speed of high-voltage circuit breakers. The device includes:

[0010] Non-contact sensor deployment module: Non-contact sensors are deployed near the high-voltage circuit breaker to monitor the movement status of the high-voltage circuit breaker. Such sensors can be laser speed sensors.

[0011] Data collection module: When the high-voltage circuit breaker starts to move, the non-contact sensor detects the movement of the high-voltage circuit breaker and collects relevant data, including the movement speed, direction of movement, and distance of movement of the high-voltage circuit breaker.

[0012] Data communication module: The collected data is sent to the reinforcement learning module through the communication module. This communication module can be wired or wireless, depending on the specific application scenario and requirements.

[0013] Data processing and meta-learning + reinforcement learning module: After receiving the data, the meta-learning + reinforcement learning module will process and learn through a pre-set algorithm. During this process, the meta-learning + reinforcement learning module will continuously optimize the speed measurement algorithm to improve the accuracy and efficiency of speed measurement.

[0014] Speed ​​calculation module: The data processed and learned by the meta-learning + reinforcement learning module is sent to the calculation module, which calculates the real-time speed of the high-voltage circuit breaker based on this data.

[0015] Result output module: The calculated high-voltage circuit breaker speed can be displayed on a display device or sent to other devices or systems via a communication module;

[0016] The deployment module for the non-contact sensor specifically includes:

[0017] Based on the specific conditions of the high-voltage circuit breaker and the speed measurement requirements, select an appropriate non-contact sensor, such as a laser speed sensor.

[0018] Determine the deployment location of the non-contact sensor so that it can effectively monitor the movement status of the high-voltage circuit breaker. This location should be as close as possible to the high-voltage circuit breaker, but without interfering with its normal movement. At the same time, the safety of the sensor also needs to be considered to avoid potential damage to the sensor when the high-voltage circuit breaker moves.

[0019] After determining the deployment location, the non-contact sensor is installed at that location. The installation process should ensure the stability of the sensor to prevent the sensor's position from shifting when the high-voltage circuit breaker moves.

[0020] After installation, the non-contact sensor needs to be tested to ensure that it can work properly and accurately detect the movement status of the high-voltage circuit breaker. The testing process can be carried out by simulating the movement of the high-voltage circuit breaker and observing the sensor's response.

[0021] The data collection module specifically includes:

[0022] Non-contact sensors monitor the motion state of high-voltage circuit breakers in real time and collect relevant data, including motion speed, motion direction, and motion distance, which are set as state space S. Each specific state is represented as s∈S, and s=[v,d,D] is defined, where v represents motion speed, d represents motion direction, and D represents motion distance. These data will be used as input to the meta-learning model.

[0023] After establishing the "state space," the collected raw dataset {v1,v2,v3...d1,d2,d3...D1,D2,D3...} needs preliminary processing, including data cleaning, formatting, and normalization, to obtain the processed dataset {v1′,v′2,v3′...d1′,d2′,d3′...D1′,D2′,D3′...}, which facilitates subsequent analysis and use. This process can be considered the data preprocessing stage, preparing for the training of the reinforcement learning model. i d represents the operating speed of the high-voltage circuit breaker in the i-th cycle; i D represents the specific direction of movement of the high-voltage circuit breaker in the i-th cycle; i This represents the actual travel distance of the high-voltage circuit breaker in the i-th cycle; v i ′ represents the velocity of the high-voltage circuit breaker collected after data processing in the i-th cycle; d i ′ represents the specific direction of motion of the high-voltage circuit breaker after data processing in the i-th cycle; D i ′ represents the actual movement distance of the high-voltage circuit breaker after data processing in the i-th cycle;

[0024] The collected data and extracted features are used to train a reinforcement learning model. The action space is usually represented as A. An action vector needs to be defined to capture all possible decisions, hence A = [a, b, c], where a represents the chosen speed measurement algorithm, b represents the algorithm's parameters, and c represents the method of processing the data. In the action space A = [a1, a2... b1, b2... c1, c2...], a1 represents the time difference-based method (using two or more sensors to measure the time difference of an object's passage to calculate speed), a2 represents the frequency-based method (using sensors to measure the frequency of an object's passage to calculate speed); b1 represents the distance between the sensors, b2 represents the time window for measuring the frequency; c1 represents using a filter to remove noise, and c2 represents using a smoother to reduce data fluctuations. During training, the model learns how to extract the most useful information from the data through continuous trial and error and feedback to improve the accuracy and efficiency of speed measurement. This process can be described by the following Bellman equation:

[0025]

[0026] Where Q(v) i A i ) indicates that in state v i Take action A i Action value function, p(v i+1 ',r1∣v i ,ai ) indicates that in state v i Take action a i Then, transition to state v i+1 And the probability of receiving a reward r1, p(v i+1 ',r2∣v i ,b i ) indicates that in state v i Take action b i Then, transition to state v i+1 And the probability of receiving a reward of r2, p(v i+1 ',r3∣v i ,c i ) indicates that in state v i Take action c i Then, transition to state v i+1 'and the probability of obtaining a return of r3; γ1, γ2, and γ3 are their respective discount factors; π(a i+1 '∣v i+1 '), π(b) i+1 '∣v i+1 '), π(c i+1 '∣v i+1 ') indicates that in state v i+1 'The following actions are taken respectively a i+1 '、b i+1 '、c i+1 The strategy; α, β, and λ represent the actual weights corresponding to different actions; Q π (v i+1 ',a i+1 '), Q π (v i+1 ',b i+1 '), Q π (v i+1 ',c i+1 ') respectively represent the next time period in state v i+1 Take action a respectively i+1 '、b i+1 '、c i+1 'Action value function;

[0027] Reinforcement learning models are used to evaluate the quality of the collected data. If the data quality is not high, such as excessive noise or insufficient coverage, the settings of the non-contact sensors or the data collection strategy can be adjusted to improve the data quality.

[0028] Based on the data quality assessment results and feedback from the reinforcement learning model, the data collection strategy is optimized. This optimization process is continuous, aiming to continuously improve the quality and validity of the data to enhance the accuracy and efficiency of speed measurement. This process can be achieved by adjusting the reinforcement learning model's policy π(A). i |v i This is achieved by, given a state v i Next, select the action value function Q(v) that will make the action value function Q(v) work. i A i The largest action A i ;

[0029] The data communication module specifically includes:

[0030] The collected data needs to be encoded and encapsulated for transmission via the communication module. During this process, each state s needs to be associated with its corresponding action p and reward r to facilitate learning and decision-making by the reinforcement learning model. The state space s in time period i is defined. i =[q i ,w i ,e i ], where q i The format representing the data (such as binary, ASCII, etc.), w i Represents the data encapsulation protocol (such as TCP / IP, UDP, etc.), e i The compression method of the data (e.g., no compression, ZIP compression, GZIP compression, etc.) defines the action space p in time period i. i =[t i ,u i ,o i ], where t i This represents changing the data format, u i This represents a change in the data encapsulation protocol, o i This represents a change in the way data is compressed;

[0031] This process can be described by the following formula:

[0032]

[0033] in, Indicates that in state s i Take action P i The action-value function, r is the reward, π t (t i+1 '∣s i+1 '), π u (u i+1 '∣s i+1 '), π o (o i+1'∣s i+1 ') represent states s and s respectively. i+1 'Take action t' i+1 '、u i+1 '、o i+1 The strategy can be ε-greedy, Softmax, or UCB (Upper Confidence Bound) depending on the actual requirements; α, β, and λ represent the actual weights corresponding to different actions. These represent the states in the next time period, s and s respectively. i+1 Take action t respectively i+1 '、u i+1 '、o i+1 'Action value function;

[0034] The encoded and encapsulated data is sent to the reinforcement learning module via the communication module. During data transmission, it is necessary to ensure data integrity and real-time performance to guarantee that the reinforcement learning model can accurately receive the data and make timely learning and decisions. This process requires continuous updating and optimization of the policy π. p (p i+1 '∣s i+1 '), so that the action value function maximum;

[0035] After the reinforcement learning module receives data, it needs to decode and depackage it to restore the original data format and structure. Then, the reinforcement learning model can use this data for learning and decision-making. In this process, the action-value function needs to be updated based on the newly received data. To reflect new status and reporting information;

[0036] The data processing and meta-learning + reinforcement learning module specifically includes:

[0037] After receiving the data, the meta-learning + reinforcement learning module first performs data parsing to restore the original format and structure of the data. The dataset θ={v1″,v′2′,v′3′...d1″,d2″,d3″...D1″,D′2′,D3″}, which was previously processed by data cleaning, formatting, and normalization, is set as the initial parameters for meta-learning, and the embedded meta-learning module is run.

[0038] For each task i, perform one (or more) gradient descent steps on the model to obtain a new parameter set θ. i If we perform a gradient descent step, the new parameters can be calculated using the following formula:

[0039]

[0040] Where α is the learning rate. η is the gradient of the loss function of model F with respect to task i under parameter set θ. i This is to compensate for the difference between the calculation results of the device module and the actual parameter values ​​required by task i. This value is calculated by statistically analyzing past numerical characteristics.

[0041] Then, we update our initial parameter set θ by minimizing the expected loss after one step of gradient descent on all tasks, which can be achieved by the following formula:

[0042]

[0043] Where β is the meta-learning rate. It is the expected loss of all tasks after one step of gradient descent. η is the compensation between the computational results of the device module and the actual required parameter values. This value is calculated by statistically analyzing past numerical characteristics. If necessary, multiple gradient descent steps can be performed to find a better initial parameter set θ. * The initial parameter set is then transmitted to the outer reinforcement learning module for further processing.

[0044] Based on the better initial parameter set θ obtained after analysis * To further calculate the action value function, this process requires considering the current state parameter set θ. * Given actions g∈[h,j,k], calculate the action-value function Q. π (θ, g) represents the value of taking an action in the current state, where h represents choosing different data processing methods, j represents choosing different learning algorithms, and k represents adjusting the model parameters. This process can be described by the following formula:

[0045]

[0046] in Indicates in parameter set θ * The action value function of the action g is taken; π h (h'∣θ * '), π j (j'∣θ * '), π k (k'∣θ * ') respectively represent the parameters in the parameter set θ * The strategy is to take actions h', j', and k'; γ1, γ2, and γ3 are the corresponding discount factors. They represent the parameters in the set θ respectively. * The action value functions for actions h', j', and k' are respectively taken below;

[0047] According to the new parameter set θ * and action value function The reinforcement learning module updates its policy π. h (h'∣θ * '), π j (j'∣θ * '), π k (k'∣θ * '), so that the action value function This process can be viewed as the learning process of a reinforcement learning model. Through continuous trial and error and feedback, the model will gradually optimize its strategy to improve the accuracy and efficiency of speed measurement.

[0048] According to the updated strategy π(g'∣θ) * The reinforcement learning module will decide on the current parameter set θ. * The action g to be taken next might be adjusting the settings of the non-contact sensor or changing the data collection strategy to improve the quality and validity of the data. During this process, our action value function needs to be updated based on the new action g and the reward r. To reflect new actions and reward information;

[0049] The speed calculation module specifically includes:

[0050] After receiving the data sent by the meta-learning + reinforcement learning module, the data is first parsed to restore its original format and structure. This process is to ensure the integrity and accuracy of the data and to provide accurate input for subsequent calculations.

[0051] Based on the parsed data, the current state and action are identified. During this process, the action value function Q(s,a;θ) and the state value function V(s;w) need to be calculated based on the current state and action. Using the Actor-Critic method, the updated formulas are as follows:

[0052] Critic Update: Update the value function V using the TD error δ.

[0053] δ t =r t +γ*V(s t+1 ;w)-V(s t ;w)

[0054]

[0055] Actor Update: Update policy π using TD error δ.

[0056]

[0057] Based on the strategy π and the action value function Q, the calculation module can calculate the real-time speed of the high-voltage circuit breaker. This process can be viewed as the decision-making process of the meta-learning model. By selecting the action that maximizes the action value function, the optimal speed calculation result can be obtained. This process can be expressed as:

[0058] a t =argmaxaQ(s t ,a;θ)

[0059] This formula indicates that in state s t Next, select the action a that maximizes the action value function Q. t As the optimal action;

[0060] The calculated real-time speed of the high-voltage circuit breaker is fed back to the meta-learning + reinforcement learning module as part of the new state. Then, the meta-learning + reinforcement learning module will adjust the speed based on the new state s. t+1 and returns t+1 Then, update its policy π and action value function Q. The update formula for this process is as follows:

[0061] δ t+1 =r t+1 +γ*V(s t+2 ;w)-V(s t+1 ;w)

[0062]

[0063]

[0064] This formula describes how the meta-learning model adapts to the new state s. t+1 and returns t+1 The method for updating its policy π and action value function Q;

[0065] The meta-learning + reinforcement learning module updates its policy π and action value function Q based on the new state and reward to further optimize the results of speed calculation. This process uses the same update formula as the Actor-Critic method as the update formula of the E4 step, except that the time step of the state and reward is pushed back one step.

[0066] The speed calculation module verifies the calculated speed to ensure its accuracy and reliability. If the verification result does not meet the preset standard, the data processing and learning strategies can be adjusted to improve the accuracy of the speed calculation. In this process, the predictive ability of the meta-learning model can be used to evaluate the accuracy of the speed calculation by comparing the predicted speed with the actual speed.

[0067] If the verification results do not meet the preset standards, the speed calculation module can send a signal to the meta-learning module to indicate that the current speed calculation model needs further optimization. The meta-learning module will then adjust the update strategy for each task based on this signal, thereby improving the speed calculation accuracy of the model. This process can be viewed as finding a better speed calculation strategy in the policy space, and can be represented by the following formula:

[0068]

[0069] Where β is the meta-learning rate, E i [L i (F(θ i ′))] is the expected loss of all tasks after one step of gradient descent;

[0070] In this way, a complete cycle of learning, calculation, feedback, and optimization is completed, enabling the entire system to continuously optimize itself and improve the calculation accuracy and efficiency of the high-voltage circuit breaker speed.

[0071] The result output module specifically includes:

[0072] Before outputting the results, the calculated velocity data needs to be processed, including data formatting and unit conversion, to meet the output requirements. In this process, the decision-making ability of a reinforcement learning model can be used to determine the data processing method by selecting the optimal action. This process can be described by the following formula:

[0073]

[0074] Among them, a * Q represents the optimal action. π (s,a) represents the action-value function for taking action a in state s. This formula describes how to choose the optimal action a given state s. In this case, state s can represent the current data processing requirements (data format, units, precision), action a can represent different data processing methods (data formatting), and the action-value function Q... π (s,a) represents the effect of the data processing method;

[0075] The processed speed data can be displayed on a display device. In this process, the decision-making ability of a reinforcement learning model can be used to determine the display method and format by selecting the optimal action. This can be achieved by displaying the speed value, the speed change curve, or speed statistics. This process can also be described by the formula mentioned above, where state s represents the current display requirement, action a represents different display methods, and the action-value function Q...π (s,a) can represent the effect of adopting a certain display method;

[0076] The processed speed data can also be sent to other devices or systems via the communication module. In this process, the decision-making ability of a reinforcement learning model can be used to determine the communication method and protocol by selecting the optimal action. This can be achieved by choosing between wired or wireless transmission, or by selecting different communication protocols such as TCP / IP, Bluetooth, or Zigbee. This process can also be described by the formula mentioned above, where state s represents the current communication requirement, action a represents different communication methods, and the action-value function Q... π (s,a) can represent the effect of adopting a certain communication method.

[0077] Compared with existing technologies, the above technical solution has the following advantages:

[0078] This invention relates to a non-contact high-voltage circuit breaker speed detection device. This device is suitable for speed detection of high-voltage circuit breakers and, for the first time, integrates with current research on the "communication-sensing-computing" integration. Through a process of data acquisition, processing, meta-learning + reinforcement learning, computation, and output, it achieves real-time, high-precision, and non-contact speed detection of high-voltage circuit breakers. The device utilizes meta-learning and reinforcement learning methods to optimize data processing and speed calculation strategies, improving the accuracy and efficiency of speed measurement. Simultaneously, the non-contact design of this device avoids the safety risks that traditional contact-based speed measurement methods may pose to high-voltage circuit breakers. Therefore, this device has broad application prospects and can be applied to scenarios such as power system operation monitoring, equipment maintenance, and fault diagnosis. Attached Figure Description

[0079] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0080] Figure 1 This is a schematic diagram of a non-contact high-voltage circuit breaker speed detection device disclosed in this invention. Detailed Implementation

[0081] In modern power systems, the performance of high-voltage circuit breakers directly affects the safety and stability of the power grid. Among these performance indicators, the breaking speed of a high-voltage circuit breaker is a crucial performance metric. However, traditional testing methods often suffer from problems such as high contact requirements, low accuracy, and low efficiency. To address these issues, this invention proposes a non-contact high-voltage circuit breaker speed detection device.

[0082] The core idea of ​​this invention lies in integrating the concepts of "communication-sensing-computing" and keeping pace with the development of artificial intelligence technology. It introduces meta-learning and reinforcement learning into the speed detection of high-voltage circuit breakers, achieving real-time, high-precision, and non-contact speed detection. Specifically, this invention divides the entire detection process into five steps: "acquisition-processing-meta-learning + reinforcement learning-computation-output." Each step is optimized by a reinforcement learning algorithm. Furthermore, the data processing and meta-learning + reinforcement learning modules incorporate meta-learning algorithms to obtain the optimal initial parameter set, thereby improving detection accuracy and efficiency.

[0083] See Figure 1 This invention provides a non-contact high-voltage circuit breaker speed detection device, suitable for speed detection of high-voltage circuit breakers. Combining the relevant content of the "communication-sensing-computing" integrated research project, it achieves real-time, high-precision, non-contact detection of high-voltage circuit breaker speed through a process of acquisition-processing-meta-learning + reinforcement learning-computation-output. The device includes:

[0084] Non-contact sensor deployment module: Non-contact sensors are deployed near the high-voltage circuit breaker to monitor the movement status of the high-voltage circuit breaker. Such sensors can be laser speed sensors.

[0085] Data collection module: When the high-voltage circuit breaker starts to move, the non-contact sensor detects the movement of the high-voltage circuit breaker and collects relevant data, including the movement speed, direction of movement, and distance of movement of the high-voltage circuit breaker.

[0086] Data communication module: The collected data is sent to the reinforcement learning module through the communication module. This communication module can be wired or wireless, depending on the specific application scenario and requirements.

[0087] Data processing and meta-learning + reinforcement learning module: After receiving the data, the meta-learning + reinforcement learning module will process and learn through a pre-set algorithm. During this process, the meta-learning + reinforcement learning module will continuously optimize the speed measurement algorithm to improve the accuracy and efficiency of speed measurement.

[0088] Speed ​​calculation module: The data processed and learned by the meta-learning + reinforcement learning module is sent to the calculation module, which calculates the real-time speed of the high-voltage circuit breaker based on this data.

[0089] Result output module: The calculated high-voltage circuit breaker speed can be displayed on a display device or sent to other devices or systems via a communication module;

[0090] The deployment module for the non-contact sensor specifically includes:

[0091] Based on the specific conditions of the high-voltage circuit breaker and the speed measurement requirements, select an appropriate non-contact sensor, such as a laser speed sensor.

[0092] Determine the deployment location of the non-contact sensor so that it can effectively monitor the movement status of the high-voltage circuit breaker. This location should be as close as possible to the high-voltage circuit breaker, but without interfering with its normal movement. At the same time, the safety of the sensor also needs to be considered to avoid potential damage to the sensor when the high-voltage circuit breaker moves.

[0093] After determining the deployment location, the non-contact sensor is installed at that location. The installation process should ensure the stability of the sensor to prevent the sensor's position from shifting when the high-voltage circuit breaker moves.

[0094] After installation, the non-contact sensor needs to be tested to ensure that it can work properly and accurately detect the movement status of the high-voltage circuit breaker. The testing process can be carried out by simulating the movement of the high-voltage circuit breaker and observing the sensor's response.

[0095] The data collection module specifically includes:

[0096] Non-contact sensors monitor the motion state of high-voltage circuit breakers in real time and collect relevant data, including motion speed, motion direction, and motion distance, which are set as state space S. Each specific state is represented as s∈S, and s=[v,d,D] is defined, where v represents motion speed, d represents motion direction, and D represents motion distance. These data will be used as input to the meta-learning model.

[0097] After establishing the "state space," the collected raw dataset {v1,v2,v3...d1,d2,d3...D1,D2,D3...} needs preliminary processing, including data cleaning, formatting, and normalization, to obtain the processed dataset {v1′,v′2,v3′...d1′,d2′,d3′...D1′,D2′,D3′...}, which facilitates subsequent analysis and use. This process can be considered the data preprocessing stage, preparing for the training of the reinforcement learning model. i d represents the operating speed of the high-voltage circuit breaker in the i-th cycle; i D represents the specific direction of movement of the high-voltage circuit breaker in the i-th cycle; iThis represents the actual travel distance of the high-voltage circuit breaker in the i-th cycle; v i ′ represents the velocity of the high-voltage circuit breaker collected after data processing in the i-th cycle; d i ′ represents the specific direction of motion of the high-voltage circuit breaker after data processing in the i-th cycle; D i ′ represents the actual movement distance of the high-voltage circuit breaker after data processing in the i-th cycle;

[0098] The collected data and extracted features are used to train a reinforcement learning model. The action space is usually represented as A. An action vector needs to be defined to capture all possible decisions, hence A = [a, b, c], where a represents the chosen speed measurement algorithm, b represents the algorithm's parameters, and c represents the method of processing the data. In the action space A = [a1, a2... b1, b2... c1, c2...], a1 represents the time difference-based method (using two or more sensors to measure the time difference of an object's passage to calculate speed), a2 represents the frequency-based method (using sensors to measure the frequency of an object's passage to calculate speed); b1 represents the distance between the sensors, b2 represents the time window for measuring the frequency; c1 represents using a filter to remove noise, and c2 represents using a smoother to reduce data fluctuations. During training, the model learns how to extract the most useful information from the data through continuous trial and error and feedback to improve the accuracy and efficiency of speed measurement. This process can be described by the following Bellman equation:

[0099]

[0100] Where Q(v) i A i ) indicates that in state v i Take action A i Action value function, p(v i+1 ',r1∣v i ,a i ) indicates that in state v i Take action a i Then, transition to state v i+1 And the probability of receiving a reward r1, p(v i+1 ',r2∣v i ,b i ) indicates that in state v i Take action b i Then, transition to state v i+1 And the probability of receiving a reward of r2, p(v i+1 ',r3∣v i ,c i ) indicates that in state v i Take action ci Then, transition to state v i+1 'and the probability of obtaining a return of r3; γ1, γ2, and γ3 are their respective discount factors; π(a i+1 '∣v i+1 '), π(b) i+1 '∣v i+1 '), π(c i+1 '∣v i+1 ') indicates that in state v i+1 'The following actions are taken respectively a i+1 '、b i+1 '、c i+1 The strategy; α, β, and λ represent the actual weights corresponding to different actions; Q π (v i+1 ',a i+1 '), Q π (v i+1 ',b i+1 '), Q π (v i+1 ',c i+1 ') respectively represent the next time period in state v i+1 Take action a respectively i+1 '、b i+1 '、c i+1 'Action value function;

[0101] Reinforcement learning models are used to evaluate the quality of the collected data. If the data quality is not high, such as excessive noise or insufficient coverage, the settings of the non-contact sensors or the data collection strategy can be adjusted to improve the data quality.

[0102] Based on the data quality assessment results and feedback from the reinforcement learning model, the data collection strategy is optimized. This optimization process is continuous, aiming to continuously improve the quality and validity of the data to enhance the accuracy and efficiency of speed measurement. This process can be achieved by adjusting the reinforcement learning model's policy π(A). i |v i This is achieved by, given a state v i Next, select the action value function Q(v) that will make the action value function Q(v) work. i A i The largest action A i ;

[0103] The data communication module specifically includes:

[0104] The collected data needs to be encoded and encapsulated for transmission via the communication module. During this process, each state s needs to be associated with its corresponding action p and reward r to facilitate learning and decision-making by the reinforcement learning model. The state space s in time period i is defined. i =[q i ,w i ,e i ], where q i The format representing the data (such as binary, ASCII, etc.), w i Represents the data encapsulation protocol (such as TCP / IP, UDP, etc.), e i The compression method of the data (e.g., no compression, ZIP compression, GZIP compression, etc.) defines the action space p in time period i. i =[t i ,u i ,o i ], where t i This represents changing the data format, u i This represents a change in the data encapsulation protocol, o i This represents a change in the way data is compressed;

[0105] This process can be described by the following formula:

[0106]

[0107] in, Indicates that in state s i Take action P i The action-value function, r is the reward, π t (t i+1 '∣s i+1 '), π u (u i+1 '∣s i+1 '), π o (o i+1 '∣s i+1 ') represent states s and s respectively. i+1 'Take action t' i+1 '、u i+1 '、o i+1 The strategy can be ε-greedy, Softmax, or UCB (Upper Confidence Bound) depending on the actual requirements; α, β, and λ represent the actual weights corresponding to different actions. These represent the states in the next time period, s and s respectively. i+1 Take action t respectively i+1 '、u i+1 '、o i+1 'Action value function;

[0108] The encoded and encapsulated data is sent to the reinforcement learning module via the communication module. During data transmission, it is necessary to ensure data integrity and real-time performance to guarantee that the reinforcement learning model can accurately receive the data and make timely learning and decisions. This process requires continuous updating and optimization of the policy π. p (p i+1 '∣s i+1 '), so that the action value function maximum;

[0109] After the reinforcement learning module receives data, it needs to decode and depackage it to restore the original data format and structure. Then, the reinforcement learning model can use this data for learning and decision-making. In this process, the action-value function needs to be updated based on the newly received data. To reflect new status and reporting information;

[0110] The data processing and meta-learning + reinforcement learning module specifically includes:

[0111] After receiving the data, the meta-learning + reinforcement learning module first performs data parsing to restore the original format and structure of the data. The dataset θ={v1″,v′2′,v′3′...d1″,d2″,d3″...D1″,D′2′,D3″}, which was previously processed by data cleaning, formatting, and normalization, is set as the initial parameters for meta-learning, and the embedded meta-learning module is run.

[0112] For each task i, perform one (or more) gradient descent steps on the model to obtain a new parameter set θ. i If we perform a gradient descent step, the new parameters can be calculated using the following formula:

[0113]

[0114] Where α is the learning rate. η is the gradient of the loss function of model F with respect to task i under parameter set θ. i This is to compensate for the difference between the calculation results of the device module and the actual parameter values ​​required by task i. This value is calculated by statistically analyzing past numerical characteristics.

[0115] Then, we update our initial parameter set θ by minimizing the expected loss after one step of gradient descent on all tasks, which can be achieved by the following formula:

[0116]

[0117] Where β is the meta-learning rate. It is the expected loss of all tasks after one step of gradient descent. η is the compensation between the computational results of the device module and the actual required parameter values. This value is calculated by statistically analyzing past numerical characteristics. If necessary, multiple gradient descent steps can be performed to find a better initial parameter set θ. * The initial parameter set is then transmitted to the outer reinforcement learning module for further processing.

[0118] Based on the better initial parameter set θ obtained after analysis * To further calculate the action value function, this process requires considering the current state parameter set θ. * Given actions g∈[h,j,k], calculate the action-value function Q. π (θ, g) represents the value of taking an action in the current state, where h represents choosing different data processing methods, j represents choosing different learning algorithms, and k represents adjusting the model parameters. This process can be described by the following formula:

[0119]

[0120] in Indicates in parameter set θ * The action value function of the action g is taken; π h (h'∣θ * '), π j (j'∣θ * '), π k (k'∣θ * ') respectively represent the parameters in the parameter set θ * The strategy is to take actions h', j', and k'; γ1, γ2, and γ3 are the corresponding discount factors. They represent the parameters in the set θ respectively. * The action value functions for actions h', j', and k' are respectively taken below;

[0121] According to the new parameter set θ * and action value function The reinforcement learning module updates its policy π. h (h'∣θ * '), π j (j'∣θ * '), π k (k'∣θ * '), so that the action value function This process can be viewed as the learning process of a reinforcement learning model. Through continuous trial and error and feedback, the model will gradually optimize its strategy to improve the accuracy and efficiency of speed measurement.

[0122] According to the updated strategy π(g'∣θ) * The reinforcement learning module will decide on the current parameter set θ. * The action g to be taken next might be adjusting the settings of the non-contact sensor or changing the data collection strategy to improve the quality and validity of the data. During this process, our action value function needs to be updated based on the new action g and the reward r. To reflect new actions and reward information;

[0123] The speed calculation module specifically includes:

[0124] After receiving the data sent by the meta-learning + reinforcement learning module, the data is first parsed to restore its original format and structure. This process is to ensure the integrity and accuracy of the data and to provide accurate input for subsequent calculations.

[0125] Based on the parsed data, the current state and action are identified. During this process, the action value function Q(s,a;θ) and the state value function V(s;w) need to be calculated based on the current state and action. Using the Actor-Critic method, the updated formulas are as follows:

[0126] Critic Update: Update the value function V using the TD error δ.

[0127] δ t =r t +γ*V(s t+1 ;w)-V(s t ;w)

[0128]

[0129] Actor Update: Update policy π using TD error δ.

[0130]

[0131] Based on the strategy π and the action value function Q, the calculation module can calculate the real-time speed of the high-voltage circuit breaker. This process can be viewed as the decision-making process of the meta-learning model. By selecting the action that maximizes the action value function, the optimal speed calculation result can be obtained. This process can be expressed as:

[0132] a t =argmaxaQ(s t ,a;θ)

[0133] This formula indicates that in state s t Next, select the action a that maximizes the action value function Q. t As the optimal action;

[0134] The calculated real-time speed of the high-voltage circuit breaker is fed back to the meta-learning + reinforcement learning module as part of the new state. Then, the meta-learning + reinforcement learning module will adjust the speed based on the new state s. t+1 and returns t+1 Then, update its policy π and action value function Q. The update formula for this process is as follows:

[0135] δ t+1 =r t+1 +γ*V(s t+2 ;w)-V(s t+1 ;w)

[0136]

[0137]

[0138] This formula describes how the meta-learning model adapts to the new state s. t+1 and returns t+1 The method for updating its policy π and action value function Q;

[0139] The meta-learning + reinforcement learning module updates its policy π and action value function Q based on the new state and reward to further optimize the results of speed calculation. This process uses the same update formula as the Actor-Critic method as the update formula of the E4 step, except that the time step of the state and reward is pushed back one step.

[0140] The speed calculation module verifies the calculated speed to ensure its accuracy and reliability. If the verification result does not meet the preset standard, the data processing and learning strategies can be adjusted to improve the accuracy of the speed calculation. In this process, the predictive ability of the meta-learning model can be used to evaluate the accuracy of the speed calculation by comparing the predicted speed with the actual speed.

[0141] If the verification results do not meet the preset standards, the speed calculation module can send a signal to the meta-learning module to indicate that the current speed calculation model needs further optimization. The meta-learning module will then adjust the update strategy for each task based on this signal, thereby improving the speed calculation accuracy of the model. This process can be viewed as finding a better speed calculation strategy in the policy space, and can be represented by the following formula:

[0142]

[0143] Where β is the meta-learning rate, E i [L i (F(θ i ′))] is the expected loss of all tasks after one step of gradient descent;

[0144] In this way, a complete cycle of learning, calculation, feedback, and optimization is completed, enabling the entire system to continuously optimize itself and improve the calculation accuracy and efficiency of the high-voltage circuit breaker speed.

[0145] The result output module specifically includes:

[0146] Before outputting the results, the calculated velocity data needs to be processed, including data formatting and unit conversion, to meet the output requirements. In this process, the decision-making ability of a reinforcement learning model can be used to determine the data processing method by selecting the optimal action. This process can be described by the following formula:

[0147]

[0148] Among them, a * Q represents the optimal action. π (s,a) represents the action-value function for taking action a in state s. This formula describes how to choose the optimal action a given state s. In this case, state s can represent the current data processing requirements (data format, units, precision), action a can represent different data processing methods (data formatting), and the action-value function Q... π (s,a) represents the effect of the data processing method;

[0149] The processed speed data can be displayed on a display device. In this process, the decision-making ability of a reinforcement learning model can be used to determine the display method and format by selecting the optimal action. This can be achieved by displaying the speed value, the speed change curve, or speed statistics. This process can also be described by the formula mentioned above, where state s represents the current display requirement, action a represents different display methods, and the action-value function Q... π (s,a) can represent the effect of adopting a certain display method;

[0150] The processed speed data can also be sent to other devices or systems via the communication module. In this process, the decision-making ability of a reinforcement learning model can be used to determine the communication method and protocol by selecting the optimal action. This can be achieved by choosing between wired or wireless transmission, or by selecting different communication protocols such as TCP / IP, Bluetooth, or Zigbee. This process can also be described by the formula mentioned above, where state s represents the current communication requirement, action a represents different communication methods, and the action-value function Q... π (s,a) can represent the effect of adopting a certain communication method.

[0151] Compared with existing technologies, the above technical solution has the following advantages:

[0152] This invention divides the entire detection process into five steps: "acquisition-processing-meta-learning + reinforcement learning-computation-output," which corresponds to the current research direction of integrated "communication-sensing-computation." This division allows each step to be optimized independently, improving the efficiency and flexibility of the entire detection process. Furthermore, this division makes the entire detection process clearer, easier to understand and operate, thus improving the operability and usability of the detection process.

[0153] This invention employs meta-learning and reinforcement learning algorithms for decision-making and optimization, enabling the entire detection process to proceed automatically without human intervention. This not only significantly improves detection efficiency but also avoids the influence of human factors on the detection results, thereby enhancing accuracy. Furthermore, the application of reinforcement learning and meta-learning algorithms makes the detection process more adaptable and robust, capable of handling various complex detection environments and situations.

[0154] This invention employs a non-contact detection method, avoiding physical damage to high-voltage circuit breakers and improving detection safety. Simultaneously, the non-contact detection method makes the detection process simpler and more efficient, reducing its complexity and difficulty. Furthermore, the use of digital data processing and output improves the usability and readability of the detection results, making them easier to analyze and utilize.

[0155] This invention takes into account the real-time changes in the power grid topology and employs a meta-learning + reinforcement learning algorithm for dynamic optimization. This allows the detection results to better reflect the actual operating status of the high-voltage circuit breaker, improving the practicality of the detection. This dynamic optimization method also makes the detection process more adaptable and flexible, capable of handling various complex and changing detection environments and situations.

[0156] The various sections in this manual are described in a progressive manner, with each section focusing on the differences from the others. Similar or identical parts can be referred to each other.

[0157] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A non-contact high-voltage circuit breaker speed detection device, characterized in that, include: Non-contact sensor deployment module: Non-contact sensors are deployed near the high-voltage circuit breaker to monitor the movement status of the high-voltage circuit breaker. Such sensors can be laser speed sensors. Data collection module: When the high-voltage circuit breaker starts to move, the non-contact sensor detects the movement of the high-voltage circuit breaker and collects relevant data, including the movement speed, direction of movement, and distance of movement of the high-voltage circuit breaker. Data communication module: The collected data is sent to the reinforcement learning module through the communication module. This communication module can be wired or wireless, depending on the specific application scenario and requirements. Data processing and meta-learning + reinforcement learning module: After receiving the data, the meta-learning + reinforcement learning module will process and learn through a pre-set algorithm. During this process, the meta-learning + reinforcement learning module will continuously optimize the speed measurement algorithm to improve the accuracy and efficiency of speed measurement. Speed ​​calculation module: The data processed and learned by the meta-learning + reinforcement learning module is sent to the calculation module, which calculates the real-time speed of the high-voltage circuit breaker based on this data. Result output module: The calculated high-voltage circuit breaker speed can be displayed on a display device or sent to other devices or systems via a communication module; The deployment module for the non-contact sensor specifically includes: Based on the specific conditions of the high-voltage circuit breaker and the speed measurement requirements, select an appropriate non-contact sensor, such as a laser speed sensor. Determine the deployment location of the non-contact sensor so that it can effectively monitor the movement status of the high-voltage circuit breaker. This location should be as close as possible to the high-voltage circuit breaker, but without interfering with its normal movement. At the same time, the safety of the sensor also needs to be considered to avoid potential damage to the sensor when the high-voltage circuit breaker moves. After determining the deployment location, the non-contact sensor is installed at that location. The installation process should ensure the stability of the sensor to prevent the sensor's position from shifting when the high-voltage circuit breaker moves. After installation, the non-contact sensor needs to be tested to ensure that it can work properly and accurately detect the movement status of the high-voltage circuit breaker. The testing process can be carried out by simulating the movement of the high-voltage circuit breaker and observing the sensor's response. The data collection module specifically includes: Non-contact sensors monitor the motion state of high-voltage circuit breakers in real time and collect relevant data, including motion speed, motion direction, and motion distance, which are set as state space S. Each specific state is represented as s∈S, and s=[v,d,D] is defined, where v represents motion speed, d represents motion direction, and D represents motion distance. These data will be used as input to the meta-learning model. After establishing the "state space," the collected raw dataset {v1,v2,v3...d1,d2,d3...D1,D2,D3...} needs preliminary processing, including data cleaning, formatting, and normalization, to obtain the processed dataset {v′1,v′2,v′3...d′1,d′2,d′3...D′1,D′2,D′3...}, which facilitates subsequent analysis and use. This process can be considered the data preprocessing stage, preparing for the training of the reinforcement learning model. i d represents the operating speed of the high-voltage circuit breaker in the i-th cycle; i D represents the specific direction of movement of the high-voltage circuit breaker in the i-th cycle; i v′ represents the actual travel distance of the high-voltage circuit breaker in the i-th cycle. i d represents the velocity collected by the high-voltage circuit breaker after data processing in the i-th cycle; i ′ represents the specific direction of motion of the high-voltage circuit breaker after data processing in the i-th cycle; D i ′ represents the actual movement distance of the high-voltage circuit breaker after data processing in the i-th cycle; The collected data and extracted features are used to train a reinforcement learning model. The action space is usually represented as A. An action vector needs to be defined to capture all possible decisions, hence A = [a, b, c], where a represents the chosen speed measurement algorithm, b represents the algorithm's parameters, and c represents the method of processing the data. In the action space A = [a1, a2... b1, b2... c1, c2...], a1 represents the time difference-based method (using two or more sensors to measure the time difference of an object's passage to calculate speed), a2 represents the frequency-based method (using sensors to measure the frequency of an object's passage to calculate speed); b1 represents the distance between the sensors, b2 represents the time window for measuring the frequency; c1 represents using a filter to remove noise, and c2 represents using a smoother to reduce data fluctuations. During training, the model learns how to extract the most useful information from the data through continuous trial and error and feedback to improve the accuracy and efficiency of speed measurement. This process can be described by the following Bellman equation: Where Q(v) i A i ) indicates that in state v i Take action A i Action value function, p(v i+1 ',r1∣v i ,a i ) indicates that in state v i Take action a i Then, transition to state v i+1 And the probability of receiving a reward r1, p(v i+1 ',r2∣v i ,b i ) indicates that in state v i Take action b i Then, transition to state v i+1 And the probability of receiving a reward of r2, p(v i+1 ',r3∣v i ,c i ) indicates that in state v i Take action c i Then, transition to state v i+1 'and the probability of obtaining a return of r3; γ1, γ2, and γ3 are their respective discount factors; π(a i+1 '∣v i+1 '), π(b) i+1 '∣v i+1 '), π(c i+1 '∣v i+1 ') indicates that in state v i+1 'The following actions are taken respectively a i+1 '、b i+1 '、c i+1 The strategy; α, β, and λ represent the actual weights corresponding to different actions; Q π (v i+1 ',a i+1 '), Q π (v i+1 ',b i+1 '), Q π (v i+1 ',c i+1 ') respectively represent the next time period in state v i+1 Take action a respectively i+1 '、b i+1 '、c i+1 'Action value function; Reinforcement learning models are used to evaluate the quality of the collected data. If the data quality is not high, such as excessive noise or insufficient coverage, the settings of the non-contact sensors or the data collection strategy can be adjusted to improve the data quality. Based on the data quality assessment results and feedback from the reinforcement learning model, the data collection strategy is optimized. This optimization process is continuous, aiming to continuously improve the quality and validity of the data to enhance the accuracy and efficiency of speed measurement. This process can be achieved by adjusting the reinforcement learning model's policy π(A). i |v i This is achieved by, given a state v i Next, select the action value function Q(v) that will make the action value function Q(v) work. i A i The largest action A i ; The data communication module specifically includes: The collected data needs to be encoded and encapsulated for transmission via the communication module. During this process, each state s needs to be associated with its corresponding action p and reward r to facilitate learning and decision-making by the reinforcement learning model. The state space s in time period i is defined. i =[q i ,w i ,e i ], where q i The format representing the data (such as binary, ASCII, etc.), w i Represents the data encapsulation protocol (such as TCP / IP, UDP, etc.), e i The compression method of the data (e.g., no compression, ZIP compression, GZIP compression, etc.) defines the action space p in time period i. i =[t i ,u i ,o i ], where t i This represents changing the data format, u i This represents a change in the data encapsulation protocol, o i This represents a change in the way data is compressed; This process can be described by the following formula: in, Indicates that in state s i Take action P i The action-value function, r is the reward, π t (t i+1 '∣s i+1 '), π u (u i+1 '∣s i+1 '), π o (o i+1 '∣s i+1 ') represent states s and s respectively. i+1 'Take action t' i+1 '、u i+1 '、o i+1 The strategy can be ε-greedy, Softmax, or UCB (Upper Confidence Bound) depending on the actual requirements; α, β, and λ represent the actual weights corresponding to different actions. These represent the states in the next time period, s and s respectively. i+1 Take action t respectively i+1 '、u i+1 '、o i+1 'Action value function; The encoded and encapsulated data is sent to the reinforcement learning module via the communication module. During data transmission, it is necessary to ensure data integrity and real-time performance to guarantee that the reinforcement learning model can accurately receive the data and make timely learning and decisions. This process requires continuous updating and optimization of the policy π. p (p i+1 '∣s i+1 '), so that the action value function maximum; After the reinforcement learning module receives data, it needs to decode and depackage it to restore the original data format and structure. Then, the reinforcement learning model can use this data for learning and decision-making. In this process, the action-value function needs to be updated based on the newly received data. To reflect new status and reporting information; The data processing and meta-learning + reinforcement learning module specifically includes: After receiving the data, the meta-learning + reinforcement learning module first performs data parsing to restore the original format and structure of the data. The dataset θ={v″1,v″2,v″3...d″1,d″2,d″3...D″1,D″2,D″3}, which was previously processed by data cleaning, formatting, normalization, etc., is set as the initial parameters for meta-learning, and the embedded meta-learning module is run. For each task i, perform one (or more) gradient descent steps on the model to obtain a new parameter set θ. i If we perform a gradient descent step, the new parameters can be calculated using the following formula: Where α is the learning rate. η is the gradient of the loss function of model F with respect to task i under parameter set θ. i This is to compensate for the difference between the calculation results of the device module and the actual parameter values ​​required by task i. This value is calculated by statistically analyzing past numerical characteristics. Then, we update our initial parameter set θ by minimizing the expected loss after one step of gradient descent on all tasks, which can be achieved by the following formula: Where β is the meta-learning rate. It is the expected loss of all tasks after one step of gradient descent. η is the compensation between the computational results of the device module and the actual required parameter values. This value is calculated by statistically analyzing past numerical characteristics. If necessary, multiple gradient descent steps can be performed to find a better initial parameter set θ. * The initial parameter set is then transmitted to the outer reinforcement learning module for further processing. Based on the better initial parameter set θ obtained after analysis * To further calculate the action value function, this process requires considering the current state parameter set θ. * Given actions g∈[h,j,k], calculate the action-value function Q. π (θ, g) represents the value of taking an action in the current state, where h represents choosing different data processing methods, j represents choosing different learning algorithms, and k represents adjusting the model parameters. This process can be described by the following formula: in Indicates in parameter set θ * The action value function of the action g is taken; π h (h'∣θ * '), π j (j'∣θ * '), π k (k'∣θ * ') respectively represent the parameters in the parameter set θ * The strategy is to take actions h', j', and k'; γ1, γ2, and γ3 are the corresponding discount factors. They represent the parameters in the set θ respectively. * The action value functions for actions h', j', and k' are respectively taken below; According to the new parameter set θ * and action value function The reinforcement learning module updates its policy π. h (h'∣θ * '), π j (j'∣θ * '), π k (k'∣θ * '), so that the action value function This process can be viewed as the learning process of a reinforcement learning model. Through continuous trial and error and feedback, the model will gradually optimize its strategy to improve the accuracy and efficiency of speed measurement. According to the updated strategy π(g'∣θ) * The reinforcement learning module will decide on the current parameter set θ. * The action g to be taken next might be adjusting the settings of the non-contact sensor or changing the data collection strategy to improve the quality and validity of the data. During this process, our action value function needs to be updated based on the new action g and the reward r. To reflect new actions and reward information; The speed calculation module specifically includes: After receiving the data sent by the meta-learning + reinforcement learning module, the data is first parsed to restore its original format and structure. This process is to ensure the integrity and accuracy of the data and to provide accurate input for subsequent calculations. Based on the parsed data, the current state and action are identified. During this process, the action value function Q(s,a;θ) and the state value function V(s;w) need to be calculated based on the current state and action. Using the Actor-Critic method, the updated formulas are as follows: Critic Update: Update the value function V using the TD error δ. δ t =r t +γ*V(s t+1 ;w)-V(s t ;w) Actor Update: Update policy π using TD error δ. Based on the strategy π and the action value function Q, the calculation module can calculate the real-time speed of the high-voltage circuit breaker. This process can be viewed as the decision-making process of the meta-learning model. By selecting the action that maximizes the action value function, the optimal speed calculation result can be obtained. This process can be expressed as: a t argmaxaQ(s t ,a)θ) This formula indicates that in state s t Next, select the action a that maximizes the action value function Q. t As the optimal action; The calculated real-time speed of the high-voltage circuit breaker is fed back to the meta-learning + reinforcement learning module as part of the new state. Then, the meta-learning + reinforcement learning module will adjust the speed based on the new state s. t+1 and returns t+1 Then, update its policy π and action value function Q. The update formula for this process is as follows: δ t+1 =r t+1 +γ*V(s t+2 ;w)-V(s t+1 ;w) This formula describes how the meta-learning model adapts to the new state s. t+1 and returns t+1 The method for updating its policy π and action value function Q; The meta-learning + reinforcement learning module updates its policy π and action value function Q based on the new state and reward to further optimize the results of speed calculation. This process uses the same update formula as the Actor-Critic method as the update formula of the E4 step, except that the time step of the state and reward is pushed back one step. The speed calculation module verifies the calculated speed to ensure its accuracy and reliability. If the verification result does not meet the preset standard, the data processing and learning strategies can be adjusted to improve the accuracy of the speed calculation. In this process, the predictive ability of the meta-learning model can be used to evaluate the accuracy of the speed calculation by comparing the predicted speed with the actual speed. If the verification results do not meet the preset standards, the speed calculation module can send a signal to the meta-learning module to indicate that the current speed calculation model needs further optimization. The meta-learning module will then adjust the update strategy for each task based on this signal, thereby improving the speed calculation accuracy of the model. This process can be viewed as finding a better speed calculation strategy in the policy space, and can be represented by the following formula: Where β is the meta-learning rate, E i [L i (F(θ′ i ))] is the expected loss of all tasks after one step of gradient descent; In this way, a complete cycle of learning, calculation, feedback, and optimization is completed, enabling the entire system to continuously optimize itself and improve the calculation accuracy and efficiency of the high-voltage circuit breaker speed. The result output module specifically includes: Before outputting the results, the calculated velocity data needs to be processed, including data formatting and unit conversion, to meet the output requirements. In this process, the decision-making ability of a reinforcement learning model can be used to determine the data processing method by selecting the optimal action. This process can be described by the following formula: Among them, a * Q represents the optimal action. π (s,a) represents the action-value function for taking action a in state s. This formula describes how to choose the optimal action a given state s. In this case, state s can represent the current data processing requirements (data format, units, precision), action a can represent different data processing methods (data formatting), and the action-value function Q... π (s,a) represents the effect of the data processing method; The processed speed data can be displayed on a display device. In this process, the decision-making ability of a reinforcement learning model can be used to determine the display method and format by selecting the optimal action. This can be achieved by displaying the speed value, the speed change curve, or speed statistics. This process can also be described by the formula mentioned above, where state s represents the current display requirement, action a represents different display methods, and the action-value function Q... π (s,a) can represent the effect of adopting a certain display method; The processed speed data can also be sent to other devices or systems via the communication module. In this process, the decision-making ability of a reinforcement learning model can be used to determine the communication method and protocol by selecting the optimal action. This can be achieved by choosing between wired or wireless transmission, or by selecting different communication protocols such as TCP / IP, Bluetooth, or Zigbee. This process can also be described by the formula mentioned above, where state s represents the current communication requirement, action a represents different communication methods, and the action-value function Q... π (s,a) can represent the effect of adopting a certain communication method.