Multi-parameter collaborative optimization method for mechatronic coupling system based on grpo reinforcement learning algorithm

By using the GRPO reinforcement learning algorithm to monitor and optimize multiple parameters of electromechanical coupling systems in real time, the problem of handling nonlinear coupling relationships in traditional methods is solved, the stability and power quality of the system are improved, and the operating performance of the electromechanical coupling system is significantly enhanced.

CN120930511BActive Publication Date: 2026-01-13STATE GRID TIANJIN ELECTRIC POWER CO BINHAI POWER SUPPLY BRANCH +2
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511452893.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-13
Publication Date
2026-01-13
Estimated Expiration
2045-10-13

AI Technical Summary

Technical Problem

Traditional electromechanical coupling system control methods are difficult to effectively handle the nonlinear coupling relationship between multiple parameters and the dynamic changes in system operating conditions, resulting in limited system optimization effects and difficulty in achieving optimal stability and power quality.

Method used

The GRPO reinforcement learning algorithm is used to monitor multiple operating parameters of the electromechanical coupling system in real time. Through online learning and decision-making capabilities, the system parameters are intelligently adjusted to achieve dynamic optimization. A state vector and action space are constructed, and a reward function is designed to maximize long-term cumulative reward. Offline pre-training and online optimization are performed.

Benefits of technology

It improves the operational stability and power quality of electromechanical coupling systems, suppresses system vibration, stabilizes temperature and voltage, reduces harmonic content, and enhances overall operational performance and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120930511B_ABST
    Figure CN120930511B_ABST
Patent Text Reader

Abstract

The present disclosure belongs to the technical field of live-line work robots of power distribution networks, and provides a multi-parameter collaborative optimization method for an electromechanical coupling system based on a GRPO reinforcement learning algorithm, which comprises the following steps: acquiring multi-parameter operation data of the electromechanical coupling system during operation; pre-processing the acquired operation data; constructing a state vector based on the pre-processed operation data; constructing a GRPO reinforcement learning collaborative optimization model based on the GRPO reinforcement learning algorithm; inputting the state vector into the GRPO reinforcement learning collaborative optimization model; and the GRPO reinforcement learning collaborative optimization model obtains an optimization strategy through interactive learning with the electromechanical coupling system. The framework based on the reinforcement learning of the present disclosure enables the optimization strategy to adapt to changes in system operation conditions, environmental disturbances and parameter drift caused by component aging in real time, and maintains the robustness of the optimization effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure belongs to the field of live-line working robot technology for power distribution networks, specifically involving a multi-parameter collaborative optimization method, system, electronic device, storage medium, and program product for electromechanical coupling systems based on the GRPO reinforcement learning algorithm. Background Technology

[0002] The core energy transmission device of a live-line working robot is an electromechanical coupling system consisting of a motor, an insulated transmission rod, and a generator, used to achieve critical power transmission and energy conversion. The operating performance of this type of electromechanical coupling system, especially its dynamic stability and the quality of its output power, is affected by multiple mutually coupled parameters (such as vibration, temperature, voltage, current, and rotational speed).

[0003] Traditional electromechanical coupling system control methods are usually based on simplified models or PID controllers, which are difficult to effectively handle the nonlinear coupling relationship between multiple parameters and the dynamic changes in system operating conditions. This results in limited system optimization effects and makes it difficult to achieve optimal stability and power quality.

[0004] With the increasing demands on live-line working robots, a significant technical challenge has emerged: how to intelligently and collaboratively adjust system parameters based on real-time monitoring data to achieve the optimal operating state of the entire electromechanical coupling system. For example, existing intelligent optimization methods, such as model-based optimization or traditional machine learning methods, may require precise system models or large amounts of labeled data, and their online adaptability and real-time decision-making capabilities need improvement. Summary of the Invention

[0005] To address the aforementioned issues, this disclosure provides a multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO (Group Relative Policy Optimization) reinforcement learning algorithm. By real-time monitoring of multiple operating parameters of the electromechanical coupling system, and utilizing the online learning and decision-making capabilities of the GRPO reinforcement learning algorithm, the control strategy of the prime mover is intelligently adjusted to achieve dynamic optimization of the operating state of the entire electromechanical coupling system, thereby improving the system's operational stability and the quality of generated power.

[0006] In a first aspect, embodiments of this disclosure provide a multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm, including:

[0007] Acquire multi-parameter operating data of the electromechanical coupling system during operation;

[0008] The acquired operational data is preprocessed; a state vector is constructed based on the preprocessed operational data.

[0009] Construct a GRPO reinforcement learning collaborative optimization model based on the GRPO reinforcement learning algorithm;

[0010] The state vector is input into the GRPO reinforcement learning collaborative optimization model; the GRPO reinforcement learning collaborative optimization model obtains the optimization strategy through interactive learning with the electromechanical coupling system.

[0011] Furthermore,

[0012] The electromechanical coupling system is specifically constructed as follows: an AC asynchronous motor driven by a frequency converter serves as the power source, which is transmitted through a transmission rod with predetermined insulation properties to drive a synchronous generator to generate electricity.

[0013] Furthermore,

[0014] The constructed state vector includes:

[0015] Based on the preprocessed multi-parameter runtime data, key parameters are selected to construct a state vector for comprehensive characterization. t The operating status of the electromechanical coupling system at all times.

[0016] Furthermore,

[0017] A GRPO reinforcement learning collaborative optimization model is constructed based on the GRPO reinforcement learning algorithm, including:

[0018] The state space is designed to treat the electromechanical coupling system and its operational dynamics as a reinforcement learning environment.

[0019] Define the actions that the GRPO reinforcement learning algorithm can perform. The set is the action space;

[0020] Design a reward function to quantify the actions performed. The system then transitions from state Transferred to The resulting immediate optimization effect;

[0021] Maximizing the expected value of long-term cumulative reward is used as the objective function of the GRPO reinforcement learning collaborative optimization model.

[0022] Furthermore,

[0023] The set of state vectors in the time-frequency domain constitutes the state space.

[0024] Furthermore,

[0025] Actions It is defined as the set of minute adjustments to the target speed of the motor, the set of adjustments to the torque limit, and the adjustment of the gain parameter of the directly output PID controller.

[0026] Furthermore,

[0027] action The form can be discrete or continuous.

[0028] Furthermore,

[0029] In the discrete action space A set of preset adjustment instructions; in the continuous motion space, It is a vector.

[0030] Furthermore,

[0031] The objective function is:

[0032] ,

[0033] in: π It is a strategy of the collaborative optimization model; It is the expected value; γ It is a discount factor between 0 and 1, representing the importance of future rewards relative to current rewards; The above reward function represents the state... Execute action Later transitioned to a new state Instant rewards received at that time.

[0034] Furthermore,

[0035] The learning process of the GRPO reinforcement learning collaborative optimization model includes offline pre-training and online optimization.

[0036] Furthermore,

[0037] The specific iterative process of online optimization is as follows:

[0038] The current state of the electromechanical coupling system is obtained through sensor data and preprocessing. ;

[0039] The GRPO reinforcement learning collaborative optimization model is based on the current policy. Output Action ;

[0040] Actions It is converted into a control signal for the motor controller and executed;

[0041] Calculated instant reward ;

[0042] Update the policy network parameters of the GRPO reinforcement learning collaborative optimization model. .

[0043] Furthermore,

[0044] The multi-parameter collaborative optimization method further includes: adjusting the input and output parameters of the motor based on the optimization strategy learned online by the GRPO reinforcement learning collaborative optimization model.

[0045] Secondly, this disclosure also provides a multi-parameter collaborative optimization system for electromechanical coupling systems based on the GRPO reinforcement learning algorithm, including: a data acquisition module, a state vector construction module, a model construction module, and a decision module;

[0046] The data acquisition module is used to acquire multi-parameter operating data of the electromechanical coupling system during operation;

[0047] A state vector construction module is used to preprocess the acquired running data and construct a state vector based on the preprocessed running data.

[0048] The model building module is used to build GRPO reinforcement learning collaborative optimization models based on the GRPO reinforcement learning algorithm.

[0049] The decision module is used to input the state vector into the GRPO reinforcement learning collaborative optimization model; the GRPO reinforcement learning collaborative optimization model obtains the optimization strategy through interactive learning with the electromechanical coupling system.

[0050] Thirdly, based on the same inventive concept, this disclosure also provides an electronic device, including at least one processor and at least one memory electrically connected;

[0051] The memory is electrically connected to the processor, wherein the memory stores instructions that can be executed by at least one of the processors, the instructions being executed by at least one of the processors to enable at least one of the processors to execute the aforementioned multi-parameter collaborative optimization method for electromechanical coupled systems based on the GRPO reinforcement learning algorithm.

[0052] Fourthly, based on the same inventive concept, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program;

[0053] When the computer program is executed by the processor, it describes the aforementioned multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm.

[0054] Fifthly, based on the same inventive concept, this disclosure also provides a computer program product.

[0055] The computer program product is stored in at least one storage medium;

[0056] The computer program product includes several instructions to cause at least one electronic device to execute the aforementioned multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm.

[0057] Compared with the prior art, this disclosure has the following advantages:

[0058] 1. Based on the framework of reinforcement learning, especially by utilizing online learning mechanisms, the optimization strategy can adapt in real time to changes in system operating conditions, environmental interference, and parameter drift caused by component aging, thus maintaining the robustness of the optimization effect.

[0059] 2. Through intelligent and precise control of the motor, system vibration can be effectively suppressed, operating temperature can be stabilized, the stability of generator output voltage and frequency can be improved, and harmonic content can be reduced, thereby significantly improving the overall operating performance and reliability of the electromechanical coupling system.

[0060] Other features and advantages of this disclosure will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the disclosure. The objects and other advantages of this disclosure may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description

[0061] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0062] Figure 1 A flowchart illustrating a multi-parameter collaborative optimization method for an electromechanical coupling system based on the GRPO reinforcement learning algorithm according to an embodiment of the present disclosure is shown.

[0063] Figure 2 A schematic diagram of an online optimization loop for GRPO reinforcement learning according to an embodiment of the present disclosure is shown;

[0064] Figure 3 A schematic diagram of an electronic device structure according to an embodiment of the present disclosure is shown. Detailed Implementation

[0065] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0066] Reinforcement learning (RL) learns decision-making strategies autonomously through environmental interaction and trial-and-error mechanisms, making it particularly suitable for solving online decision-making and optimization problems in complex dynamic systems.

[0067] Group Relative Policy Optimization (GRPO), as an advanced reinforcement learning algorithm, can handle high-dimensional state and action spaces, effectively learn and utilize complex nonlinear coupling relationships between multiple parameters, and achieve synergistic optimization of multiple objectives such as system stability and power quality, rather than the single-parameter or decoupled control of traditional methods. Through the relative optimization mechanism of policy groups, it shows potential in reducing computational resource requirements and improving the efficiency of learning complex patterns, making it particularly suitable for handling complex system optimization tasks with multidimensional state and action spaces.

[0068] Applying the GRPO reinforcement learning algorithm to the multi-parameter collaborative optimization of electromechanical coupling systems can overcome the limitations of traditional methods and achieve more efficient and intelligent control of system operation status.

[0069] The electromechanical coupling system of this disclosure includes an electric motor, an insulated transmission rod, and a generator, used to provide power and drive for a live-line working robot in a power distribution network.

[0070] The electromechanical coupling system of this embodiment is specifically constructed as follows: an AC asynchronous motor driven by a frequency converter serves as a power source, which is transmitted through a transmission rod with predetermined insulation properties to drive a synchronous generator to generate electricity.

[0071] To achieve comprehensive monitoring of the operating status of the electromechanical coupling system, a sensor network needs to be deployed at key locations within the system. Specifically: on the generator side, voltage and current transformers, as well as frequency and waveform monitoring equipment, should be installed, supplemented by temperature sensor one to monitor temperature rise in key areas; on the insulated transmission rod, acceleration sensors should be installed at key locations to monitor vibration, and temperature sensor two should be installed to monitor surface temperature; on the motor side, corresponding sensors need to be installed to monitor its actual speed, output torque, stator current, and temperature, etc.

[0072] The embodiments of this disclosure select the GRPO reinforcement learning algorithm as the core optimization controller, mainly utilizing its ability to handle complex dynamic systems, multivariable coupling, and online adaptive optimization problems.

[0073] Figure 1 A flowchart illustrating a multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm according to an embodiment of this disclosure is shown, as follows: Figure 1 As shown in this embodiment, the multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm includes:

[0074] S1, acquire multi-parameter operating data of the electromechanical coupling system during operation.

[0075] The operating data includes parameters reflecting the state of the generator and insulated transmission rod, such as vibration, temperature, generator output voltage, and generator output current, as well as parameters reflecting the state of the motor, such as motor speed and torque.

[0076] S2, preprocess the acquired running data; construct a state vector based on the preprocessed running data.

[0077] S21, the preprocessing includes data cleaning, signal filtering, feature extraction, normalization, and standardization.

[0078] Feature extraction involves obtaining more characterizing parameters of a system's deeper state from raw time-domain data. Taking vibration signals as an example, the root mean square (RMS), peak value, and kurtosis factor can be calculated, or the main frequency components and their amplitudes can be extracted through Fast Fourier Transform (FFT) analysis. For voltage and current signals, the effective value, frequency deviation, and total harmonic distortion (THD) are calculated. Parameters such as temperature, rotational speed, and torque can be used directly or smoothed.

[0079] Normalization or standardization is used to eliminate the influence of different physical dimensions on the subsequent neural network processing in the GRPO model.

[0080] S22, Construct the state vector .

[0081] Based on the preprocessed multi-parameter runtime data, key parameters are selected to construct a state vector. Used for comprehensive characterization t The operating state of the electromechanical coupling system at all times, i.e. express t The operating state of a constant electromechanical coupling system, or simply its state. .

[0082] In this embodiment:

[0083] ,

[0084] in: This is the effective value of the generator voltage. For frequency deviation, For total harmonic distortion of current, For generator temperature, The RMS value of the transmission rod vibration. For the vibration amplitude at a specific frequency, For the temperature of the transmission rod, This refers to the motor speed. This represents the motor torque.

[0085] S3, Constructing a GRPO reinforcement learning collaborative optimization model based on the GRPO reinforcement learning algorithm.

[0086] S31, Design state space.

[0087] The electromechanical coupling system and its operational dynamics are considered as the environment for reinforcement learning. The preprocessed multidimensional feature data forms the basis for the GRPO reinforcement learning collaborative optimization model to perceive the environment.

[0088] Specifically, data collected from vibration sensors, temperature sensors, voltage transformers, current transformers, and motor encoders are extracted in the time and frequency domain and combined into a multi-dimensional vector set to dynamically reflect the mechanical vibration state, thermal state, electrical output characteristics, and potential anomalies of the transmission components of the electromechanical coupling system.

[0089] In other words, the state space ( S ), which is the set of state vectors in the instantaneous frequency domain.

[0090] S32, designing motion space ( A ).

[0091] Parameterizing the control commands of the motor controller, including the actions... Defined as the set of minute adjustments to the target speed of the electric motor. A set of torque limit adjustment amounts And the gain parameter adjustment amount of the PID controller is directly output. Among them, Indicates the target rotational speed. Indicates the adjustment factor. This indicates torque limitation.

[0092] action The specific control adjustment amount applied to the motor includes: target speed change. Δω Torque change ΔT Or actions This corresponds to adjusting specific parameters of the motor controller, including PID parameters.

[0093] action The form can be discrete or continuous.

[0094] If a discrete action space is used, the actions can be... Defined as a set of preset adjustment commands, such as {hold, slight increase in speed, slight decrease in speed, slight increase in torque limit, slight decrease in torque limit};

[0095] If a continuous motion space is used, then the motion... It can be a vector, whose elements directly correspond to the target rotational speed. or its change Δω and output torque limit or torque change ΔT The specific value.

[0096] S33, Design the reward function.

[0097] reward function ( R The design of ) is the core of guided optimization, used to quantify the actions to be performed. The system then transitions from state Transferred to The resulting immediate optimization effect.

[0098] To achieve multi-objective collaborative optimization, the reward function is often designed as a weighted combination of multiple indicators.

[0099] In this embodiment:

[0100] ,

[0101] in, It is a reward component based on whether the vibration level and temperature are within a safe range; It is a bonus component based on the generator's output voltage stability, frequency stability, and harmonic content; It is to perform an action Potential penalties for control costs or adjustments; , , These are the weight coefficients for each reward and penalty item, and priorities are set according to specific optimization objectives.

[0102] Positive rewards are given for improving or maintaining the system's state within the excellent range, while negative rewards (penalties) are given for deteriorating the state or deviating from the target range.

[0103] S34, Construct the objective function for the GRPO reinforcement learning collaborative optimization model.

[0104] The objective function aims to maximize the expected value of the long-term cumulative reward.

[0105] ,

[0106] in: π It is a strategy of the collaborative optimization model; It is the expected value; γ It is a discount factor between 0 and 1, representing the importance of future rewards relative to current rewards; The above reward function represents the state... Execute action Later transitioned to a new state Instant rewards received at that time.

[0107] S4, the state vector is input into the GRPO reinforcement learning collaborative optimization model; the GRPO reinforcement learning collaborative optimization model obtains the optimization strategy through interactive learning with the electromechanical coupling system.

[0108] The learning process of the GRPO reinforcement learning collaborative optimization model includes offline pre-training and online optimization.

[0109] S41, In the offline phase, preliminary training can be performed using simulation models or historical data.

[0110] S42, the online optimization phase is even more critical. The GRPO reinforcement learning collaborative optimization model continuously interacts with the electromechanical coupling system during actual operation. For example... Figure 2 As shown, the specific iterative process of online optimization is as follows:

[0111] S421 obtains the current state of the electromechanical coupling system through sensor data and preprocessing. .

[0112] S422, the GRPO reinforcement learning collaborative optimization model, based on the current policy... Output Action .

[0113] S423, the action It is converted into a control signal for the motor controller and executed.

[0114] S424, the electromechanical coupling system transitions to a new state. And obtain sensor data in that state.

[0115] S425, Calculate the immediate reward obtained. .

[0116] S426, Update the policy network parameters of the GRPO reinforcement learning collaborative optimization model. .

[0117] This online learning process enables the GRPO reinforcement learning collaborative optimization model to continuously adapt to environmental changes and model drift, constantly optimizing its control strategy to maximize long-term cumulative rewards, thereby achieving the goal of collaboratively optimizing system stability and power quality.

[0118] Finally, the input and output parameters of the motor are adjusted based on the optimization strategy learned online by the GRPO reinforcement learning collaborative optimization model.

[0119] The GRPO reinforcement learning collaborative optimization model obtains optimization strategies through interactive learning with the electromechanical coupling system, which is illustrated through the following two specific embodiments.

[0120] Specific Implementation Example 1: Coordinated Optimization of Vibration Suppression and Voltage Stability.

[0121] In the scenario of Example 1, assuming that during the operation of the electromechanical coupling system, the sensor detects a significant increase in the vibration intensity at a key measuring point of the insulated transmission rod, which is reflected in the state vector. The root mean square value of vibration in And vibration amplitude at a specific frequency (a certain natural frequency of the system) All exceeded the normal threshold. At the same time, the effective value of the generator output voltage... and frequency deviation While power quality indicators are still within acceptable ranges, slight fluctuations may have occurred, indicating that system stability is threatened.

[0122] Faced with this situation The GRPO reinforcement learning collaborative optimization model relies on its internal policy network formed through online learning and offline pre-training. It makes decisions. The policy network has learned the mapping relationship between different states and optimal actions. It evaluates the potential long-term rewards of various possible actions in the current state. Action space ( A This can include adjusting the target speed of the motor. Adjusting torque limit Multiple options are available.

[0123] Based on the learned strategy, the GRPO reinforcement learning co-optimization model can infer that the current high vibration is likely due to the operating speed being close to the system's resonance point. Therefore, a small speed adjustment, for example, performing an action... or , corresponding to a specific This value has a high probability of shifting the system's operating point out of the resonant region, thereby effectively reducing vibration. Simultaneously, the policy network also considers the potential impact of this action on power quality. If the learned policy indicates that a small speed adjustment typically does not cause a significant deviation of voltage or frequency from the target value, the collaborative optimization model ultimately chooses to execute the speed adjustment action. .

[0124] Next, the optimal control action command is converted into an action. The specific control signal is sent to the motor controller, and the motor responds to the command by adjusting its speed. The system then transitions to the new state. At this point, the sensor re-acquires and processes the data. If the new state... show: and A significant decrease, returning to normal or acceptable levels; at the same time... and It remains stable, or recovers after only minor, brief fluctuations.

[0125] At this point, calculate the immediate reward. Based on the aforementioned reward function design Due to the significant improvement in vibration, This item will contribute a large positive value; due to the good power quality, The item also contributes positive value or is close to zero; action cost (depending on This is typically a small negative value. Overall, instant rewards... r t It will be a significantly positive number.

[0126] The GRPO reinforcement learning algorithm utilizes this positive reward r t This is to update its internal model parameters. The experience gained from this success... It will enhance the policy network In the process, the action of fine-tuning the rotational speed is selected under conditions similar to high vibration. The probability of this. Through continuous accumulation of such experience and updates, the coordinated control capabilities of electromechanical coupling systems for vibration suppression and voltage stabilization have been continuously improved.

[0127] Specific Implementation Example 2: In-depth Analysis of the Coordinated Control of High Temperature and Harmonics.

[0128] In the scenario of Example 2, the system state This manifests as: temperature of key generator components (such as windings) The total harmonic distortion of the generator output current remains consistently above the preset safety threshold, and is also monitored. THD IThis also exceeds the limits allowed by power quality standards. Other parameters such as vibration, voltage, and frequency are still within the normal range.

[0129] The GRPO reinforcement learning collaborative optimization model receives data containing high... and high state vector Its policy network Begin evaluating optional actions. The action space can include adjusting speed, adjusting torque limits, etc.

[0130] Based on the strategies it has learned, the collaborative optimization model may recognize that reducing the motor's output torque corresponds to the following action: That is, a specific A negative value can directly reduce the generator load, thereby reducing its heat generation and helping to alleviate high-temperature problems. Simultaneously, changing the operating point of the motor and generator (changes in load rate) may also affect the system's nonlinear characteristics, potentially improving harmonic levels.

[0131] After comprehensive evaluation by the strategy network, if it is determined that, compared to other actions, such as adjusting the engine speed, which may have uncertain harmonic effects or slow temperature effects, reducing the torque limit is the optimal choice to obtain a higher expected cumulative reward under the current state, then that action is output. .

[0132] Next, the optimal control action command is converted. a t The specific control signal is sent to the motor controller, limiting the motor's output torque and causing a reduction in the generator load. The system enters a new state. Ideally, the new status would display: It has begun to show a downward trend; The value has also decreased.

[0133] Calculate the reward at this time Due to temperature decline, Positive rewards for contributions; due to harmonics reduce, It also contributes positive rewards. If the positive contributions of these two items significantly exceed the action cost... The total reward It remains positive.

[0134] The GRPO reinforcement learning algorithm utilizes this positive reward This reinforces the corresponding strategies. This successful interaction makes the co-optimization model more inclined to adopt a control strategy that appropriately reduces torque limits when encountering conditions of both high temperature and high harmonics in the future. Through continuous learning, the co-optimization model can learn how to coordinately manage temperature and power quality by adjusting load and other means while meeting basic power generation needs.

[0135] In this embodiment of the disclosure, the GRPO reinforcement learning collaborative optimization model adopts an offline pre-training and online learning mechanism, which allows the model to be fully trained based on simulation data or historical data before deployment, and to continuously optimize the strategy online based on the actual operating effect after deployment. During the actual operation of the electromechanical coupling system, the strategy of the GRPO reinforcement learning collaborative optimization model is continuously updated and fine-tuned online using newly collected data and obtained reward feedback to adapt to the slow drift of system parameters and changes in real-time operating conditions.

[0136] Based on the above method, this disclosure also provides a multi-parameter collaborative optimization system for electromechanical coupling systems based on the GRPO reinforcement learning algorithm, corresponding to the above method, including a data acquisition module, a state vector construction module, a model construction module, and a decision module;

[0137] The data acquisition module is used to acquire multi-parameter operating data of the electromechanical coupling system during operation;

[0138] A state vector construction module is used to preprocess the acquired running data and construct a state vector based on the preprocessed running data.

[0139] The model building module is used to build GRPO reinforcement learning collaborative optimization models based on the GRPO reinforcement learning algorithm.

[0140] The decision module is used to input the state vector into the GRPO reinforcement learning collaborative optimization model; the GRPO reinforcement learning collaborative optimization model obtains the optimization strategy through interactive learning with the electromechanical coupling system.

[0141] Based on the same inventive concept as the above-disclosed content, this disclosure also provides an electronic device. For example... Figure 3 As shown, the electronic device of this disclosure includes at least one processor and at least one memory electrically connected to each other. The memory is electrically connected to the processor, wherein the memory stores instructions executable by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm as described above.

[0142] It should be noted that the electrical connection between the above-mentioned units does not necessarily mean the connection between lines. The indirect connection method can be applied to the embodiments of this disclosure as long as it achieves the purpose of this disclosure.

[0143] Based on the same inventive concept, this disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm as described above.

[0144] Based on the same inventive concept, this disclosure also provides a computer program product, which is stored in at least one storage medium; the computer program product includes several instructions to cause at least one computer device to execute the multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm as described above.

[0145] Although the present disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present disclosure.

Claims

1. A multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm, characterized in that, The method includes, Acquire multi-parameter operating data of the electromechanical coupling system during operation; The acquired runtime data is preprocessed; constructing a state vector based on the pre-processed operating data , representing t the operating state of the electromechanical coupling system at the instant in time, in short the state, , wherein: is the generator voltage RMS value, is the frequency deviation, is the current total harmonic distortion, is the generator temperature, is the drive rod vibration RMS value, is the specific frequency vibration amplitude, is the drive rod temperature, is the motor speed, is the motor torque; Constructing a GRPO reinforcement learning collaborative optimization model based on the GRPO reinforcement learning algorithm includes: designing the state space, treating the electromechanical coupling system and its operational dynamics as the reinforcement learning environment; and defining the actions that the GRPO reinforcement learning algorithm can execute. The set of actions forms the action space, which contains the actions. Defined as the set of minute adjustments to the target speed of the electric motor. A set of torque limit adjustment amounts And the gain parameter adjustment amount of the PID controller is directly output. wherein, represents a target rotational speed, represents an adjustment coefficient, represents a torque limit; design reward function R for quantifying the immediate optimization effect brought by the system after transitioning from a state to to after performing an action , wherein, is a reward component based on the vibration level, whether the temperature is in the safe interval; is a reward component based on the generator output voltage stability, frequency stability and harmonic content; is a penalty term of the control cost or adjustment range that the execution of the action may bring; are weight coefficients of each reward and penalty term, and priority is set according to specific optimization objectives;​​ Maximizing the expected value of long-term cumulative reward is used as the objective function of the GRPO reinforcement learning collaborative optimization model; The state vector is input into the GRPO reinforcement learning collaborative optimization model; the GRPO reinforcement learning collaborative optimization model obtains the optimization strategy through interactive learning with the electromechanical coupling system.

2. The multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm according to claim 1, characterized in that, The electromechanical coupling system is specifically constructed as follows: an AC asynchronous motor driven by a frequency converter serves as the power source, which is transmitted through a transmission rod with predetermined insulation properties to drive a synchronous generator to generate electricity.

3. The multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm according to claim 1, characterized in that, The constructed state vector includes: Based on the pretreated multi-parameter operation data, key parameters are selected to form a state vector for comprehensively representing t The running state of the electromechanical coupling system.

4. The multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm according to claim 1, characterized in that, The set of state vectors in the time-frequency domain constitutes the state space.

5. The multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm according to claim 1, characterized in that, Action The action is defined as a set of small adjustment amounts for the target rotation speed of the motor, a set of adjustment amounts for the torque limit, and an adjustment amount for the gain parameter of the direct output PID controller.

6. The multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm according to any one of claims 1-5, characterized in that, Actions The form is discrete or continuous.

7. The multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm according to any one of claims 1-5, characterized in that, In discrete action space, action is a vector. In continuous action space, action is a vector.

8. The multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm according to claim 1, characterized in that, The objective function is: , in: π It is a strategy of the collaborative optimization model; It is the expected value; γ It is a discount factor between 0 and 1, representing the importance of future rewards relative to current rewards; The above reward function represents the state... Execute action Later transitioned to a new state Instant rewards received at that time.

9. The multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm according to claim 1, characterized in that, The learning process of the GRPO reinforcement learning collaborative optimization model includes offline pre-training and online optimization.

10. The multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm according to claim 9, characterized in that, The specific iterative process of online optimization is as follows: By sensor data and pre-processing, the current electromechanical coupling system state is obtained ; GRPO reinforcement learning collaborative optimization model according to current policy output action ; The action is converted into a control signal for the motor controller and executed; Computing the obtained instant prize ; Updating policy network parameters of a grpo reinforcement learning co-optimization model .

11. The multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm according to claim 1, characterized in that, The multi-parameter collaborative optimization method further includes: adjusting the input and output parameters of the motor based on the optimization strategy learned online by the GRPO reinforcement learning collaborative optimization model.

12. A multi-parameter collaborative optimization system for electromechanical coupling systems based on the GRPO reinforcement learning algorithm, characterized in that, The system includes a data acquisition module, a state vector construction module, a model construction module, and a decision-making module; The data acquisition module is used to acquire multi-parameter operating data of the electromechanical coupling system during operation; a state vector construction module, configured to preprocess the acquired operation data, and construct a state vector based on the preprocessed operation data , representing t an operation state of the electromechanical coupling system at the moment, referred to as a state, , wherein: Vgen is the generator voltage RMS value, fdev is the frequency deviation, Ith is the current total harmonic distortion, Tgen is the generator temperature, RMS is the drive rod vibration RMS value, Vfreq is the vibration amplitude at a specific frequency, Trod is the drive rod temperature, N is the motor speed, T is the motor torque; The model building module is used to construct a GRPO reinforcement learning co-optimization model based on the GRPO reinforcement learning algorithm. This includes: designing the state space, treating the electromechanical coupling system and its dynamics as the reinforcement learning environment; and defining the actions that the GRPO reinforcement learning algorithm can execute. The set of actions forms the action space, which contains the actions. Defined as the set of minute adjustments to the target speed of the electric motor. A set of torque limit adjustment amounts And the gain parameter adjustment amount of the PID controller is directly output. wherein, denotes a target rotational speed, denotes an adjustment coefficient, denotes a torque limit; design reward function R for quantifying the immediate optimization effect brought by the system after performing an action from a state to a state , in, It is a reward component based on whether the vibration level and temperature are within a safe range; It is a bonus component based on the generator's output voltage stability, frequency stability, and harmonic content; It is to perform an action Potential penalties for control costs or adjustments; , , These are the weight coefficients for each reward and penalty item, and priorities are set according to specific optimization objectives; Maximizing the expected value of long-term cumulative reward is used as the objective function of the GRPO reinforcement learning collaborative optimization model; The decision module is used to input the state vector into the GRPO reinforcement learning collaborative optimization model; the GRPO reinforcement learning collaborative optimization model obtains the optimization strategy through interactive learning with the electromechanical coupling system.

13. An electronic device, characterized in that, Includes at least one processor and at least one memory electrically connected; The memory is electrically connected to the processor, wherein the memory stores instructions executable by at least one of the processors, the instructions being executed by at least one of the processors to enable at least one of the processors to perform the multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm as described in any one of claims 1-11.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program. When the computer program is executed by the processor, it implements the multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm as described in any one of claims 1-11.

15. A computer program product, characterized in that, The computer program product is stored in at least one storage medium; The computer program product includes several instructions to cause at least one electronic device to execute the multi-parameter collaborative optimization method for electromechanical coupling systems based on the GRPO reinforcement learning algorithm as described in any one of claims 1-11.

Citation Information

Patent Citations

  • Multi-machine navigation method and system capable of keeping connectivity and medium

    CN114706384A

  • Air conditioner energy-saving control method and system based on reinforcement learning and digital twinborn model

    CN118031385A