Planetary gear system digital twin model updating method based on deep reinforcement learning

The digital twin model of the planetary gear system is constructed through deep reinforcement learning methods. The vibration signal frequency domain analysis and rigid-flexible coupling model are used to solve the problem of long calculation time and the inability to retain search experience in the existing technology, and the rapid parameter inversion and model update of the planetary gear system are realized.

CN120449642APending Publication Date: 2025-08-08SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510435229.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing digital twin model construction and update methods of planetary gear system have the problem of long calculation time and the inability to retain search experience, which is difficult to meet the real-time needs. Especially in the complex operating conditions of planetary gear systems, existing methods such as particle swarm algorithms need to be re-searched, resulting in a long time.

Method used

The planetary gear system digital twin model is constructed by deep reinforcement learning method. Through the rigid-flexible coupling model and vibration signal frequency domain analysis, the deep reinforcement learning algorithm is used to perform parameter inversion, and combined with the first third-order meshing frequency multiplication of the vibration signal frequency domain and its edge frequency amplitude as optimization goals, a reinforcement learning reward function is constructed, and the learning experience is stored in the target value network to achieve rapid optimization of parameters.

Benefits of technology

On the premise of ensuring accuracy, the parameter inversion time is significantly reduced, from hour to minute level, improving parameter inversion efficiency, and is suitable for rapid model updates of planetary gear systems and parameter optimization under similar operating conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120449642A_ABST
    Figure CN120449642A_ABST
Patent Text Reader

Abstract

The invention discloses a planetary gear system digital twin model updating method based on deep reinforcement learning. The planetary gear system digital twin model updating method comprises the steps of S1, simplifying a planetary gear system; s2, describing the kinetic model by using a lumped parameter method, and constructing a planetary gear train oscillatory differential equation set; s3, a frequency response function family is calculated through finite element analysis, and a planetary gearbox rigid-flexible coupling model, namely a digital twinborn body, is constructed; s4, vibration response of the planetary gearbox shell under the stable rotating speed is obtained through a vibration signal collecting device; s5, defining an action space and a state space, and establishing a reinforcement learning environment; s6, solving the oscillatory differential equation set to obtain a simulation response, and constructing a reward function according to the simulation response; step S7, setting neural network parameters; step S8, optimizing by using a deep reinforcement learning algorithm; and step S9, loop iteration is carried out, the optimal physical parameters of inversion are output, and updating of the digital twinborn body is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the research field of digital twin applications in equipment operation and maintenance, and in particular to a method for updating a digital twin model of a planetary gear system based on deep reinforcement learning. Background Art

[0002] Planetary gear systems offer advantages such as compact structure, high transmission ratio, long service life, and light weight. However, due to their small size, high power-to-weight ratio, harsh operating environment, and complex and variable operating conditions, they are prone to failure. Failure to promptly detect and correct planetary gear system failures can easily result in severe economic losses or even casualties.

[0003] Digital twin technology digitally creates virtual models of physical entities. A digital twin is a digital representation of a real-world entity or system, and can be used to understand, predict, optimize, and control it. To ensure the proper functioning of planetary gearboxes, digital twin technology can be incorporated into the health monitoring of planetary gear trains. By building a digital model of the planetary gear system in the virtual world, the system's operation can be monitored in real time through interaction between the digital twin and the physical entity.

[0004] Digital twin technology requires a high level of virtual model precision. This requires not only an accurate description of the physical entity's model structure but also reliable model parameters. System parameter identification is fundamental to building an accurate digital twin model. However, the complex physical structure and operating conditions of planetary gear trains make accurate parameter identification difficult. Therefore, determining parameters through inversion analysis has become a key method for revising and updating twin models.

[0005] Currently, commonly used optimization methods and heuristic algorithms include genetic algorithms and particle swarm optimization (Zhao Kewei. Research and Application of Digital Twin Model Construction and Update Methods for Planetary Gear Systems [D]. South China University of Technology, 2024). These methods cannot accumulate search experience, so when encountering similar problems, a significant amount of time is still required to search from scratch. Deep reinforcement learning can be considered an optimization method, training an intelligent agent to find the optimal action strategy. Deep reinforcement learning can accumulate the agent's search experience in the environment and quickly optimize similar problems. Therefore, it has been widely used in high-dimensional perception and decision-making control, achieving good results (Zhou Xiaotian. Research on Deep Reinforcement Learning Methods for Wind Turbine Parameter Optimization [D]. Shanghai Jiao Tong University, 2018). Planetary gear systems often operate continuously under the same operating conditions. Therefore, researching a deep reinforcement learning-based digital twin model update method for planetary gear systems is of great significance to ensuring the healthy operation of planetary gear systems.

[0006] The existing method for constructing and updating the digital twin model of a planetary gear system (Research on the method for constructing and updating the digital twin model of a planetary gear system, School of Mechanical and Automotive Engineering, South China University of Technology, Zhao Kewei) and its application use the particle swarm optimization algorithm to perform parameter inversion update on the parameters of the constructed digital twin model of the planetary gear system. The particle swarm algorithm has a certain parameter update accuracy, but the iterative calculation time is long (about 3-4 hours), and the search experience cannot be retained. When encountering similar working conditions again, it needs to be searched again, which is time-consuming overall.

[0007] The existing deep reinforcement learning method for wind turbine parameter optimization design (Research on Deep Reinforcement Learning Method for Wind Turbine Parameter Optimization Design, Shanghai Jiao Tong University: School of Mechanical and Power Engineering, Zhou Xiaotian) uses the DQN algorithm in deep reinforcement learning to optimize the design parameters of wind turbines. The results show that the deep reinforcement learning algorithm can achieve better optimization effects while significantly reducing the optimization time. The present invention refers to its ideas when constructing a reinforcement learning environment. When constructing a parameter inversion method for a digital twin model of a planetary gear system, it pays more attention to the physical properties and structural characteristics of the planetary gear system, and is applicable to the planetary gear system.

[0008] At present, the application of deep reinforcement learning in the mechanical field is still concentrated on diagnostic decision-making, and there is little literature on parameter inversion optimization. In terms of combining planetary gear systems with deep reinforcement learning, there is currently no research on combining the two for model parameter inversion. Summary of the Invention

[0009] The purpose of the present invention is to overcome the shortcomings and deficiencies of the above-mentioned prior art, and to provide a method for updating a digital twin model of a planetary gear system based on deep reinforcement learning. The digital twin model is used to evaluate the operating status parameters of the planetary gearbox, so it is necessary to ensure high accuracy and computational efficiency. Although the particle swarm algorithm can guarantee accuracy, it is difficult to meet the real-time nature of the calculation, and the calculated parameters can only indicate the operating status of a few hours ago. The purpose of introducing deep reinforcement learning is to value its use of the neural network to save experience after training the value network, so that similar parameter inversion problems can be carried out in the future. Under the premise of ensuring a certain accuracy, the inversion speed can be reduced from hours to minutes.

[0010] The model updating method of the present invention constructs a rigid-flexible coupling model of the planetary gear system, takes the first three-order meshing frequency multiplications and their sideband amplitudes in the frequency domain of the shell vibration signal as the optimization target, uses the deep reinforcement learning method to find the optimal solution, and stores the learning experience in the target value network to realize the parameter update of the digital twin model of the planetary gear system and the rapid optimization of parameters under similar working conditions.

[0011] The present invention is achieved through at least one of the following technical solutions.

[0012] The digital twin model updating method of a planetary gear system based on deep reinforcement learning includes the following steps:

[0013] Step S1, obtaining parameters of each gear of the planetary gear system;

[0014] Step S2: using the lumped parameter method to describe the dynamic model, and constructing the vibration differential equations of the planetary gear train using the D'Alembert principle;

[0015] Step S3: Calculate the frequency response function family through finite element analysis, construct a rigid-flexible coupling model of the planetary gearbox, i.e., a digital twin model of the planetary gearbox, and simulate the radial vibration acceleration response of the planetary gearbox housing under normal operating conditions through numerical simulation;

[0016] Step S4: Install a vibration acceleration sensor on the existing planetary gearbox housing and use a vibration signal acquisition device to obtain a vibration response signal at a stable speed;

[0017] Step S5: Determine the dynamic model parameters that need to be updated, and based on the set physical parameter value range, define the action space and state space, and establish a deep reinforcement learning environment that includes the rigid-flexible coupling dynamic model of the planetary gear system;

[0018] Step S6: Call the vibration differential equation to calculate the simulation response, take the difference between the amplitudes of the virtual and real signals at the key frequencies in the frequency domain as the optimization target, and construct a reinforcement learning reward function;

[0019] Step S7: Setting the agent type and related network parameters;

[0020] Step S8: Optimize using a deep reinforcement learning algorithm, store process parameters in a replay memory unit, and use the current value network to update the target value network at regular time steps. Each update of the current value network is defined as a round of training.

[0021] Step S9: Check whether the training termination conditions are met; if the conditions are not met, perform the next round of value-action pair update; if the conditions are met, output the inverted optimal physical parameters to achieve the update of the digital twin model of the planetary gearbox.

[0022] Furthermore, in step S2, the gear meshing process is regarded as a time-varying mass-spring-damper system using the lumped parameter method. The sun gear-planet gear meshing pair and the planet gear-ring gear meshing pair are discretized, and the vibration differential equation is established according to the D'Alembert principle: Where q is the system generalized displacement matrix; and They represent the system generalized velocity array and the system generalized acceleration array respectively; M is the system mass matrix; C(t) is the system time-varying damping matrix; K(t) is the system time-varying stiffness matrix; T is the system external torque matrix, and t is the time series of the meshing process.

[0023] Furthermore, in step S3, finite element analysis is performed on the finite element analysis software to obtain a family of frequency response functions;

[0024] The vibration differential equation in step S2 is numerically solved by the backward difference method, and the excitation force distributed to the inner ring gear teeth and the corresponding impulse response function are convolved and summed to obtain the radial vibration acceleration response of the planetary gearbox housing.

[0025] Furthermore, step S4 specifically includes the following steps:

[0026] S4.1. Install a piezoelectric accelerometer on the planetary gearbox housing to measure the radial vibration acceleration signal of the planetary gearbox. Connect the sensor, signal acquisition equipment, and host computer for digital-to-analog conversion and subsequent analysis of the signal. Connect the servo motor, distribution box, and host computer to control the motor speed.

[0027] S4.2. Set the sampling frequency and signal channel, set the motor speed, set the duration of each signal acquisition, and perform real-time acquisition of vibration signals.

[0028] Furthermore, in step S5, the dynamic model parameters that need to be updated include sun gear density, planet gear density, inner gear density, sun gear elastic modulus, planet gear elastic modulus, inner gear elastic modulus, sun gear support radial stiffness, planet gear support radial stiffness, inner gear support radial stiffness, sun gear support circumferential stiffness, planet gear support circumferential stiffness, inner gear support circumferential stiffness, attenuation coefficient, and output torque;

[0029] After obtaining the upper and lower limits of the dynamic model parameters that need to be updated based on experience, the upper and lower limits of each parameter can be normalized to the [0,1] space:

[0030]

[0031] Where p is a positive integer, s p represents the pth value in the state space, x p is the original value of the pth parameter, x pmax 、x pmin are the upper and lower limit values of the pth parameter, respectively, and the state space set S={s1,s2,…,s n}, s n Represents the nth value of the state space, where n is the number of parameters;

[0032] The action space is defined as a discrete 2n value, action space A = {1,2,…,2n}, actions 1 to n correspond to increasing the 1st to nth variables by a step value, and actions n+1 to 2n correspond to decreasing the 1st to nth variables by a step value. When an action causes the variable to exceed the boundary range of [0,1], the action is not executed, thereby discretizing the action space.

[0033] Furthermore, in step S6, the dynamic model is called to restore the state space to the parameter range and substitute it into the vibration differential equation constructed in step S2 to obtain the simulated acceleration frequency domain signal:

[0034] x p =s p (x pmax -x pmin )+x pmin ;

[0035] where s p represents the pth value in the state space, x p is the original value of the pth parameter, x pm 、x pmin are the upper and lower limit values of the pth parameter respectively; the vibration response signal obtained from the experiment is converted from the time domain to the frequency domain using the fast Fourier transform, the first three-order meshing frequencies and their sidebands are selected as the key frequencies, the spectrum of the measured signal spectrum and the simulation signal spectrum are calibrated using the energy center of gravity method, and the amplitude difference between the measured signal and the simulation signal at the key frequency is analyzed; the simulation signal extraction frequency F is expressed as F = (f1, f2, ..., f m ), the measured signal extraction frequency F r Indicated as F r =(f r1 ,f r2 ,…,f rm ), the amplitude Y of the simulated signal is expressed as Y=(y1,y2,…,y m ), the measured signal amplitude Y is expressed as Y r =(y r1 ,y r2 ,…,y rm ), and build a reinforcement learning reward function based on the measured signal and the simulated signal:

[0036]

[0037] Among them, R is the reward function, A, α, β are constant coefficients, y rj The average value of m is the number of key frequencies, f j Extract the jth value of the frequency for the simulation signal, f rj Extract the jth value of the frequency for the measured signal, yj Extract the jth value of the amplitude of the simulation signal, y rj Extract the j-th value of the amplitude of the measured signal, where j is a positive integer in the interval [1, m].

[0038] Furthermore, in step S7, the neural network parameters are set according to the defined action space and state space, including the number of network layers, the number of neurons in each layer, the activation function type, the loss function type, the learning rate, and the discount factor.

[0039] Furthermore, in step S8, the agent adjusts its action strategy according to the environmental feedback, and the discounted cumulative reward G r Expressed as:

[0040]

[0041] Among them, R r represents the reward obtained by the agent in round r, γ is the discount factor, which represents the value ratio of the expected reward in the future at the current moment, and k is an integer in the interval [0, +∞];

[0042] During the interaction, the agent continuously adjusts its strategy π according to the environment reward to find the best strategy π * , expressed as:

[0043]

[0044] Among them, E π represents the expected reward value obtained by the agent under strategy π, s r represents the state value of the agent in round r, s represents any state, and S represents the state space;

[0045] The agent strategy is stored in the form of a value network. According to the settings, the current value network is used to update the target value network every N time steps. Each update of the current value network is defined as a round of training.

[0046] Furthermore, in step S9, check whether the training results meet the model update requirements; if not, continue to update the state-action pairs and start again from step S8; if satisfied, terminate the digital twin model update process of this section of collected signals, output the dynamic parameters obtained by inversion, and start again from step S4.

[0047] A computer device of the present invention includes: a memory and a processor and a computer program stored in the memory. When the computer program is executed on the processor, the method for updating the digital twin model of a planetary gear system based on deep reinforcement learning is implemented.

[0048] Compared with the prior art, the present invention has the following advantages and effects:

[0049] The present invention sets the parameter optimization target as the first three meshing frequency multiples and their sideband amplitudes of the vibration signal, which can greatly reduce the interference of noise on parameter inversion and achieve more accurate model update; the first three meshing frequency multiples and their sidebands are selected for practical considerations, and the range is limited to the first three meshing frequency multiples with the largest energy. The modulation sidebands are clear and the amplitude changes are obvious. The amount of calculation is reduced while ensuring the accuracy of parameter inversion, and a more accurate model update is achieved; a rigid-flexible coupling model is used to describe the planetary gearbox dynamic system while retaining accuracy and efficiency, and a deep reinforcement learning method is used for system parameter inversion. While completing the rapid optimization of parameters, the training experience can be retained, which saves training time for encountering similar working conditions again and improves the efficiency of parameter inversion. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 This is a flowchart of an embodiment of a digital twin model update of a planetary gear system based on deep reinforcement learning.

[0051] Figure 2 1 is a cluster diagram of the ring gear-sensor frequency response function in an embodiment.

[0052] Figure 3 It is an order spectrum diagram of the simulation of the steady-state response of the radial vibration acceleration of the planetary gearbox housing in the embodiment.

[0053] Figure 4 It is the simulation spectrum diagram near the meshing order.

[0054] Figure 5 This is a flowchart of the DDQN algorithm.

[0055] Figure 6 It is a reward curve diagram of the deep reinforcement learning training process.

[0056] Figure 7 It is a comparison diagram of the amplitudes of virtual and real signals near the meshing order.

[0057] Figure 8 It is a comparison diagram of the amplitudes of the virtual and real signals near the second-order meshing order.

[0058] Figure 9 It is a comparison diagram of the amplitudes of the virtual and real signals near the third-order meshing order. DETAILED DESCRIPTION

[0059] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0060] like Figure 1 As shown, this embodiment is based on a planetary gear system digital twin model updating method based on deep reinforcement learning, which is mainly used for updating key parameters of the planetary gearbox dynamic model.

[0061] Taking a planetary gearbox with stable speed and load under fault-free conditions as an example, the specific steps include the following:

[0062] Step S1: To simplify the model planetary gear system, the following assumptions are made: ignoring the axial degrees of freedom of the gears; disregarding the inertia and degrees of freedom introduced by the input and output ends; disregarding the internal loads of the gearbox; isotropic planetary gear support stiffness and damping; all damping involved is linear viscous damping; the gear tooth meshing error is zero; and there is no speed fluctuation or load fluctuation. The parameters of the gears involved in the planetary gear system are shown in Table 1.

[0063] Table 1 Planetary gear system parameters

[0064]

[0065]

[0066] Step S2: using the lumped parameter method to describe the dynamic model, and constructing the vibration differential equations of the planetary gear train using the D'Alembert principle;

[0067] The lumped parameter method considers the gear meshing process as a time-varying mass-spring-damper system and simplifies the planetary gear train into an 18-degree-of-freedom dynamic model. Based on the relative motion relationship of the gear pair in the meshing line direction, the radial and circumferential stiffness and damping caused by the bearing support and the influence of the driving torque are considered. Using the D'Alembert principle, the differential equations of motion of the sun gear are expressed as follows:

[0068]

[0069] Among them, m s , I s and r bs Respectively represent the mass, moment of inertia and base circle radius of the sun gear; x s 、y s Represent the displacement of the sun gear in the x-axis and y-axis directions respectively; Respectively represent the displacement speed of the sun gear in the x-axis and y-axis directions; Respectively represent the displacement acceleration of the sun gear in the x-axis and y-axis directions; k spi and c spi Respectively represent the meshing stiffness and meshing damping of the i-th planetary gear-sun gear; δ spiIt represents the projection of the relative displacement between the sun gear and the i-th planet gear in the direction of the meshing line; represents the projection of the relative speed between the sun gear and the i-th planet gear in the direction of the meshing line; k s and c s They represent the radial stiffness and radial damping of the sun gear support respectively; k st and c st Respectively represent the circumferential stiffness and circumferential damping of the sun gear support; ω c Indicates the planet carrier speed; is the angular acceleration of the sun gear around the z-axis; is the angular velocity of the sun gear around the z axis; u s is the angular displacement of the sun gear around the z axis; T in represents the driving torque; It represents the angle between the meshing line and the x-axis.

[0070] Similarly, the motion differential equations of the internal gear ring are expressed as follows:

[0071]

[0072] Among them, m r , I r and r br Respectively represent the mass, moment of inertia and base circle radius of the inner gear ring; x r 、y r Respectively represent the displacement of the inner gear ring in the x-axis and y-axis directions; Respectively represent the displacement speed of the inner gear ring in the x-axis and y-axis directions; Respectively represent the displacement acceleration of the inner gear ring in the x-axis and y-axis directions; k rp and c rpi Respectively represent the meshing stiffness and meshing damping of the i-th planetary gear-annular gear; δ rpi represents the projection of the relative displacement between the inner ring gear and the i-th planetary gear in the direction of the meshing line; represents the projection of the relative speed between the inner ring gear and the i-th planetary gear in the direction of the meshing line; k r and c r They represent the radial stiffness and radial damping of the inner gear ring respectively; k rt and c rt They represent the circumferential stiffness and circumferential damping of the inner gear ring respectively; is the angular acceleration of the inner gear ring around the z axis; is the angular velocity of the inner gear ring around the z axis; u r is the angular displacement of the inner gear ring around the z axis; It represents the angle between the meshing line and the x-axis.

[0073] Similarly, based on the contact relationship between the planetary gear, the sun gear, and the inner ring gear, the motion differential equations of the i-th planetary gear are expressed as follows:

[0074]

[0075] Among them, m pi , I pi and r bpi Respectively represent the mass, moment of inertia and base circle radius of the i-th planetary gear; x pi 、y pi Respectively represent the displacement of the i-th planetary gear in the x-axis and y-axis directions; Respectively represent the displacement speed of the i-th planetary gear in the x-axis and y-axis directions; Respectively represent the displacement acceleration of the i-th planetary gear in the x-axis and y-axis directions; is the angular acceleration of the i-th planetary gear around the z-axis; k pi and c pi represent the radial stiffness and radial damping of the i-th planet gear respectively.

[0076] Similarly, the planet carrier is connected to the three planetary gears and is subjected to the output load torque. The motion differential equations of the planet carrier are expressed as follows:

[0077]

[0078] Among them, m c , I c and r c Respectively represent the mass, moment of inertia and center distribution circle radius of the planet carrier; x c 、y c Respectively represent the displacement of the planet carrier in the x-axis and y-axis directions; Respectively represent the displacement speed of the planet carrier in the x-axis and y-axis directions; Respectively represent the displacement acceleration of the planet carrier in the x-axis and y-axis directions; δ xcpi 、v ycpi They represent the projections of the relative displacement between the planet carrier and the inner ring gear in the x-axis and y-axis directions respectively; They represent the projection of the relative speed between the planet carrier and the inner gear ring in the x-axis and y-axis directions respectively; k c and c c They represent the radial stiffness and radial damping of the planet carrier support respectively; is the angular acceleration of the planet carrier around the z axis; is the angular velocity of the planet carrier around the z axis; u c is the angular displacement of the planet carrier around the z axis; k ct and c ct They represent the circumferential stiffness and circumferential damping of the planet carrier respectively; T outIndicates the load moment.

[0079] The above differential equation of motion of the planetary gear train can be organized into a matrix form and expressed as: Where q is the system generalized displacement matrix; and They represent the system generalized velocity array and the system generalized acceleration array respectively; M is the system mass matrix; C(t) is the system time-varying damping matrix; K(t) is the system time-varying stiffness matrix; T is the system external torque matrix, and t is the time series of the meshing process.

[0080] Step S3: Calculate the vibration frequency transfer function cluster through finite element analysis, construct a rigid-flexible coupling model of the planetary gearbox, that is, a digital twin model of the planetary gearbox, make the housing flexible, obtain a time-varying flexible transmission path of the gear meshing excitation, and improve the integrity and accuracy of the simulation response; simulate the radial vibration acceleration response of the planetary gearbox housing under normal operating conditions through numerical simulation.

[0081] The digital twin model of a planetary gearbox consists of simulated excitation and a simulated transfer path. Since this step involves obtaining the frequency response function, the ring gear and upper cover of the planetary gearbox participate in the planetary gearbox response as non-excitation transmission path components. The excitation is generated by the constructed set of vibration differential equations. To simplify the analysis, only the ring gear and upper cover are subjected to finite element analysis.

[0082] As an example, a finite element analysis of the inner ring gear and upper cover assembly of a planetary gearbox was performed in ANSYS Workbench. The Harmonic Response module was used to obtain a family of frequency response functions. The mesh was divided into 60,169 units and 109,145 unit nodes using tetrahedral elements. The sensor position was used as the response point, and the numbers were incremented in a counterclockwise order starting from the inner ring gear tooth closest to the sensor. A unit pulse excitation force was applied to each tooth meshing surface in the positive direction of the sensor. The obtained family of frequency response functions is as follows: Figure 2 As shown in the figure, the natural frequency distribution of the first few orders of the inner gear ring is mainly concentrated in the high-frequency area after 2500 Hz, and there is only one low-amplitude natural frequency component of 1300 Hz in the low-frequency area.

[0083] Use the ode23s solver in MATLAB to numerically solve the differential equations of motion of the planetary gear system (that is, the differential equations of motion of the sun gear, inner ring gear, planetary gear, and planet carrier in step S2 are combined to form the differential equations of motion of the planetary gear system, and all differential equations are combined in matrix form). Set the output time series, set the relative error tolerance and absolute error tolerance of the solver, set the initial values of all independent variables to zero, input the dynamic parameters shown in Table 2, and obtain the meshing excitation force F between the planetary gear and the inner ring gear. rp , the meshing excitation force F between the planetary gear and the sun gear sp :

[0084]

[0085] Among them, F rp (t) and c rp (t) are the meshing stiffness and damping of the planetary gear and the internal gear ring; F sp (t) and c sp (t) are the planet-sun gear meshing stiffness and meshing damping respectively; δ rp (t) and are the projections of the relative displacement and relative speed of the planetary gear and the internal gear ring on the meshing line; δ sp (t) and are the projections of the relative displacement and relative speed of the planetary gear and the sun gear on the meshing line, and t is the time series of the meshing process.

[0086] Table 2 Initial parameters of planetary gear system dynamics

[0087] property sun gear planetary gear Internal gear ring planetary carrier Material 45 steel POM M90 HT300 HT300 <![CDATA[Density / (kg·m -3 )]]> 7850 1410 7300 7300 Mass / (kg) 0.187 0.100 3.231 2.109 <![CDATA[Moment of inertia / (kg·m 2 )]]> <![CDATA[9.901×10 -6 ]]> <![CDATA[8.285×10 -5 ]]> <![CDATA[2.035×10 -4 ]]> <![CDATA[4.018×10 -5 ]]> <![CDATA[Young's modulus / (N·m -2 )]]> <![CDATA[2.1×10 11 ]]> <![CDATA[2.76×10 9 ]]> <![CDATA[1.3×10 11 ]]> <![CDATA[1.3×10 11 ]]> Poisson's ratio 0.31 0.35 0.25 0.25

[0088] The frequency domain form S(f) of the radial vibration acceleration response of the planetary gearbox housing is obtained by convolving and summing the excitation force distributed to the inner ring gear teeth with the corresponding impulse response function:

[0089]

[0090] Among them, F rp (f) is the meshing force F between the planetary gear and the internal gear ring rp (t) frequency domain signal, F sp (f) is the meshing force F between the planetary gear and the sun gear sp (t) frequency domain signal, H w (f) is the transmission path corresponding to the wth tooth of the inner gear ring, z is the number of teeth of the inner gear ring, and σ is F sp (f) Attenuation coefficient of the transmission path passing through the planetary gear.

[0091] The solver obtains the steady-state response order spectrum of the planetary gearbox housing radial vibration acceleration with the output rotation frequency as the order, as shown in the following figure: Figure 3 As shown in the figure, the order spectrum shows obvious meshing frequency harmonics and modulation sidebands. At the same time, the meshing impact excites the natural frequency of the gearbox housing, and a resonant frequency component appears. The spectrum near the meshing order is as follows: Figure 4 Since the number of teeth on the inner ring gear is not a multiple of 3, the three planetary gears are arranged unevenly, so modulation sidebands appear at intervals of the output frequency and its multiples.

[0092] Step S4: Install a vibration acceleration sensor on the existing planetary gearbox housing and use a vibration signal acquisition device to obtain a vibration response signal at a stable speed. This signal is used to extract the amplitude at the key frequency in step S6 as part of the reward function construction. The steps specifically include the following:

[0093] S4.1. Install a PCB piezoelectric acceleration sensor on the planetary gearbox housing to measure the radial vibration acceleration signal of the planetary gearbox. Connect the WLS-9163 acquisition card to the host computer; connect the B2 servo motor, distribution box and the host computer.

[0094] S4.2. Set the sampling frequency and signal acquisition channel. As an embodiment, the motor speed is set to 2000 rpm, and the vibration signal is collected in real time. The single signal acquisition time is 3 seconds.

[0095] Step S5: Determine the dynamic model parameters that need to be updated, define the action space and state space based on the set physical parameter value range, and establish a deep reinforcement learning environment that includes the rigid-flexible coupling dynamic model of the planetary gear system; the deep reinforcement learning environment includes the action space and the state space, and the rigid-flexible coupling dynamic model of the planetary gear system is intended to show that both the action space and the state space are obtained from the dynamic model.

[0096] The dynamic model parameters that need to be updated include sun gear density, planet gear density, ring gear density, sun gear elastic modulus, planet gear elastic modulus, ring gear elastic modulus, sun gear support radial stiffness, planet gear support radial stiffness, ring gear support radial stiffness, sun gear support circumferential stiffness, planet gear support circumferential stiffness, ring gear support circumferential stiffness, attenuation coefficient, and output torque. The range of manually selected parameters is mainly determined by empirical values and manufacturer-provided parameters, as shown in Table 3. After empirically obtaining the upper and lower limits of the dynamic model parameters that need to be updated, the upper and lower limits of each parameter can be normalized to the [0,1] space:

[0097]

[0098] Where p is a positive integer, sp represents the pth value in the state space, x p is the original value of the pth parameter, x pm 、x pmin are the upper and lower limit values of the pth parameter, respectively, and the state space set S={s1,s2,…,s n}, n is the number of parameters, s n Represents the nth value in the state space, and the initial state of each parameter is a random value in [0,1].

[0099] The action space is defined as a discrete 2n value. Actions 1 to n correspond to increasing the 1st to nth variables by a step value, and actions n+1 to 2n correspond to decreasing the 1st to nth variables by a step value. When an action causes the variable to exceed the boundary range of [0,1], the action is not performed. The action space is discretized to obtain the action space A = {1,2,…,2n}.

[0100] Table 3 Inversion parameter ranges for planetary gear systems

[0101] parameter scope parameter scope <![CDATA[Inner gear ring density / (kg·m -3 )]]> [7000,7500] <![CDATA[Sun gear density / (kg·m -3 )]]> [7000,7500] <![CDATA[Planet gear density / (kg·m -3 )]]> [1200,1600] Internal gear ring elastic modulus / (GPa) [200,300] Sun gear elastic modulus / (GPa) [150,250] Planet gear elastic modulus / (GPa) [2.0,3.0] <![CDATA[Radial stiffness of internal gear ring / (N·m -1 )]]> <![CDATA[[2.5×10 8 ,3.5×10 8 ]]]> <![CDATA[Circumferential stiffness of internal gear ring / (N·m -1 )]]> <![CDATA[[1.5×10 8 ,2.5×10 8 ]]]> <![CDATA[Radial stiffness of sun gear / (N·m -1 )]]> <![CDATA[[1.3×10 5 ,2.3×10 5 ]]]> <![CDATA[Circumferential stiffness of sun gear / (N·m -1 )]]> <![CDATA[[5.6×10 5 ,6.6×10 5 ]]]> <![CDATA[Radial stiffness of planetary gear / (N·m -1 )]]> <![CDATA[[3.2×10 7 ,4.2×10 7 ]]]> <![CDATA[Circumferential stiffness of planetary gear / (N·m -1 )]]> <![CDATA[[0.5×10 5 ,1.5×10 5 ]]]> <![CDATA[Radial stiffness of the planet carrier / (N·m -1 )]]> <![CDATA[[6.5×10 7 ,7.5×10 7 ]]]> <![CDATA[Circumferential stiffness of planet carrier / (N·m -1 )]]> <![CDATA[[0.9×10 6 ,1.9×10 6 ]]]> Transfer function scaling factor [0.25,0.75] Fourier series coefficient 1 [1,4] Fourier series coefficient 2 [-3,3] Fourier series coefficient 3 [-2,2] Output torque / (N·m) [10,150] — —

[0102] Step S6: Call the dynamic model to calculate the simulation response, take the difference between the amplitudes of the virtual and real signals at the key frequencies in the frequency domain as the optimization target, and construct a reinforcement learning reward function;

[0103] The state space is restored to the parameter range as follows:

[0104] x p =s p (x pmax -x pmin )+x pmin ;

[0105] Among them, s p is the pth parameter value in the state space, x pmax 、x pm are the upper and lower limit values of the pth physical parameter respectively. The dynamic model is called to substitute the parameters into the calculation to obtain the simulated acceleration frequency domain signal. The vibration time domain signal collected from the experiment is converted to the frequency domain using the fast Fourier transform, and the first three-order meshing frequencies and their sidebands are selected as the optimization targets. The energy center of gravity method is used to perform spectrum correction on the measured signal spectrum and the simulated signal spectrum respectively, and the difference in amplitude between the measured signal and the simulated signal at these key frequency components is analyzed. The extracted frequency of the simulation signal is expressed as F = (f1, f2, ..., f m ), the measured signal extraction frequency is expressed as F r =(f r1 ,f r1 ,…,f rm), the amplitude of the simulated signal is expressed as Y=(y1,y2,…,y m ), the amplitude of the measured signal is expressed as Y r =(y r1 ,y r2 ,…,y rm ). Construct a reinforcement learning reward function based on the measured signal and the simulated signal:

[0106]

[0107] Among them, R is the reward function, A, α, β are constant coefficients, y rj The average value of m is the number of key frequencies, f m Extract the mth value of the frequency for the simulation signal, f rm Extract the mth value of the frequency of the measured signal, y m Extract the mth value of the amplitude of the simulation signal, y rm The mth value of the amplitude extracted from the measured signal, j is a positive integer in the interval [1, m], f rj Extract the jth value of the frequency for the simulation signal, f j Extract the jth value of the frequency for the measured signal, y rj Extract the jth value of the amplitude of the measured signal, y j Extract the j-th value of the amplitude for the simulated signal.

[0108] Step S7: Set the agent type and related network parameters.

[0109] Use the Reinforcement Learning Toolbox in MATLAB to set relevant network parameters according to the action space and state space defined in the previous steps to build a deep reinforcement learning process. The algorithm diagram is shown below. Figure 5 After starting training, set the number of rounds r = 0, randomly initialize the action space a0 and state space s0, and initialize the network parameters. The agent selects action a according to the ε-greedy strategy. t . Execute action a r After that, get the state s r+1 , and call the digital twin model of the planetary gear system to calculate the reward R for round r r 。 Will (s r ,a r ,R r ,s r+1) is stored in the experience pool, and the target network parameters are updated every N moments. Gradient descent is applied to the loss function to update the current network parameters. The optimal parameter set is determined to determine whether the number of training rounds or rewards meets the training requirements. If so, the optimal parameter set is output. Otherwise, the number of rounds is increased by 1, and the agent continues training by selecting actions based on the policy. Since both the action space and the state space are discretized, the deep reinforcement learning algorithm DDQN is used for inversion, and the number of training rounds is set to 3000. The network is set to 2 layers, with 100 neurons per layer, and the ReLU function is used as the activation function.

[0110] The loss function uses the mean square error (MSE), which is defined as follows:

[0111]

[0112] Among them, H is the number of key frequencies, y h The hth value of the amplitude of the simulated signal is extracted, y rh Extract the hth value of the amplitude for the measured signal.

[0113] The agent learning rate is set to 0.008, the discount factor is set to 0.9, and the ε-greedy algorithm coefficient is set to 0.99.

[0114] Step S8 uses the existing DDQN algorithm in deep reinforcement learning, which is described in Hasselt HV, Guez A, Silver D. Deep Reinforcement Learning with Double Q-Learning[C]. Proceedings of the AAAI Conference on Artificial Intelligence, 2016: 2094–2100. The agent adjusts its action strategy based on environmental feedback, and the discounted cumulative reward G is used. r It can be expressed as:

[0115]

[0116] Among them, R r represents the reward obtained by the agent in round r, γ is the discount factor, which represents the value ratio of the expected reward in the future at the current moment, and k is an integer in the interval [0, +∞].

[0117] During the interaction, the agent continuously adjusts its strategy π according to the environment reward to find the best strategy π * , expressed as:

[0118]

[0119] Among them, E πrepresents the expected reward value obtained by the agent under strategy π, s r represents the state value of the agent in round r, s represents any state, and S represents the state space.

[0120] The process parameters are stored in the replay memory unit: the agent strategy is stored in the form of a value network, and the agent is set to use the current value network parameters to update the target value network parameters every 8 time steps. Each update of the current value network is defined as a round of training.

[0121] Step S9: Check whether the training process has converged and whether the training results meet the model update requirements. Specifically, check whether the output reward of each iteration meets the requirements or whether the number of iterations meets the requirements to determine the termination conditions. If the number of training rounds or the accuracy requirements are not met, continue to update the state-action pair and start again from step S8. If they are met, terminate the digital twin model update process of this section of the collected signal, output the inverted dynamic parameters, and start again from step S4.

[0122] The training process is as follows Figure 6 As shown in the figure, after 3000 rounds of training, the error is reduced to 0.45, which takes 11.1 hours. The amplitude comparison between the simulated signal obtained by inversion and the real signal near the meshing order is shown in the figure. Figure 7 As shown, the amplitude comparison near the second-order meshing order is as follows Figure 8 As shown, the amplitude comparison near the third-order meshing order is as follows Figure 9 As shown in the figure, it can be seen that the simulated signal and the real signal have the same frequency components in the first three meshing orders and their sidebands, and are also relatively similar in amplitude. It can be considered that the digital twin model of the planetary gear system has been updated.

[0123] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the above embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be considered as equivalent replacement methods and are included in the scope of protection of the present invention.

Claims

1. A digital twin model updating method for a planetary gear system based on deep reinforcement learning, characterized in that: The following steps are involved: Step S1, obtaining parameters of each gear of the planetary gear system; Step S2: using the lumped parameter method to describe the dynamic model, and constructing the vibration differential equations of the planetary gear train using the D'Alembert principle; Step S3: Calculate the frequency response function family through finite element analysis, construct a rigid-flexible coupling model of the planetary gearbox, i.e., a digital twin model of the planetary gearbox, and simulate the radial vibration acceleration response of the planetary gearbox housing under normal operating conditions through numerical simulation; Step S4: Install a vibration acceleration sensor on the existing planetary gearbox housing and use a vibration signal acquisition device to obtain a vibration response signal at a stable speed; Step S5: Determine the dynamic model parameters that need to be updated, and based on the set physical parameter value range, define the action space and state space, and establish a deep reinforcement learning environment that includes the rigid-flexible coupling dynamic model of the planetary gear system; Step S6: Call the vibration differential equation to calculate the simulation response, take the difference between the amplitudes of the virtual and real signals at the key frequencies in the frequency domain as the optimization target, and construct a reinforcement learning reward function; Step S7: Setting the agent type and related network parameters; Step S8: Optimize using a deep reinforcement learning algorithm, store process parameters in a replay memory unit, and use the current value network to update the target value network at regular time steps. Each update of the current value network is defined as a round of training. Step S9: Check whether the training termination conditions are met; if the conditions are not met, perform the next round of value-action pair update; if the conditions are met, output the inverted optimal physical parameters to achieve the update of the digital twin model of the planetary gearbox.

2. The method for updating a digital twin model of a planetary gear system based on deep reinforcement learning according to claim 1, characterized in that: In step S2, the gear meshing process is regarded as a time-varying mass-spring-damper system using the lumped parameter method. The sun gear-planet gear meshing pair and the planet gear-ring gear meshing pair are discretized, and the vibration differential equation is established according to the D'Alembert principle: Where q is the system generalized displacement matrix; and They represent the system generalized velocity matrix and the system generalized acceleration matrix respectively; M is the system mass matrix; C(t) is the system time-varying damping matrix; K(t) is the system time-varying stiffness matrix; t is the system external torque matrix, and t is the time series of the meshing process.

3. The method for updating a digital twin model of a planetary gear system based on deep reinforcement learning according to claim 1, characterized in that: In step S3, finite element analysis is performed using finite element analysis software to obtain a family of frequency response functions; The vibration differential equation in step S2 is numerically solved by the backward difference method, and the excitation force distributed to the inner ring gear teeth and the corresponding impulse response function are convolved and summed to obtain the radial vibration acceleration response of the planetary gearbox housing.

4. The method for updating a digital twin model of a planetary gear system based on deep reinforcement learning according to claim 1, wherein: Step S4 specifically includes the following steps: S4.

1. Install a piezoelectric accelerometer on the planetary gearbox housing to measure the radial vibration acceleration signal of the planetary gearbox. Connect the sensor, signal acquisition equipment, and host computer for digital-to-analog conversion and subsequent analysis of the signal. Connect the servo motor, distribution box, and host computer to control the motor speed. S4.

2. Set the sampling frequency and signal channel, set the motor speed, set the duration of each signal acquisition, and perform real-time acquisition of vibration signals.

5. The method for updating a digital twin model of a planetary gear system based on deep reinforcement learning according to claim 1, wherein: In step S5, the dynamic model parameters that need to be updated include the sun gear density, the planet gear density, the inner ring gear density, the sun gear elastic modulus, the planet gear elastic modulus, the inner ring gear elastic modulus, the sun gear support radial stiffness, the planet gear support radial stiffness, the inner ring gear support radial stiffness, the sun gear support circumferential stiffness, the planet gear support circumferential stiffness, the inner ring gear support circumferential stiffness, the attenuation coefficient, and the output torque; After obtaining the upper and lower limits of the dynamic model parameters that need to be updated based on experience, the upper and lower limits of each parameter can be normalized to the [0,1] space: Where p is a positive integer, s p represents the pth value in the state space, x p is the original value of the pth parameter, x pmax 、x pmin are the upper and lower limit values of the pth parameter, respectively, and the state space set S={s1,s2,…,s n }, s n Represents the nth value of the state space, where n is the number of parameters; The action space is defined as a discrete 2n value, action space A = {1,2,…,2n}, actions 1 to n correspond to increasing the 1st to nth variables by a step value, and actions n+1 to 2n correspond to decreasing the 1st to nth variables by a step value. When an action causes the variable to exceed the boundary range of [0,1], the action is not executed, thereby discretizing the action space.

6. The method for updating a digital twin model of a planetary gear system based on deep reinforcement learning according to claim 1, characterized in that: In step S6, the dynamic model is called to restore the state space to the parameter range and substitute it into the vibration differential equation constructed in step S2 to obtain the simulated acceleration frequency domain signal: x p =s p (x pmax -x pmin )+x pmin ; where s p represents the pth value in the state space, x p is the original value of the pth parameter, x pmax 、x pmin are the upper and lower limit values of the pth parameter respectively; the vibration response signal obtained from the experiment is converted from the time domain to the frequency domain using the fast Fourier transform, the first three-order meshing frequencies and their sidebands are selected as the key frequencies, the spectrum of the measured signal spectrum and the simulation signal spectrum are calibrated using the energy center of gravity method, and the amplitude difference between the measured signal and the simulation signal at the key frequency is analyzed; the simulation signal extraction frequency F is expressed as F = (f1, f2, ..., f m ), the measured signal extraction frequency F r Indicated as F r =(f r1 ,f r2 ,…,f rm ), the amplitude Y of the simulated signal is expressed as Y=(y1,y2,…,y m ), the measured signal amplitude Y is expressed as Y r =(y r1 ,y r2 ,…,y rm ), and build a reinforcement learning reward function based on the measured signal and the simulated signal: Among them, R is the reward function, A, α, β are constant coefficients, y rj The average value of m is the number of key frequencies, f j Extract the jth value of the frequency for the simulation signal, f rj Extract the jth value of the frequency for the measured signal, y j Extract the jth value of the amplitude of the simulation signal, y rj Extract the j-th value of the amplitude of the measured signal, where j is a positive integer in the interval [1, m].

7. The method for updating a digital twin model of a planetary gear system based on deep reinforcement learning according to claim 1, characterized in that: In step S7, the neural network parameters are set according to the defined action space and state space, including the number of network layers, the number of neurons in each layer, the activation function type, the loss function type, the learning rate, and the discount factor.

8. The method for updating a digital twin model of a planetary gear system based on deep reinforcement learning according to claim 1, characterized in that: In step S8, the agent adjusts its action strategy based on the environmental feedback, and the discounted cumulative reward G r Expressed as: Among them, R r represents the reward obtained by the agent in round r, γ is the discount factor, which represents the value ratio of the expected reward in the future at the current moment, and k is an integer in the interval [0, +∞]; During the interaction, the agent continuously adjusts its strategy π according to the environment reward to find the best strategy π * , expressed as: Among them, E π represents the expected reward value obtained by the agent under strategy π, s r represents the state value of the agent in round r, s represents any state, and S represents the state space; The agent strategy is stored in the form of a value network. According to the settings, the current value network is used to update the target value network every N time steps. Each update of the current value network is defined as a round of training.

9. The method for updating a digital twin model of a planetary gear system based on deep reinforcement learning according to any one of claims 1 to 8, characterized in that: In step S9, check whether the training results meet the model update requirements; if not, continue to update the state-action pairs and start again from step S8; if satisfied, terminate the digital twin model update process of this section of collected signals, output the dynamic parameters obtained by inversion, and start again from step S4.

10. A computer device, characterized in that: It comprises: a memory and a processor and a computer program stored in the memory. When the computer program is executed on the processor, it implements the method for updating the digital twin model of a planetary gear system based on deep reinforcement learning as described in any one of claims 1 to 9.

Citation Information

Cited By

  • Reinforcement learning compensation method and system for gear multi-body dynamics constraint

    CN120745471A

  • Planetary gearbox gear ring load identification method based on vibration response inversion

    CN121580525A

  • A planet gear box ring gear load identification method based on vibration response inversion

    CN121580525B