Bearing lubrication parameter self-adaptive regulation and control system based on reinforcement learning

By combining the improved Dreamer model with physical consistency loss terms and risk penalty terms, the problems of poor dynamic adaptability and insufficient safety in bearing lubrication control are solved, achieving high-precision and high-safety adaptive control of lubrication parameters, thereby improving bearing service life and system reliability.

CN122018299APending Publication Date: 2026-05-12上海祎榕实业有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
上海祎榕实业有限公司
Filing Date
2026-03-18
Publication Date
2026-05-12

Smart Images

  • Figure CN122018299A_ABST
    Figure CN122018299A_ABST
Patent Text Reader

Abstract

The invention discloses a bearing lubrication parameter adaptive regulation and control system based on reinforcement learning, and the system comprises a bearing operation state collection module which is used for synchronizing multi-source data; the state vector construction module is used for preprocessing the multi-source data and combining the multi-source data into a state vector; the world model learning module is used for outputting the state and uncertainty of the next moment by using an improved Dreamer model and introducing a physical consistency loss item; the optimal action module is used for carrying out multi-step prospective planning through a behavior network and outputting an optimal regulation and control action; the regulation and control execution and reward evaluation module is used for executing regulation and control and calculating a real reward value based on multi-index feedback; and the model self-optimization module is used for continuously optimizing the model based on experience playback and time sequence difference errors. The problems that a traditional method is poor in dynamic adaptability, incredible in decision and low in safety are solved, and high safety and intelligent self-adaptive regulation and control of bearing lubrication are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of intelligent manufacturing and reinforcement learning, and in particular to an adaptive control system for bearing lubrication parameters based on reinforcement learning. Background Technology

[0002] Reinforcement learning algorithms, due to their ability to make autonomous decisions and optimize through trial and error in complex environments, have shown great potential in recent years in fields such as robot control, resource scheduling, and industrial process optimization, and are considered a key technology path for achieving intelligent equipment operation and maintenance. However, in the critical industrial scenario of bearing lubrication, practical applications face many challenges, such as the large demand for model training data, the physical uninterpretability of the decision-making process, and high requirements for safety and stability, which severely restricts the deployment effectiveness of traditional reinforcement learning.

[0003] Currently, most bearing lubrication control methods rely on fixed parameter thresholds or simple PID control logic, which are difficult to adapt to the dynamic lubrication requirements under varying operating conditions and loads. This leads to frequent occurrences of insufficient or excessive lubrication, making it impossible to achieve the optimal balance between energy efficiency and lifespan. Some systems that introduce intelligent algorithms have decision models that are like black boxes, outputting only based on data correlation and ignoring the underlying physical laws such as shaft dynamics and fluid lubrication that bearings follow. This results in a lack of physical interpretability in their decision-making results, making them prone to dangerous actions that violate common sense when faced with unseen operating conditions, which seriously restricts their application in scenarios with high reliability requirements.

[0004] Furthermore, most existing reinforcement learning-based control methods employ fixed reward function designs, failing to effectively quantify and mitigate the decision-making risks of the model under high uncertainty. This can lead to the system adopting risky strategies that damage bearings during the exploration process. Simultaneously, these methods lack effective mechanisms for integrating domain knowledge into model training, resulting in low model learning efficiency. They require massive amounts of interactive data to converge, making them unsuitable for the sparse and expensive data conditions of real-world industrial environments. This severely impacts the practical value and deployment feasibility of the models in real-world production environments.

[0005] Therefore, how to provide a patient continuous health management system based on reinforcement learning strategy optimization is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] One objective of this invention is to propose an adaptive control system for bearing lubrication parameters based on reinforcement learning. This invention fully integrates key steps such as bearing operating state perception, state vector construction, improved Dreamer model decision-making, physical consistency constraints, and risk avoidance, constructing an intelligent lubrication control process with state vector standardization, embedding of physical laws into the world model, quantification of risk in forward planning, and self-optimization of the control strategy. By introducing a physical consistency loss term, this invention incorporates prior knowledge such as the shaft center trajectory equation and Reynolds equation into the world model training, solving the problems of physical inexplicability and violation of mechanisms in the decision-making process of traditional models. By introducing a risk penalty term in the behavioral network planning, it quantifies and avoids risky decisions under high uncertainty, improving the system's safety and robustness. This invention possesses advantages such as interpretable decision-making mechanisms, high precision of control strategies, strong safety adaptability, and efficient model training, significantly improving the energy efficiency and reliability of bearing lubrication under varying operating conditions, thereby effectively solving the problems of poor dynamic adaptability, "black box" decision-making, and high safety risks in existing methods.

[0007] An adaptive control system for bearing lubrication parameters based on reinforcement learning according to an embodiment of the present invention includes the following modules: The bearing operating status acquisition module is used to simultaneously acquire data from multiple sources. The state vector construction module is used to preprocess multi-source data and combine them sequentially into state vectors; The world model learning module is used to input the state vector into the encoder network of the improved Dreamer model to obtain a low-dimensional latent state vector, and input the predicted low-dimensional latent state vector of the world model to predict the next time step. It outputs a prediction uncertainty metric and calculates the physical consistency loss term by decoding the predicted low-dimensional latent state vector and comparing it with the theoretical physical state vector of the next time step. The optimal action module is used to perform multi-step forward planning in the world model through the behavior network, generate multiple candidate future action trajectories, calculate the final cumulative reward value by subtracting the sum of the physical consistency loss term and the risk penalty term from the basic cumulative reward value, and output a two-dimensional action vector. The control execution and reward evaluation module is used to convert the two-dimensional action vector into a control signal, collect the instantaneous change of bearing housing temperature and the main frequency amplitude of vibration frequency domain characteristics, multiply the instantaneous change of bearing housing temperature, the main frequency amplitude of vibration frequency domain characteristics and the real-time power consumption of the lubrication system by the corresponding preset weight coefficients, sum them, and take the negative number of the sum to obtain the real reward value. The model self-optimization module stores an experience tuple consisting of a state vector, a two-dimensional action vector, and a true reward value into an experience replay pool. It updates the network weights of the improved Dreamer model by minimizing the temporal difference error, and continuously performs self-optimization of the improved Dreamer model.

[0008] Optionally, modules can be integrated using the following methods: S1. Through multi-source sensors, synchronously collect multi-source data such as vibration time-domain signal, bearing housing temperature, lubricating oil film pressure value and real-time spindle speed during bearing operation; S2. Perform fast Fourier transform on the vibration time domain signal to extract the vibration frequency domain features, normalize the bearing housing temperature, lubricating oil film pressure value and real-time spindle speed, and combine the processed vibration frequency domain features, temperature, pressure and speed data into a state vector in sequence. S3. Input the state vector into the encoder network of the improved Dreamer model to obtain a low-dimensional latent state vector, and input it into the world model to predict the predicted low-dimensional latent state vector of the next moment. Output the prediction uncertainty measure, and calculate the physical consistency loss term by comparing the decoded predicted state with the theoretical physical state vector of the next moment. S4. Through a behavioral network, perform multi-step forward planning in the world model to generate multiple candidate future action trajectories and calculate the final cumulative reward value by subtracting the sum of the physical consistency loss term and the risk penalty term from the basic cumulative reward value. Output a two-dimensional action vector that maximizes the final cumulative reward value and includes the target fuel supply pressure value and the target fuel supply frequency value. S5. Convert the two-dimensional motion vector into a control signal, collect the instantaneous change in bearing housing temperature and the main frequency amplitude of vibration frequency domain characteristics, multiply the instantaneous change in bearing housing temperature, the main frequency amplitude of vibration frequency domain characteristics and the real-time power consumption of the lubrication system by the corresponding preset weight coefficients and sum them, and take the negative value of the sum to obtain the real reward value. S6. Store the experience tuple consisting of the state vector, two-dimensional action vector and real reward value into an experience replay pool. Randomly sample a batch of experience tuples from the experience replay pool. Update the network weights of the improved Dreamer model by minimizing the temporal difference error, and continuously perform self-optimization of the improved Dreamer model.

[0009] Optionally, S1 specifically includes: S11. Physically install the vibration sensor, temperature sensor, oil film pressure sensor and speed sensor at the predetermined measuring points on the bearing housing or spindle, and start data acquisition from all sensors simultaneously through a synchronous trigger signal. S12. Synchronously read the vibration time-domain signal output by the vibration sensor, the bearing housing temperature output by the temperature sensor, the lubricating oil film pressure value output by the oil film pressure sensor, and the real-time spindle speed output by the speed sensor.

[0010] Optionally, S2 specifically includes: S21. Perform a fast Fourier transform on the vibration time-domain signal to convert the vibration time-domain signal from the time domain to the frequency domain, extract the amplitude spectrum in the frequency domain as the vibration frequency domain feature, read the bearing housing temperature, lubricating oil film pressure value and spindle real-time speed respectively, and use the preset maximum and minimum values ​​to perform linear normalization processing on each data item to map the numerical range to between 0 and 1. S22. The normalized bearing housing temperature, lubricating oil film pressure value and real-time spindle speed data are spliced ​​with the vibration frequency domain characteristics in a preset order to generate a state vector of preset dimensions.

[0011] Optionally, S3 specifically includes: S31. The state vector is input into the encoder network consisting of a preset number of fully connected layers of the improved Dreamer model. The first fully connected layer calculates the state vector with the preset first weight matrix and the first bias vector, and then performs a nonlinear transformation through the ReLU activation function to obtain the first layer encoding features. The second fully connected layer receives the first layer encoding features and continues to calculate the preset second weight matrix and the second bias vector. S32. The computation output dimension of the repeatedly stacked fully connected layers is a low-dimensional latent state vector with a preset number of fully connected layers. The policy network receives the low-dimensional latent state vector at the current time, calculates it through forward propagation, multiplies the low-dimensional latent state vector with a preset weight matrix, adds a preset bias vector, and processes it through a hyperbolic tangent activation function to obtain two basic action vectors with values ​​between negative one and one. S33. Read the minimum and maximum values ​​of the preset target oil supply pressure and target oil supply frequency, and linearly map the first value of the basic motion vector from the range of negative one to one to the range of the minimum to maximum value of the target oil supply pressure and target oil supply frequency to obtain a basic oil supply pressure value and a basic oil supply frequency value. Repeat all numerical calculations on the basic motion vector to generate a two-dimensional motion vector with actual physical meaning. S34. Input the low-dimensional latent state vector into the preset world model of the improved Dreamer model. The world model consists of a state transition network, a reward network and an uncertainty network. The state transition network receives the low-dimensional latent state vector at the current time and the two-dimensional action vector executed at the previous time. It is calculated through a gated recurrent unit network and outputs the predicted low-dimensional latent state vector for the next time. S35. The reward network receives the low-dimensional latent state vector and the two-dimensional action vector at the current time and combines them into a long vector. It calculates and outputs a scalar instant reward value through a fully connected layer with preset weights and biases. The uncertainty network receives the low-dimensional latent state vector and the two-dimensional action vector at the current time and concatenates them into a long vector. It calculates and outputs a prediction uncertainty measure through another fully connected layer with preset weights and biases. S36. In the training process of the world model, a physical consistency loss term is introduced. The physical consistency loss term inputs the low-dimensional latent state vector predicted by the world model in the next time step into the decoder network, which consists of two layers and a preset fully connected layer, to restore the low-dimensional latent state vector to the predicted physical state vector. S37. Based on the bearing's shaft center trajectory equation and Reynolds equation, calculate a theoretical physical state vector for the next moment using numerical solution methods. Calculate the squared Euclidean distance between the predicted physical state vector and the theoretical physical state vector for the next moment, and use it as the physical consistency loss term. Optionally, S37 specifically includes: S371. Create a two-dimensional array to store the pressure at each point on the oil film grid. Initialize the values ​​at all positions to zero and enter a loop. The loop continues until the preset number of loop stops is met. In each iteration of the loop, traverse each position of this two-dimensional array and read the pressure values ​​at the four adjacent positions above, below, to the left, and to the right of the grid position currently being calculated. S372. Obtain the preset bearing radius clearance value and the offset of the shaft center in the vertical direction, and record the angle corresponding to the current grid point on the circumference. Subtract the product of the offset and the cosine of the angle from the radius clearance value to obtain the oil film thickness of the current grid point. S373. Read the current spindle speed value and multiply it by a constant of sixty to convert the unit from minutes to seconds, and multiply it by the product of pi and the preset journal radius to obtain the linear velocity of the journal surface. S374. Subtract the left pressure value from the right pressure value of the current grid point read in the two-dimensional array to obtain the circumferential pressure difference. Subtract the lower pressure value from the upper pressure value to obtain an axial pressure difference. Multiply the circumferential pressure difference by the preset circumferential grid spacing coefficient to obtain the first term. Multiply the axial pressure difference by the preset axial grid spacing coefficient to obtain the second term. Multiply the linear velocity by the preset lubricating oil viscosity and divide by the square of the oil film thickness to obtain the third term. Add the first, second and third terms together to obtain the total, which is the pressure update term. S375. Multiply the pressure update item by a preset relaxation coefficient greater than 0 and less than 1 to obtain the adjustment amount. Add this adjustment amount to the original pressure value at the current grid position. The new value is the updated pressure, which overwrites the original value. S376. After updating the pressure value for all positions in the two-dimensional array, a global iteration is completed. Check the difference between the old and new pressure values ​​for all positions. If the largest difference is less than the preset difference threshold, stop the loop. Otherwise, repeat the next global iteration. After the loop stops, accumulate all the pressure values ​​for all positions in the two-dimensional array and multiply them by the area represented by a single grid to obtain the total oil film support force. S377. Subtract the known weight of the shaft from the total oil film support force to obtain a net force, and divide it by a preset bearing stiffness coefficient to obtain a displacement. Add this displacement to the original coordinates of the shaft center to obtain the theoretical new coordinates of the shaft center at the next moment. This new coordinate and the two-dimensional array of the final pressure value are sequentially spliced ​​together to form the theoretical physical state vector at the next moment.

[0012] Optionally, S4 specifically includes: S41. Input the low-dimensional potential state vector at the current moment into the behavior network of the improved Dreamer model. The behavior network performs multi-step look-ahead planning in the world model and generates a number of candidate future action trajectories multiplied by the preset planning steps and the number of candidate actions per step. S42. In the planning and optimization process of the behavior network, a risk penalty term related to the prediction uncertainty measure is introduced. For each candidate future action trajectory, the prediction uncertainty measure obtained by the uncertainty network at each step in the preset planning steps is accumulated to obtain the total uncertainty of the current future action trajectory. S43. Multiply the total uncertainty of the current future action trajectory by a preset risk coefficient that is greater than zero to obtain the risk penalty term of the current future action trajectory. Add the physical consistency loss term of the current candidate future action trajectory to the risk penalty term to obtain the final risk penalty term. S44. Multiply the scalar instant reward value obtained by each candidate future action trajectory at each step in the preset planning steps by a preset discount factor that decreases with the number of steps, and sum them up to obtain a basic cumulative reward value. Subtract the final comprehensive risk penalty to obtain the final cumulative reward value. S45. The behavioral network searches and selects the candidate future action trajectory that maximizes the final cumulative reward value from all candidate future action trajectories, and uses the first candidate action vector as the final output, which is a two-dimensional action vector containing the target fuel supply pressure value and the target fuel supply frequency value.

[0013] Optionally, the multi-step forward planning specifically includes: The number of planned steps and the number of candidate actions per step are preset. Starting from the current moment, for each step, the behavior network calls a random number generator to generate two independent random numbers that conform to the standard normal distribution. These two random numbers are multiplied by a preset exploration intensity coefficient that determines the size of the exploration range. These two scaled random numbers are then added to the base fuel supply pressure value and the base fuel supply frequency value, respectively. The range of the summed base oil supply pressure and base oil supply frequency is limited. If it exceeds the preset maximum value, it is set to the maximum value. If it is less than the minimum value, it is set to the minimum value. A candidate two-dimensional action vector is generated. The process is repeated, and a preset number of candidate action vectors are generated each time using a different random number. Each candidate action vector and the current low-dimensional latent state vector are input into the world model to obtain a preset number of predicted low-dimensional latent state vectors and scalar instant reward values, and to generate a preset number of future action trajectories multiplied by the preset number of planning steps and the number of candidate actions per step.

[0014] Optionally, S5 specifically includes: S51. Multiply the target oil supply pressure value in the two-dimensional motion vector by a preset pressure conversion coefficient to obtain the first control signal for driving the proportional pressure valve. Multiply the target oil supply frequency value in the two-dimensional motion vector by a preset frequency conversion coefficient to obtain the second control signal for driving the variable frequency pump. S52. After executing the first control signal and the second control signal, wait for a preset time interval, re-acquire the bearing housing temperature, subtract the bearing housing temperature before the control is executed from the newly acquired bearing housing temperature to obtain the instantaneous change in bearing housing temperature, perform a fast Fourier transform on the re-acquired vibration time domain signal, find the frequency point with the largest amplitude in the obtained vibration frequency domain features and read the amplitude as the main frequency amplitude of the vibration frequency domain features. S53. Read the real-time operating voltage and current of the proportional pressure valve and variable frequency pump in the lubrication system, and multiply the real-time operating voltage and real-time operating current to obtain the real-time power consumption of the lubrication system. S54. Multiply the instantaneous change in bearing housing temperature by a preset temperature weighting coefficient to obtain the first term; multiply the dominant frequency amplitude of the vibration frequency domain characteristic by a preset vibration weighting coefficient to obtain the second term; multiply the real-time power consumption of the lubrication system by a preset power consumption weighting coefficient to obtain the third term; and take the negative sum of the first, second, and third terms to generate the true reward value.

[0015] Optionally, S6 specifically includes: S61. Store the experience tuple consisting of state vector, two-dimensional action vector and real reward value into an experience replay pool. Randomly sample a batch of experience tuples from the experience replay pool. Input the state vector and two-dimensional action vector in each experience tuple into the world model of the improved Dreamer model to obtain a predicted low-dimensional latent state vector and scalar instantaneous reward value for the next time step. S63. Input the predicted low-dimensional latent state vector of the next time step into the behavior network of the improved Dreamer model, calculate a predicted final cumulative reward value, and add the predicted scalar instant reward value to the predicted final cumulative reward value to obtain a predicted total reward value. S64. Input the next-time state vector in each empirical tuple into the encoder network to obtain the true next-time low-dimensional latent state vector, and input it into the improved Dreamer model to calculate the true final cumulative reward value. Then add the true reward value to the true final cumulative reward value to obtain a true total reward value. S65. Calculate the squared Euclidean distance between the predicted total return value and the actual total return value to obtain a temporal difference error. Use the backpropagation algorithm and gradient descent algorithm to update the network weights of the improved Dreamer model based on the temporal difference error, and continuously perform self-optimization of the improved Dreamer model.

[0016] The beneficial effects of this invention are: First, this invention constructs a world model based on an improved Dreamer model to achieve accurate prediction of the future state of bearings. The improved Dreamer model introduces a physical consistency loss term, embedding prior physical knowledge such as the bearing's shaft trajectory equation and Reynolds equation as strong constraints into the model training process. This loss term calculates the difference between the model's predicted state and the theoretical physical state, forcing the world model's learning results to conform to the objective laws of hydrodynamic lubrication and rotor dynamics. This fundamentally solves the problem of absurd decisions that violate physical common sense may arise from purely data-driven models, greatly improving the model's prediction accuracy, physical interpretability, and generalization ability under unknown operating conditions.

[0017] Secondly, in the behavioral network for forward planning based on the world model, this invention introduces a risk penalty term to impose safety constraints on the decision-making process. This risk penalty term consists of two parts: one is the aforementioned physical consistency loss, ensuring the physical rationality of actions; the other part is positively correlated with the prediction uncertainty metric output by the uncertainty network, penalizing decisions the model is "unsure" about. This enables the behavioral network, when optimizing cumulative rewards, not only to pursue high returns but also to proactively avoid high-risk actions that may lead to abnormal physical states or model prediction failures, thereby significantly improving the safety and robustness of the control strategy and effectively preventing potential damage to the bearing during the exploration process.

[0018] In summary, by collaboratively introducing a physical consistency loss term and a risk penalty term into the improved Dreamer model, this invention constructs an intelligent decision-making core that understands both physical principles and risk avoidance, achieving high-precision, high-safety, and high-adaptive optimized control of bearing lubrication parameters in complex dynamic environments. Attached Figure Description

[0019] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 This is a structural diagram of a bearing lubrication parameter adaptive control system based on reinforcement learning proposed in this invention. Figure 2 This is a flowchart of the world model learning and physical consistency constraint based on the improved Dreamer model proposed in this invention. Figure 3 This is a flowchart of the multi-step look-ahead planning and optimal action output of the behavior network based on risk penalty terms proposed in this invention. Detailed Implementation

[0020] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0021] refer to Figures 1-3 An adaptive control system for bearing lubrication parameters based on reinforcement learning includes the following modules: The bearing operating status acquisition module is used to simultaneously acquire data from multiple sources. The state vector construction module is used to preprocess multi-source data and combine them sequentially into state vectors; The world model learning module is used to input the state vector into the encoder network of the improved Dreamer model to obtain a low-dimensional latent state vector, and input the predicted low-dimensional latent state vector of the world model to predict the next time step. It outputs a prediction uncertainty metric and calculates the physical consistency loss term by decoding the predicted low-dimensional latent state vector and comparing it with the theoretical physical state vector of the next time step. The optimal action module is used to perform multi-step forward planning in the world model through the behavior network, generate multiple candidate future action trajectories, calculate the final cumulative reward value by subtracting the sum of the physical consistency loss term and the risk penalty term from the basic cumulative reward value, and output a two-dimensional action vector. The control execution and reward evaluation module is used to convert the two-dimensional action vector into a control signal, collect the instantaneous change of bearing housing temperature and the main frequency amplitude of vibration frequency domain characteristics, multiply the instantaneous change of bearing housing temperature, the main frequency amplitude of vibration frequency domain characteristics and the real-time power consumption of the lubrication system by the corresponding preset weight coefficients, sum them, and take the negative number of the sum to obtain the real reward value. The model self-optimization module stores an experience tuple consisting of a state vector, a two-dimensional action vector, and a true reward value into an experience replay pool. It updates the network weights of the improved Dreamer model by minimizing the temporal difference error, and continuously performs self-optimization of the improved Dreamer model.

[0022] In this embodiment, the modules are interconnected using the following method: S1. Simultaneously collect multi-source data such as vibration time-domain signal, bearing housing temperature, lubricating oil film pressure value and real-time spindle speed during bearing operation through vibration sensor, temperature sensor, oil film pressure sensor and speed sensor. S2. Perform fast Fourier transform on the vibration time domain signal to extract vibration frequency domain features, normalize the bearing housing temperature, lubricating oil film pressure value and real-time spindle speed, and combine the processed vibration frequency domain features, temperature, pressure and speed data into a state vector of preset dimensions in sequence. S3. Input the state vector into the encoder network of the improved Dreamer model to obtain a low-dimensional latent state vector, and input it into the world model to predict the predicted low-dimensional latent state vector of the next moment. Output the prediction uncertainty measure, and calculate the physical consistency loss term by comparing the decoded predicted state with the theoretical physical state vector of the next moment. S4. Through a behavioral network, perform multi-step forward planning in the world model to generate multiple candidate future action trajectories and calculate the final cumulative reward value by subtracting the sum of the physical consistency loss term and the risk penalty term from the basic cumulative reward value. Output a two-dimensional action vector that maximizes the final cumulative reward value and includes the target fuel supply pressure value and the target fuel supply frequency value. S5. Convert the two-dimensional motion vector into control signals to drive the proportional pressure valve and the variable frequency pump. Within the preset time interval after the control is executed, collect the instantaneous change in bearing housing temperature and the main frequency amplitude of vibration frequency domain characteristics. Multiply the instantaneous change in bearing housing temperature, the main frequency amplitude of vibration frequency domain characteristics, and the real-time power consumption of the lubrication system by the corresponding preset weight coefficients and sum them. Take the negative value of the sum to obtain the real reward value. S6. Store the experience tuple consisting of the state vector, two-dimensional action vector and real reward value into an experience replay pool. Randomly sample a batch of experience tuples from the experience replay pool. Update the network weights of the improved Dreamer model by minimizing the temporal difference error, and continuously perform self-optimization of the improved Dreamer model.

[0023] This implementation significantly improves the accuracy, safety, and adaptability of bearing lubrication control. By simultaneously collecting multi-source data to construct a high-dimensional state vector, it provides comprehensive information support for intelligent decision-making. The core lies in the improved Dreamer model, which, by introducing a physical consistency loss term, ensures that the prediction and decision-making processes strictly follow physical laws, solving the problems of the "black box" and unreliability of data-driven models. Simultaneously, by introducing a risk penalty term, the model actively avoids high-risk actions when optimizing performance, greatly improving the safety and robustness of the control strategy. Combined with a closed-loop self-optimization mechanism based on real-world feedback, the model can continuously learn and adapt to the slowly changing characteristics of bearings and complex operating condition disturbances, ultimately achieving a leap from passive response to proactive predictive maintenance, significantly extending bearing life and improving system reliability.

[0024] In this embodiment, S1 specifically includes: S11. Physically install the vibration sensor, temperature sensor, oil film pressure sensor and speed sensor at the predetermined measuring points on the bearing housing or spindle, and start data acquisition from all sensors simultaneously through a synchronous trigger signal. S12. Synchronously read the vibration time-domain signal output by the vibration sensor, the bearing housing temperature output by the temperature sensor, the lubricating oil film pressure value output by the oil film pressure sensor, and the real-time spindle speed output by the speed sensor.

[0025] In this embodiment, S2 specifically includes: S21. Perform a fast Fourier transform on the vibration time-domain signal to convert the vibration time-domain signal from the time domain to the frequency domain, extract the amplitude spectrum in the frequency domain as the vibration frequency domain feature, read the bearing housing temperature, lubricating oil film pressure value and spindle real-time speed respectively, and use the preset maximum and minimum values ​​to perform linear normalization processing on each data item to map the numerical range to between 0 and 1. S22. The normalized bearing housing temperature, lubricating oil film pressure value and real-time spindle speed data are spliced ​​with the vibration frequency domain characteristics in a preset order to generate a state vector of preset dimensions.

[0026] In this embodiment, S3 specifically includes: S31. The state vector is input into the encoder network consisting of a preset number of fully connected layers of the improved Dreamer model. The first fully connected layer calculates the state vector with the preset first weight matrix and the first bias vector, and then performs a nonlinear transformation through the ReLU activation function to obtain the first layer encoding features. The second fully connected layer receives the first layer encoding features and continues to calculate the preset second weight matrix and the second bias vector. S32. The computation output dimension of the repeatedly stacked fully connected layers is a low-dimensional latent state vector with a preset number of fully connected layers. The policy network receives the low-dimensional latent state vector at the current time, calculates it through forward propagation, multiplies the low-dimensional latent state vector with a preset weight matrix, adds a preset bias vector, and processes it through a hyperbolic tangent activation function to obtain two basic action vectors with values ​​between negative one and one. S33. Read the minimum and maximum values ​​of the preset target oil supply pressure and target oil supply frequency, and linearly map the first value of the basic motion vector from the range of negative one to one to the range of the minimum to maximum value of the target oil supply pressure and target oil supply frequency to obtain a basic oil supply pressure value and a basic oil supply frequency value. Repeat all numerical calculations on the basic motion vector to generate a two-dimensional motion vector with actual physical meaning. S34. Input the low-dimensional latent state vector into the preset world model of the improved Dreamer model. The world model consists of a state transition network, a reward network and an uncertainty network. The state transition network receives the low-dimensional latent state vector at the current time and the two-dimensional action vector executed at the previous time. It is calculated through a gated recurrent unit network and outputs the predicted low-dimensional latent state vector for the next time. S35. The reward network receives the low-dimensional latent state vector and the two-dimensional action vector at the current time and combines them into a long vector. It calculates and outputs a scalar instant reward value through a fully connected layer with preset weights and biases. The uncertainty network receives the low-dimensional latent state vector and the two-dimensional action vector at the current time and concatenates them into a long vector. It calculates and outputs a prediction uncertainty measure through another fully connected layer with preset weights and biases. S36. In the training process of the world model, a physical consistency loss term is introduced. The physical consistency loss term inputs the low-dimensional latent state vector predicted by the world model to the decoder network. The decoder network is the inverse structure of the encoder network and consists of two layers and a preset fully connected layer. It restores the low-dimensional latent state vector to the predicted physical state vector. S37. Based on the bearing's shaft center trajectory equation and Reynolds equation, calculate a theoretical physical state vector for the next moment using numerical solution methods. Calculate the squared Euclidean distance between the predicted physical state vector and the theoretical physical state vector for the next moment, and use it as the physical consistency loss term. This implementation significantly improves the accuracy and reliability of bearing state prediction and decision-making by constructing a deeply coupled physics-data dual-driven model. The encoder compresses high-dimensional multi-source data into low-dimensional latent states containing core information, laying the foundation for efficient learning. The world model accurately predicts future states through gated cyclic units and outputs uncertainty metrics, providing a quantitative basis for risk assessment. The core innovation lies in the introduction of a physical consistency loss term. This mechanism uses prior physical knowledge such as the shaft center trajectory and Reynolds equations as strong constraints, forcing the model's predictions to conform to objective physical laws. This fundamentally overcomes the problem of absurd decisions that may arise from purely data-driven models, which violate common sense physics. This gives the model not only data fitting capabilities but also physical interpretability and strong generalization ability under unknown conditions, ensuring the scientific rigor and reliability of subsequent decisions.

[0027] In this embodiment, S37 specifically includes: S371. Create a two-dimensional array to store the pressure at each point on the oil film grid. Initialize the values ​​at all positions to zero and enter a loop. The loop continues until the preset number of loop stops is met. In each iteration of the loop, traverse each position of this two-dimensional array and read the pressure values ​​at the four adjacent positions above, below, to the left, and to the right of the grid position currently being calculated. S372. Obtain the preset bearing radius clearance value and the offset of the shaft center in the vertical direction, and record the angle corresponding to the current grid point on the circumference. Subtract the product of the offset and the cosine of the angle from the radius clearance value to obtain the oil film thickness of the current grid point. S373. Read the current spindle speed value and multiply it by a constant of sixty to convert the unit from minutes to seconds, and multiply it by the product of pi and the preset journal radius to obtain the linear velocity of the journal surface. S374. Subtract the left pressure value from the right pressure value of the current grid point read in the two-dimensional array to obtain the circumferential pressure difference. Subtract the lower pressure value from the upper pressure value to obtain an axial pressure difference. Multiply the circumferential pressure difference by the preset circumferential grid spacing coefficient to obtain the first term. Multiply the axial pressure difference by the preset axial grid spacing coefficient to obtain the second term. Multiply the linear velocity by the preset lubricating oil viscosity and divide by the square of the oil film thickness to obtain the third term. Add the first, second and third terms together to obtain the total, which is the pressure update term. S375. Multiply the pressure update item by a preset relaxation coefficient greater than 0 and less than 1 to obtain the adjustment amount. Add this adjustment amount to the original pressure value at the current grid position. The new value is the updated pressure, which overwrites the original value. S376. After updating the pressure value for all positions in the two-dimensional array, a global iteration is completed. Check the difference between the old and new pressure values ​​for all positions. If the largest difference is less than the preset difference threshold, stop the loop. Otherwise, repeat the next global iteration. After the loop stops, accumulate all the pressure values ​​for all positions in the two-dimensional array and multiply them by the area represented by a single grid to obtain the total oil film support force. S377. Subtract the known weight of the shaft from the total oil film support force to obtain a net force, and divide it by a preset bearing stiffness coefficient to obtain a displacement. Add this displacement to the original coordinates of the shaft center to obtain the theoretical new coordinates of the shaft center at the next moment. This new coordinate and the two-dimensional array of the final pressure value are sequentially spliced ​​together to form the theoretical physical state vector at the next moment.

[0028] 7. The bearing lubrication parameter adaptive control system based on reinforcement learning according to claim 2, wherein step S4 specifically includes: S41. Input the low-dimensional potential state vector at the current moment into the behavior network of the improved Dreamer model. The behavior network performs multi-step look-ahead planning in the world model and generates a number of candidate future action trajectories multiplied by the preset planning steps and the number of candidate actions per step. S42. In the planning and optimization process of the behavior network, a risk penalty term related to the prediction uncertainty measure is introduced. For each candidate future action trajectory, the prediction uncertainty measure obtained by the uncertainty network at each step in the preset planning steps is accumulated to obtain the total uncertainty of the current future action trajectory. S43. Multiply the total uncertainty of the current future action trajectory by a preset risk coefficient that is greater than zero to obtain the risk penalty term of the current future action trajectory. Add the physical consistency loss term of the current candidate future action trajectory to the risk penalty term to obtain the final risk penalty term. S44. Multiply the scalar instant reward value obtained by each candidate future action trajectory at each step in the preset planning steps by a preset discount factor that decreases with the number of steps, and sum them up to obtain a basic cumulative reward value. Subtract the final comprehensive risk penalty to obtain the final cumulative reward value. S45. The behavioral network searches and selects the candidate future action trajectory that maximizes the final cumulative reward value from all candidate future action trajectories, and uses the first candidate action vector as the final output, which is a two-dimensional action vector containing the target fuel supply pressure value and the target fuel supply frequency value.

[0029] In this embodiment, the multi-step look-ahead planning specifically includes: The number of planned steps and the number of candidate actions per step are preset. Starting from the current moment, for each step, the behavior network calls a random number generator to generate two independent random numbers that conform to the standard normal distribution. These two random numbers are multiplied by a preset exploration intensity coefficient that determines the size of the exploration range. These two scaled random numbers are then added to the base fuel supply pressure value and the base fuel supply frequency value, respectively. The range of the summed base oil supply pressure and base oil supply frequency is limited. If it exceeds the preset maximum value, it is set to the maximum value. If it is less than the minimum value, it is set to the minimum value. A candidate two-dimensional action vector is generated. The process is repeated, and a preset number of candidate action vectors are generated each time using a different random number. Each candidate action vector and the current low-dimensional latent state vector are input into the world model to obtain a preset number of predicted low-dimensional latent state vectors and scalar instant reward values, and to generate a preset number of future action trajectories multiplied by the preset number of planning steps and the number of candidate actions per step.

[0030] In this embodiment, S5 specifically includes: S51. Multiply the target oil supply pressure value in the two-dimensional motion vector by a preset pressure conversion coefficient to obtain the first control signal for driving the proportional pressure valve. Multiply the target oil supply frequency value in the two-dimensional motion vector by a preset frequency conversion coefficient to obtain the second control signal for driving the variable frequency pump. S52. After executing the first control signal and the second control signal, wait for a preset time interval, re-acquire the bearing housing temperature, subtract the bearing housing temperature before the control is executed from the newly acquired bearing housing temperature to obtain the instantaneous change in bearing housing temperature, perform a fast Fourier transform on the re-acquired vibration time domain signal, find the frequency point with the largest amplitude in the obtained vibration frequency domain features and read the amplitude as the main frequency amplitude of the vibration frequency domain features. S53. Read the real-time operating voltage and current of the proportional pressure valve and variable frequency pump in the lubrication system, and multiply the real-time operating voltage and real-time operating current to obtain the real-time power consumption of the lubrication system. S54. Multiply the instantaneous change in bearing housing temperature by a preset temperature weighting coefficient to obtain the first term; multiply the dominant frequency amplitude of the vibration frequency domain characteristic by a preset vibration weighting coefficient to obtain the second term; multiply the real-time power consumption of the lubrication system by a preset power consumption weighting coefficient to obtain the third term; and take the negative sum of the first, second, and third terms to generate the true reward value.

[0031] In this embodiment, S6 specifically includes: S61. Store the experience tuple consisting of state vector, two-dimensional action vector and real reward value into an experience replay pool. Randomly sample a batch of experience tuples from the experience replay pool. Input the state vector and two-dimensional action vector in each experience tuple into the world model of the improved Dreamer model to obtain a predicted low-dimensional latent state vector and scalar instantaneous reward value for the next time step. S63. Input the predicted low-dimensional latent state vector of the next time step into the behavior network of the improved Dreamer model, calculate a predicted final cumulative reward value, and add the predicted scalar instant reward value to the predicted final cumulative reward value to obtain a predicted total reward value. S64. Input the next-time state vector in each empirical tuple into the encoder network to obtain the true next-time low-dimensional latent state vector, and input it into the improved Dreamer model to calculate the true final cumulative reward value. Then add the true reward value to the true final cumulative reward value to obtain a true total reward value. S65. Calculate the squared Euclidean distance between the predicted total return value and the actual total return value to obtain a temporal difference error. Use the backpropagation algorithm and gradient descent algorithm to update the network weights of the improved Dreamer model based on the temporal difference error, and continuously perform self-optimization of the improved Dreamer model.

[0032] Example 1: To verify the feasibility of this invention in practice, it was applied to a large-scale precision gear machining workshop of a leading domestic heavy equipment manufacturing group. This workshop is responsible for producing key components such as high-precision marine gearboxes and wind turbine speed-increasing gears. It possesses 86 core processing equipment units, including five-axis CNC grinding machines and high-precision gear hobbing machines, with a total equipment value exceeding 320 million yuan. The spindle bearings of these machines all use oil-air lubrication systems, and their operating condition directly determines the machining accuracy and surface quality of the gears. A single bearing failure can result in direct economic losses exceeding 500,000 yuan and severely impact order delivery cycles.

[0033] The workshop environment is complex and variable, requiring frequent adjustments to processing parameters for different batches and materials of workpieces. Spindle speeds range from 500 rpm to 12,000 rpm, and cutting loads fluctuate dramatically from light to heavy. Traditional lubrication control relies primarily on PID controllers with fixed parameters, setting fixed oil supply intervals and quantities based on speed, which cannot adapt to dynamic load changes in actual processing. In actual production, problems frequently arise such as abnormally high bearing temperatures and increased vibration due to insufficient lubrication, or oil waste and oil mist contamination due to over-lubrication. The workshop's original lubrication control system recorded an average of 42 abnormal bearing temperature events and 28 vibration exceedance events per month, resulting in over 15 hours of unplanned downtime, severely hindering production efficiency.

[0034] In practical deployment, the method of this invention installs high-precision sensors at key locations on each piece of equipment to collect 12-dimensional operational data in real time, including spindle speed, bearing temperature, vibration acceleration, feed rate, and cutting power, with a sampling frequency of 1kHz. Simultaneously, static information such as equipment structural parameters, bearing model, and lubrication system specifications are incorporated into the model. After preprocessing, this multi-source heterogeneous data is used to construct a multi-dimensional state vector that comprehensively characterizes the bearing's operating state. The improved Dreamer model of this invention introduces a physical consistency loss term, embedding the shaft dynamics equations and fluid lubrication theory into the world model's learning process, ensuring that the prediction results conform to physical laws. Furthermore, by introducing a risk penalty term, it effectively avoids high-risk actions that could damage the bearing, improving system safety. Table 1 below shows the comparison data between the method of this invention and traditional PID control in precision grinding tasks during three months of continuous operation: Table 1. Performance Comparison Data of the Invention and Traditional PID Control in Precision Gear Machining

[0035] Based on the comparative data shown in Table 1, it can be seen that the bearing lubrication adaptive control method based on reinforcement learning proposed in this invention exhibits significant performance advantages over traditional PID control in precision machining scenarios, especially in terms of key indicators such as operational stability, economic benefits, and equipment reliability, which have all been comprehensively improved.

[0036] In terms of operational stability, this invention reduces the average bearing temperature by 9.1%, and the temperature standard deviation by a significant 56.3%. Peak vibration and root mean square vibration values ​​are both reduced by over 33%. This demonstrates that through forward-looking planning incorporating physical knowledge, the system effectively suppresses dynamic disturbances during processing, providing an extremely stable thermal environment for high-precision machining. Regarding economic benefits, this invention reduces daily lubricant consumption by 24.1% and unplanned downtime by 76.3% through on-demand lubrication, directly improving equipment utilization and output efficiency. The improvement in equipment reliability is particularly outstanding, with the average bearing lifespan extended by 44.5%, combined with a 2.5% increase in processing qualification rate, fully demonstrating the invention's superior ability to ensure long-term stable equipment operation and improve product quality. Overall, this invention effectively solves the core pain points of traditional control methods, such as poor dynamic adaptability, high energy consumption, and low reliability, showcasing significant industrial application value.

[0037] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A bearing lubrication parameter adaptive control system based on reinforcement learning, characterized in that, Includes the following modules: The bearing operating status acquisition module is used to simultaneously acquire data from multiple sources. The state vector construction module is used to preprocess multi-source data and combine them sequentially into state vectors; The world model learning module is used to input the state vector into the encoder network of the improved Dreamer model to obtain a low-dimensional latent state vector, and input the predicted low-dimensional latent state vector of the world model to predict the next time step. It outputs a prediction uncertainty metric and calculates the physical consistency loss term by decoding the predicted low-dimensional latent state vector and comparing it with the theoretical physical state vector of the next time step. The optimal action module is used to perform multi-step forward planning in the world model through the behavior network, generate multiple candidate future action trajectories, calculate the final cumulative reward value by subtracting the sum of the physical consistency loss term and the risk penalty term from the basic cumulative reward value, and output a two-dimensional action vector. The control execution and reward evaluation module is used to convert the two-dimensional action vector into a control signal, collect the instantaneous change of bearing housing temperature and the main frequency amplitude of vibration frequency domain characteristics, multiply the instantaneous change of bearing housing temperature, the main frequency amplitude of vibration frequency domain characteristics and the real-time power consumption of the lubrication system by the corresponding preset weight coefficients, sum them, and take the negative number of the sum to obtain the real reward value. The model self-optimization module stores an experience tuple consisting of a state vector, a two-dimensional action vector, and a true reward value into an experience replay pool. It updates the network weights of the improved Dreamer model by minimizing the temporal difference error, and continuously performs self-optimization of the improved Dreamer model.

2. The bearing lubrication parameter adaptive control system based on reinforcement learning according to claim 1, characterized in that, The modules are connected in the following way: S1. Through multi-source sensors, synchronously collect multi-source data such as vibration time-domain signal, bearing housing temperature, lubricating oil film pressure value and real-time spindle speed during bearing operation; S2. Perform fast Fourier transform on the vibration time domain signal to extract the vibration frequency domain features, normalize the bearing housing temperature, lubricating oil film pressure value and real-time spindle speed, and combine the processed vibration frequency domain features, temperature, pressure and speed data into a state vector in sequence. S3. Input the state vector into the encoder network of the improved Dreamer model to obtain a low-dimensional latent state vector, and input it into the world model to predict the predicted low-dimensional latent state vector of the next moment. Output the prediction uncertainty measure, and calculate the physical consistency loss term by comparing the decoded predicted state with the theoretical physical state vector of the next moment. S4. Through a behavioral network, perform multi-step forward planning in the world model to generate multiple candidate future action trajectories and calculate the final cumulative reward value by subtracting the sum of the physical consistency loss term and the risk penalty term from the basic cumulative reward value. Output a two-dimensional action vector that maximizes the final cumulative reward value and includes the target fuel supply pressure value and the target fuel supply frequency value. S5. Convert the two-dimensional motion vector into a control signal, collect the instantaneous change in bearing housing temperature and the main frequency amplitude of vibration frequency domain characteristics, multiply the instantaneous change in bearing housing temperature, the main frequency amplitude of vibration frequency domain characteristics and the real-time power consumption of the lubrication system by the corresponding preset weight coefficients and sum them, and take the negative value of the sum to obtain the real reward value. S6. Store the experience tuple consisting of the state vector, two-dimensional action vector and real reward value into an experience replay pool. Randomly sample a batch of experience tuples from the experience replay pool. Update the network weights of the improved Dreamer model by minimizing the temporal difference error, and continuously perform self-optimization of the improved Dreamer model.

3. The bearing lubrication parameter adaptive control system based on reinforcement learning according to claim 2, characterized in that, S1 specifically includes: S11. Physically install the vibration sensor, temperature sensor, oil film pressure sensor and speed sensor at the predetermined measuring points on the bearing housing or spindle, and start data acquisition from all sensors simultaneously through a synchronous trigger signal. S12. Synchronously read the vibration time-domain signal output by the vibration sensor, the bearing housing temperature output by the temperature sensor, the lubricating oil film pressure value output by the oil film pressure sensor, and the real-time spindle speed output by the speed sensor.

4. The bearing lubrication parameter adaptive control system based on reinforcement learning according to claim 2, characterized in that, S2 specifically includes: S21. Perform a fast Fourier transform on the vibration time-domain signal to convert the vibration time-domain signal from the time domain to the frequency domain, extract the amplitude spectrum in the frequency domain as the vibration frequency domain feature, read the bearing housing temperature, lubricating oil film pressure value and spindle real-time speed respectively, and use the preset maximum and minimum values ​​to perform linear normalization processing on each data item to map the numerical range to between 0 and 1. S22. The normalized bearing housing temperature, lubricating oil film pressure value and real-time spindle speed data are spliced ​​with the vibration frequency domain characteristics in a preset order to generate a state vector of preset dimensions.

5. The bearing lubrication parameter adaptive control system based on reinforcement learning according to claim 2, characterized in that, S3 specifically includes: S31. The state vector is input into the encoder network consisting of a preset number of fully connected layers of the improved Dreamer model. The first fully connected layer calculates the state vector with the preset first weight matrix and the first bias vector, and then performs a nonlinear transformation through the ReLU activation function to obtain the first layer encoding features. The second fully connected layer receives the first layer encoding features and continues to calculate the preset second weight matrix and the second bias vector. S32. The computation output dimension of the repeatedly stacked fully connected layers is a low-dimensional latent state vector with a preset number of fully connected layers. The policy network receives the low-dimensional latent state vector at the current time, calculates it through forward propagation, multiplies the low-dimensional latent state vector with a preset weight matrix, adds a preset bias vector, and processes it through a hyperbolic tangent activation function to obtain two basic action vectors with values ​​between negative one and one. S33. Read the minimum and maximum values ​​of the preset target oil supply pressure and target oil supply frequency, and linearly map the first value of the basic motion vector from the range of negative one to one to the range of the minimum to maximum value of the target oil supply pressure and target oil supply frequency to obtain a basic oil supply pressure value and a basic oil supply frequency value. Repeat all numerical calculations on the basic motion vector to generate a two-dimensional motion vector with actual physical meaning. S34. Input the low-dimensional latent state vector into the preset world model of the improved Dreamer model. The world model consists of a state transition network, a reward network and an uncertainty network. The state transition network receives the low-dimensional latent state vector at the current time and the two-dimensional action vector executed at the previous time. It is calculated through a gated recurrent unit network and outputs the predicted low-dimensional latent state vector for the next time. S35. The reward network receives the low-dimensional latent state vector and the two-dimensional action vector at the current time and combines them into a long vector. It calculates and outputs a scalar instant reward value through a fully connected layer with preset weights and biases. The uncertainty network receives the low-dimensional latent state vector and the two-dimensional action vector at the current time and concatenates them into a long vector. It calculates and outputs a prediction uncertainty measure through another fully connected layer with preset weights and biases. S36. In the training process of the world model, a physical consistency loss term is introduced. The physical consistency loss term inputs the low-dimensional latent state vector predicted by the world model in the next time step into the decoder network, which consists of two layers and a preset fully connected layer, to restore the low-dimensional latent state vector to the predicted physical state vector. S37. Based on the bearing's shaft center trajectory equation and Reynolds equation, calculate a theoretical physical state vector for the next moment using numerical methods. Calculate the squared Euclidean distance between the predicted physical state vector and the theoretical physical state vector for the next moment, and use it as the physical consistency loss term.

6. The bearing lubrication parameter adaptive control system based on reinforcement learning according to claim 5, characterized in that, Specifically, S37 includes: S371. Create a two-dimensional array to store the pressure at each point on the oil film grid. Initialize the values ​​at all positions to zero and enter a loop. The loop continues until the preset number of loop stops is met. In each iteration of the loop, traverse each position of this two-dimensional array and read the pressure values ​​at the four adjacent positions above, below, to the left, and to the right of the grid position currently being calculated. S372. Obtain the preset bearing radius clearance value and the offset of the shaft center in the vertical direction, and record the angle corresponding to the current grid point on the circumference. Subtract the product of the offset and the cosine of the angle from the radius clearance value to obtain the oil film thickness of the current grid point. S373. Read the current spindle speed value and multiply it by a constant of sixty to convert the unit from minutes to seconds, and multiply it by the product of pi and the preset journal radius to obtain the linear velocity of the journal surface. S374. Subtract the left pressure value from the right pressure value of the current grid point read in the two-dimensional array to obtain the circumferential pressure difference. Subtract the lower pressure value from the upper pressure value to obtain an axial pressure difference. Multiply the circumferential pressure difference by the preset circumferential grid spacing coefficient to obtain the first term. Multiply the axial pressure difference by the preset axial grid spacing coefficient to obtain the second term. Multiply the linear velocity by the preset lubricating oil viscosity and divide by the square of the oil film thickness to obtain the third term. Add the first, second and third terms together to obtain the total, which is the pressure update term. S375. Multiply the pressure update item by a preset relaxation coefficient greater than 0 and less than 1 to obtain the adjustment amount. Add this adjustment amount to the original pressure value at the current grid position. The new value is the updated pressure, which overwrites the original value. S376. After updating the pressure value for all positions in the two-dimensional array, a global iteration is completed. Check the difference between the old and new pressure values ​​for all positions. If the largest difference is less than the preset difference threshold, stop the loop. Otherwise, repeat the next global iteration. After the loop stops, accumulate all the pressure values ​​for all positions in the two-dimensional array and multiply them by the area represented by a single grid to obtain the total oil film support force. S377. Subtract the known weight of the shaft from the total oil film support force to obtain a net force, and divide it by a preset bearing stiffness coefficient to obtain a displacement. Add this displacement to the original coordinates of the shaft center to obtain the theoretical new coordinates of the shaft center at the next moment. This new coordinate and the two-dimensional array of the final pressure value are sequentially spliced ​​together to form the theoretical physical state vector at the next moment.

7. The bearing lubrication parameter adaptive control system based on reinforcement learning according to claim 2, characterized in that, S4 specifically includes: S41. Input the low-dimensional potential state vector at the current moment into the behavior network of the improved Dreamer model. The behavior network performs multi-step look-ahead planning in the world model and generates a number of candidate future action trajectories multiplied by the preset planning steps and the number of candidate actions per step. S42. In the planning and optimization process of the behavior network, a risk penalty term related to the prediction uncertainty measure is introduced. For each candidate future action trajectory, the prediction uncertainty measure obtained by the uncertainty network at each step in the preset planning steps is accumulated to obtain the total uncertainty of the current future action trajectory. S43. Multiply the total uncertainty of the current future action trajectory by a preset risk coefficient that is greater than zero to obtain the risk penalty term of the current future action trajectory. Add the physical consistency loss term of the current candidate future action trajectory to the risk penalty term to obtain the final risk penalty term. S44. Multiply the scalar instant reward value obtained by each candidate future action trajectory at each step in the preset planning steps by a preset discount factor that decreases with the number of steps, and sum them up to obtain a basic cumulative reward value. Subtract the final comprehensive risk penalty to obtain the final cumulative reward value. S45. The behavioral network searches and selects the candidate future action trajectory that maximizes the final cumulative reward value from all candidate future action trajectories, and uses the first candidate action vector as the final output, which is a two-dimensional action vector containing the target fuel supply pressure value and the target fuel supply frequency value.

8. The bearing lubrication parameter adaptive control system based on reinforcement learning according to claim 7, characterized in that, The multi-step forward planning specifically includes: The number of planned steps and the number of candidate actions per step are preset. Starting from the current moment, for each step, the behavior network calls a random number generator to generate two independent random numbers that conform to the standard normal distribution. These two random numbers are multiplied by a preset exploration intensity coefficient that determines the size of the exploration range. These two scaled random numbers are then added to the base fuel supply pressure value and the base fuel supply frequency value, respectively. The range of the summed base oil supply pressure and base oil supply frequency is limited. If it exceeds the preset maximum value, it is set to the maximum value. If it is less than the minimum value, it is set to the minimum value. A candidate two-dimensional action vector is generated. The process is repeated, and a preset number of candidate action vectors are generated each time using a different random number. Each candidate action vector and the current low-dimensional latent state vector are input into the world model to obtain a preset number of predicted low-dimensional latent state vectors and scalar instant reward values, and to generate a preset number of future action trajectories multiplied by the preset number of planning steps and the number of candidate actions per step.

9. The bearing lubrication parameter adaptive control system based on reinforcement learning according to claim 2, characterized in that, S5 specifically includes: S51. Multiply the target oil supply pressure value in the two-dimensional motion vector by a preset pressure conversion coefficient to obtain the first control signal for driving the proportional pressure valve. Multiply the target oil supply frequency value in the two-dimensional motion vector by a preset frequency conversion coefficient to obtain the second control signal for driving the variable frequency pump. S52. After executing the first control signal and the second control signal, wait for a preset time interval, re-acquire the bearing housing temperature, subtract the bearing housing temperature before the control is executed from the newly acquired bearing housing temperature to obtain the instantaneous change in bearing housing temperature, perform a fast Fourier transform on the re-acquired vibration time domain signal, find the frequency point with the largest amplitude in the obtained vibration frequency domain features and read the amplitude as the main frequency amplitude of the vibration frequency domain features. S53. Read the real-time operating voltage and current of the proportional pressure valve and variable frequency pump in the lubrication system, and multiply the real-time operating voltage and real-time operating current to obtain the real-time power consumption of the lubrication system. S54. Multiply the instantaneous change in bearing housing temperature by a preset temperature weighting coefficient to obtain the first term; multiply the dominant frequency amplitude of the vibration frequency domain characteristic by a preset vibration weighting coefficient to obtain the second term; multiply the real-time power consumption of the lubrication system by a preset power consumption weighting coefficient to obtain the third term; and take the negative sum of the first, second, and third terms to generate the true reward value.

10. The bearing lubrication parameter adaptive control system based on reinforcement learning according to claim 2, characterized in that, S6 specifically includes: S61. Store the experience tuple consisting of state vector, two-dimensional action vector and real reward value into an experience replay pool. Randomly sample a batch of experience tuples from the experience replay pool. Input the state vector and two-dimensional action vector in each experience tuple into the world model of the improved Dreamer model to obtain a predicted low-dimensional latent state vector and scalar instantaneous reward value for the next time step. S63. Input the predicted low-dimensional latent state vector of the next time step into the behavior network of the improved Dreamer model, calculate a predicted final cumulative reward value, and add the predicted scalar instant reward value to the predicted final cumulative reward value to obtain a predicted total reward value. S64. Input the next-time state vector in each empirical tuple into the encoder network to obtain the true next-time low-dimensional latent state vector, and input it into the improved Dreamer model to calculate the true final cumulative reward value. Then add the true reward value to the true final cumulative reward value to obtain a true total reward value. S65. Calculate the squared Euclidean distance between the predicted total return value and the actual total return value to obtain a temporal difference error. Use the backpropagation algorithm and gradient descent algorithm to update the network weights of the improved Dreamer model based on the temporal difference error, and continuously perform self-optimization of the improved Dreamer model.