System and method for prolonging operation life of motor based on fuzzy control

Through a system based on fuzzy control, combined with reinforcement learning and deep learning technology, the motor lubrication cycle is adjusted in real time, and the problem of insufficient lubrication accuracy in the existing technology is solved, achieving the extension of the motor operation life and the improvement of the system's adaptability.

CN120090526AActive Publication Date: 2025-06-03EAST CHINA JIAOTONG UNIVERSITY

Patent Information

Application Number
CN202510084778.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2025-06-03
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

The prior art cannot realize real-time monitoring of the motor lubrication status and on-demand filling, resulting in insufficient lubrication accuracy and affecting the motor operation life.

Method used

The system based on fuzzy control is adopted, combining the environment and state space module, reinforcement learning module, deep learning and fuzzy reasoning integration module, fuzzy rules and membership function optimization module, feedback and real-time control module, and dynamically adjust the lubrication cycle by collecting temperature and vibration data in real time to achieve accurate lubrication.

Benefits of technology

It improves lubrication accuracy and efficiency, extends the operating life of the motor, and enhances the adaptability and accuracy of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120090526A_ABST
    Figure CN120090526A_ABST
Patent Text Reader

Abstract

The invention discloses a system and method for prolonging the operation life of a motor based on fuzzy control, and the method comprises the steps: building a state space, an action space and a reward function, and defining an optimization target as an optimal lubrication period; a Q-learning algorithm is used, an optimal strategy is learned through interaction with the environment, and the lubrication period is adjusted; approaching a Q function value by using a deep Q network, and automatically extracting a fuzzy control rule and a membership function through a neural network; automatically learning a nonlinear relation between the temperature and the vibration difference value through a neural network, and adjusting a fuzzy rule and a membership function shape; and the lubrication period is dynamically adjusted according to the real-time state, and self-adaptive control is achieved. By acquiring and processing related information of a motor lubricating point, lubricating grease required by the motor during working is deduced, and the lubricating point is filled with the corresponding lubricating grease, so that precise lubrication of the motor is realized, frictional resistance is reduced, and the service life of the motor is prolonged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field:

[0001] The present invention belongs to the technical field of intelligent control, and particularly relates to a system and method for prolonging the service life of an electric motor based on fuzzy control. Background Art:

[0002] During the operation of an electric motor, friction occurs between its internal moving parts. This friction not only causes energy loss but also accelerates component wear. By adding grease, the friction coefficient between them can be significantly reduced. The traditional method of adding grease is manual lubrication, which cannot achieve continuous and efficient lubrication operations. Insufficient grease may lead to poor lubrication effects, thus shortening the service life of the electric motor. Excessive grease will also increase the resistance inside the bearing, instead reducing the operating efficiency and easily causing electric motor failures. Compared with manual lubrication, intelligent lubrication can indeed add grease more precisely and efficiently to a certain extent. However, the setting of its lubrication cycle still depends on manual operation, and it cannot achieve real-time monitoring of the lubrication state of the electric motor and refueling on demand. Therefore, the lubrication accuracy still needs to be improved.

[0003] Chinese Patent CN118009023A discloses a "distributed lubrication system". This system uploads data such as the grease quantity and battery power of the lubricator to the cloud through the NB-lot module, and the lubrication cycle can be remotely modified by issuing a modification cycle instruction through the cloud. However, this system can only be modified through a pre-set cycle. Once the lubrication accuracy is low, precise lubrication cannot be achieved, and inaccurate cycle setting may affect the service life of the electric motor.

[0004] Chinese Patent CN108847282B discloses a "system and method for expert experience reasoning based on fuzzy inference". This system calculates the disease information corresponding to the user's symptom information using fuzzy inference according to the rules in the expert experience rule library. This system does relieve the problems of low disease judgment accuracy and low triage efficiency to a certain extent. However, over-reliance on expert experience also brings limitations. Facing a dynamically changing system, it may be difficult to update the rules or adjust the inference results in a timely manner, thus causing delays.

[0005] In order to cope with the dynamically changing operating state of the electric motor, improve the lubrication accuracy, and prolong the service life of the electric motor, it is an urgent problem to introduce new technologies and methods to correct the lubrication cycle in real time according to the working state of the electric motor, so as to add grease more precisely.

[0006] The information disclosed in this background art section is only intended to increase the overall understanding of the present invention and should not be regarded as an admission or any form of implication that this information constitutes prior art already known to those of ordinary skill in the art. Summary of the Invention:

[0007] The purpose of the present invention is to provide a system and method for extending the operating life of a motor based on fuzzy control, so as to overcome the defects in the above-mentioned prior art.

[0008] To achieve the above object, the present invention provides a system for extending the operating life of a motor based on fuzzy control, including an environment and state space module, a reinforcement learning module, an integrated deep learning and fuzzy inference module, a fuzzy rule and membership function optimization module, and a feedback and real-time control module; the environment and state space module is used to define and simulate the state space for interacting with the actual environment, collect temperature and vibration sensor data, and evaluate the current system state based on these data; the reinforcement learning module is trained using a reinforcement learning algorithm to learn an adjustment strategy for selecting a lubrication period according to different states; the integrated deep learning and fuzzy inference module is used to automatically extract fuzzy rules and membership functions, improve the accuracy of fuzzy inference, and provide more accurate input information for fuzzy control; the fuzzy rule and membership function optimization module is used to optimize the fuzzy rules and the shape of the membership functions to ensure that the adjustment of the lubrication period is more accurate and effective, and further refine the fuzzy control; the feedback and real-time control module is used to dynamically adjust the lubrication period according to the collected temperature and vibration sensor data and the system state, so that the lubrication period is highly consistent with the actual operating environment.

[0009] A method for extending the operating life of a motor based on fuzzy control, the steps of which are as follows:

[0010] S01. Collect the temperature value and vibration value at the motor lubrication point through sensors, and upload them to the cloud through a 4G module. Establish a state space, an action space, and a reward function in the cloud, and define the optimization goal as the optimal lubrication period;

[0011] S02. Use the Q-learning algorithm to evaluate the quality of the reward function variables through the Q function, and learn the optimal strategy and adjust the lubrication period by continuously updating the Q function value; the reward function variables are state-action pairs;

[0012] S03. Use a deep Q network to approximate the Q function, and automatically extract fuzzy rules and membership functions through a neural network;

[0013] S04. The neural network automatically learns the non-linear relationship between the temperature and vibration differences, and adjusts the shapes of the fuzzy rules and membership functions;

[0014] S05. The system dynamically adjusts the lubrication period according to the state information obtained in real time, where the state information includes the temperature difference and the vibration difference, and issues a period command to the lubricator, and the lubricator injects oil according to the received period command.

[0015] Preferably, in the technical solution, in step S01, the state space is various types of information of the system at present, which is used to describe the current situation of the system. The state space is defined as the following set:

[0016] S = {TempDiff, VibChange, PrevAction} (1),

[0017] where S is the state space; TempDiff is the temperature difference, representing the difference between the current temperature and the target temperature; VibChange is the vibration change amount, representing the change amount of the current vibration signal; PrevAction is the previous control output, representing the adjustment in the previous lubrication cycle;

[0018] The action space defines the control behaviors that the system can choose. The action space is defined as the following set:

[0019] A = {Increase lubrication cycle, Decrease lubrication cycle, Remain unchanged} (2),

[0020] where A is the action space; the control behaviors include increasing the lubrication cycle, decreasing the lubrication cycle, and remaining unchanged;

[0021] The reward function is used to measure the effect of each control behavior. The reward function is defined as:

[0022]

[0023] where R() is the reward function; s t represents the current state, and the current state includes the temperature value and the vibration value; a t represents the action taken by the system in state s t ; R() = +10 represents high-temperature and high-vibration adjustment, R() = -10 represents low-temperature and low-vibration adjustment, and R() = 0 represents ineffective adjustment.

[0024] Preferably, in the technical solution, in step S02, the Q-function formula is expressed as:

[0025]

[0026] where Q(s t , a t ) represents the expected return after taking action a t in state s t ; s 0 represents the initial state; a 0 represents the initial action taken by the system in state s 0 ; r t represents the reward signal, which measures the immediate return after taking a certain action; γ represents the discount factor, which determines the weight of future rewards.

[0027] Preferably, in the technical solution, in step S03, the deep Q-network approximates the Q-function in the Q-learning algorithm through a deep neural network:

[0028] Q(s t ,a t ,θ)≈Q * (s t ,a t )(5),

[0029] where θ is the neural network parameter; Q() represents the Q-function value output by the deep Q-network; Q * () represents the optimal Q-function value output by the deep Q-network;

[0030] The deep Q-network makes the current Q-function value close to the true target Q-function value and minimizes the error of the Q-function:

[0031]

[0032] where L(θ) represents the loss function, which is used to measure the gap between the Q-function value output by the current deep Q-network and the target Q-function value; E[] represents the expected value; is the optimal action Q-function value in the next state; a′ represents the next action in the next state; r t+1 is the reward after executing the action a t in the current state s t , Q(s t ,a t ; θ) is the Q-function value of the current policy, and θ - is the target neural network parameter.

[0033] Preferably, in the technical solution, the training process of the deep Q-network includes the following steps:

[0034] S1. Randomly initialize the neural network parameter θ of the deep Q-network and the target neural network parameter θ - ;

[0035] S2. Through the experience replay mechanism of the deep Q-network, the system executes actions in the environment and stores the experience in the experience pool, and then randomly samples a batch of data from the experience pool for training; the experience includes the state s t , the action a t , the reward r t+1 , the next state s t+1 ;

[0036] S3. After sampling a batch of data from the experience pool each time, update the Q-function value of the deep Q-network according to the Q-function value output by the current deep Q-network and the target Q-function value calculated by the target neural network; calculate the loss through backpropagation and optimize the neural network parameters θ to approximate the true Q-function value. The update formula for the Q-function value is as follows:

[0037]

[0038] where a is the learning rate, which controls the update step size;

[0039] S4. Use the target neural network. Every certain number of steps, the target neural network parameters θ - will be updated to the neural network parameters θ of the current deep Q-network, so as to ensure that the deep Q-network will not oscillate due to unstable target values during the training process.

[0040] Preferably, in the technical solution, in step S04, the system will first automatically generate a fuzzy rule, and through the combination of reinforcement learning and deep learning, dynamically optimize the fuzzy rule, membership function, and controller parameters.

[0041] The membership function is used to map each element in the fuzzy set to the membership degree in [0, 1], indicating the degree of a certain input in a certain fuzzy set; where the fuzzy set is the set formed by the fuzzy rules;

[0042] During the automatic optimization process, through the neural network to learn the non-linear mapping between the input value and the membership function, where the input value includes the temperature difference and vibration difference, automatically adjust the shape of the membership function, so that the system can process complex inputs more accurately;

[0043] The system adjusts the controller parameters according to the Q-learning algorithm through the interaction with the environment, where the interaction is the operating state and lubrication effect of the device. When the system finds that the control strategy of "increasing the lubrication period when the temperature difference is large" can reduce the incidence of equipment failures, the reinforcement learning model will adjust the weight of the fuzzy rule according to this feedback signal, so that the system is more likely to execute the same action when encountering similar situations in the future; through the combination of deep learning and reinforcement learning, the system can automatically extract fuzzy rules from the temperature and vibration sensor data, optimize the membership function, and continuously adjust the weight of the fuzzy rule and the shape of the membership function.

[0044] Compared with the prior art, the present invention has the following beneficial effects:

[0045] By introducing reinforcement learning, such as Q-learning and deep Q-network, through reinforcement learning and deep learning techniques, the adaptive ability and accuracy of the system are continuously improved. In this process, reinforcement learning is used to optimize the control strategy, while deep learning is used to optimize the system model and fuzzy rules, enabling the system to automatically adapt to different environmental conditions and provide more accurate and efficient lubrication cycle control. Brief Description of the Drawings:

[0046] Figure 1 It is a system principle block diagram of an invention for extending the operating life of a motor based on fuzzy control;

[0047] Figure 2 It is a method flowchart of an invention for extending the operating life of a motor based on fuzzy control. Detailed Embodiments:

[0048] The following describes the detailed embodiments of the present invention in detail, but it should be understood that the protection scope of the present invention is not limited by the detailed embodiments.

[0049] Unless otherwise clearly stated, throughout the specification and claims, the term "comprising" or its variations such as "including" or "having" etc. will be understood to include the stated elements or components, without excluding other elements or other components.

[0050] As Figure 1 shown, a system for extending the operating life of a motor based on fuzzy control includes an environment and state space module, a reinforcement learning module, an integrated deep learning and fuzzy inference module, a fuzzy rule and membership function optimization module, and a feedback and real-time control module; the environment and state space module is used to define and simulate the state space for interacting with the actual environment, collect temperature and vibration sensor data, and evaluate the current system state based on these data; the reinforcement learning module is trained using reinforcement learning algorithms to learn adjustment strategies for selecting lubrication cycles according to different states; the integrated deep learning and fuzzy inference module is used to automatically extract fuzzy rules and membership functions to improve the accuracy of fuzzy inference and provide more accurate input information for fuzzy control; the fuzzy rule and membership function optimization module is used to optimize the fuzzy rules and the shape of the membership functions to ensure more precise and effective adjustment of the lubrication cycle and further refine the fuzzy control; the feedback and real-time control module is used to dynamically adjust the lubrication cycle according to the collected temperature and vibration sensor data and the system state to make the lubrication cycle highly consistent with the actual operating environment.

[0051] As Figure 2 shown, a method for extending the operating life of a motor based on fuzzy control has the following steps:

[0052] S01. Collect the temperature value and vibration value at the motor lubrication point through sensors, and upload them to the cloud through a 4G module. Establish a state space, an action space, and a reward function in the cloud, and define the optimization goal as the optimal lubrication period;

[0053] The state space is various information of the system at present, which is used to describe the current situation of the system. The state space is defined as the following set:

[0054] S = {TempDiff, VibChange, PrevAction} (1),

[0055] where S is the state space; TempDiff is the temperature difference, which represents the difference between the current temperature and the target temperature; VibChange is the vibration change amount, which represents the change amount of the current vibration signal; PrevAction is the previous control output, which represents the adjustment of the previous lubrication period;

[0056] The action space defines the control behaviors that the system can choose. The action space is defined as the following set:

[0057] A = {Increase lubrication period, Decrease lubrication period, Remain unchanged} (2),

[0058] where A is the action space; the control behaviors include increasing the lubrication period, decreasing the lubrication period, and remaining unchanged;

[0059] The reward function is used to measure the effect of each control behavior. The reward function is defined as:

[0060]

[0061] where R() is the reward function; s t represents the current state, and the current state includes the temperature value and the vibration value; a t represents the action taken by the system in the s t state; if the temperature difference and vibration difference of the equipment decrease after increasing the lubrication period and the equipment operation becomes more stable, the reward function is defined as +10; if the temperature and vibration difference of the equipment increase after decreasing the lubrication period and the equipment operation becomes unstable, the reward function is defined as -10; when the period remains unchanged, the reward function is defined as 0;;

[0062] S02. Use the Q-learning algorithm to evaluate the quality of the reward function variable through the Q function, learn the optimal strategy by continuously updating the Q function value, and adjust the lubrication period; the reward function variable is the state-action pair;

[0063] The Q function formula is expressed as:

[0064]

[0065] Among them, Q(s t , a t ) represents the expected return after taking action a t in state s t ; s 0 represents the initial state; a 0 represents the initial action taken by the system in state s 0 ; r t represents the reward signal, measuring the immediate return after taking a certain action; γ represents the discount factor, determining the weight of future rewards;

[0066] Use a table to store all possible state and action pairs. This table is the Q-table. Initialize the Q-table, and set all values in the Q-table to 0, indicating that we have no knowledge about each state-action pair yet; The system obtains state information by interacting with the environment and selects an action. The actions of the system may be increasing lubrication, reducing lubrication, or not adjusting; After each action, the environment gives feedback;

[0067] In state s 0 , the system selects "increase lubrication cycle", the system feedback reward is +10, and according to the Q-function update formula, update the Q-function; Assume that at this time, both the temperature difference and the vibration difference have changed. The system selects an action according to the new state and obtains the corresponding reward; After multiple rounds of training, the system learns to select to increase or decrease the lubrication cycle under specific temperature differences and vibration differences to optimize the equipment performance;

[0068] S03. Use a deep Q-network to approximate the Q-function and automatically extract fuzzy rules and membership functions through a neural network;

[0069] The deep Q-network approximates the Q-function in the Q-learning algorithm through a deep neural network:

[0070] Q(s t , a t , θ) ≈ Q * (s t , a t )(5),

[0071] Among them, θ is the neural network parameter; Q() represents the Q-function value output by the deep Q-network; Q * () represents the optimal Q-function value output by the deep Q-network;

[0072] The deep Q-network makes the current Q-function value close to the true target Q-function value and minimizes the error of the Q-function:

[0073]

[0074] Among them, \(L(\theta)\) represents the loss function, which is used to measure the gap between the Q-function value output by the current deep Q-network and the target Q-function value; \(E[]\) represents the expected value; is the Q-function value of the optimal action in the next state; \(a'\) represents the next action in the next state; \(r\) t+1 is the current state \(s\) t executes the action \(a\) t and the reward after that, \(Q(s\) t , \(a\) t ; \(\theta)\) is the Q-function value of the current policy, \(\theta\) - is the target neural network parameter;

[0075] The training process of the deep Q-network includes the following steps:

[0076] S1. Randomly initialize the neural network parameter \(\theta\) of the deep Q-network and the target neural network parameter \(\theta\) - ;

[0077] S2. Through the experience replay mechanism of the deep Q-network, the system executes actions in the environment and stores the experiences in the experience pool, and then randomly samples a batch of data from the experience pool for training; the experiences include the state \(s\) t , the action \(a\) t , the reward \(r\) t+1 , the next state \(s\) t+1 ;

[0078] S3. Each time a batch of data is sampled from the experience pool, update the Q-function value of the deep Q-network according to the Q-function value output by the current deep Q-network and the target Q-function value calculated by the target neural network; calculate the loss through backpropagation and optimize the neural network parameter \(\theta\) to approximate the true Q-function value. The update formula of the Q-function value is:

[0079]

[0080] Among them, \(\alpha\) is the learning rate, which controls the update step size;

[0081] S4. Adopt the target neural network. Every certain number of steps, the target neural network parameter \(\theta\) - will be updated to the neural network parameter \(\theta\) of the current deep Q-network, so as to ensure that the deep Q-network will not oscillate due to unstable target values during the training process;

[0082] S04. The neural network automatically learns the non-linear relationship between temperature and vibration difference, and adjusts the fuzzy rules and membership function shapes;

[0083] In the initial state, the system will first automatically generate a fuzzy rule, and through the combination of reinforcement learning and deep learning, dynamically optimize the fuzzy rules, membership functions, and controller parameters;

[0084] Suppose the generated fuzzy rules are: Rule 1, if the temperature difference is small and the vibration difference is small, then maintain the current lubrication period; Rule 2, if the temperature difference is large and the vibration difference is small, then reduce the lubrication period; Rule 3, if the temperature difference is small and the vibration difference is large, then increase the lubrication period;

[0085] By continuously adjusting the fuzzy rule weights, optimizing the selection of existing fuzzy rules, and gradually learning more reasonable fuzzy rules, the difficulty of manually setting fuzzy rules is reduced. If the possibility of a certain fuzzy rule causing equipment failure decreases, then the weight of that fuzzy rule will be increased, and vice versa;

[0086] The membership function is used to map each element in the fuzzy set to the membership degree in [0, 1], indicating the degree of a certain input in a certain fuzzy set; where the fuzzy set is the set formed by fuzzy rules;

[0087] Initialize the membership function, define the triangle as the initial membership function. Suppose the temperature difference varies within the range of [0, 100], and the vibration difference varies within [0, 10]. The low temperature is defined as a triangle within the range of [0, 0, 50], the medium temperature is defined as a triangle of [40, 50, 60], and the high temperature is defined as a triangle of [50, 100, 100]. Through the training of the deep neural network, the system will adjust the shapes of these membership functions. When the temperature difference is large, the peak of the membership function automatically shifts to the "high temperature" interval; when the vibration difference is large, the peak of the membership function biases towards the "high vibration" interval. Through the reward signal in reinforcement learning, the neural network parameters will be optimized by the backpropagation algorithm. Here, the neural network parameters include the shape of the membership function and the peak position;

[0088] Define the controller parameters. Each fuzzy rule has a weight, indicating the importance of that fuzzy rule in the decision-making. The weight of Rule 1 is 0.3, the weight of Rule 2 is 0.5, and the weight of Rule 3 is 0.2; The system adjusts the weight of each rule according to the reward signal. For example, when Rule 2 improves the operating efficiency of the equipment or reduces wear, the system will increase the weight of this rule, and vice versa;

[0089] Suppose the system executes the action of "reducing the lubrication period" in the state of "temperature difference = 50, vibration difference = 5". The system obtains a reward based on the result evaluation. The equipment wear is reduced and the operation is more stable, and the obtained reward is +10. Through the Q-function value update, the system will adjust the weight of the corresponding rule such as Rule 2, making the system more inclined to select this rule in a similar state;

[0090] S05. The system dynamically adjusts the lubrication period according to the state information obtained in real time. Here, the state information includes the temperature difference and the vibration difference, and issues the period instruction to the lubricator. The lubricator injects oil according to the received period instruction.

[0091] The foregoing description of the specific exemplary embodiments of the present invention is for purposes of illustration and exemplification. These descriptions are not intended to limit the invention to the precise forms disclosed, and it is apparent that many modifications and variations are possible in light of the above teachings. The purpose of selecting and describing the exemplary embodiments is to explain the specific principles of the present invention and its practical applications, so that those skilled in the art can implement and utilize the various different exemplary embodiments of the present invention, as well as various different selections and modifications. The scope of the present invention is intended to be defined by the claims and their equivalents.

Claims

1. A system for extending the life of a motor based on fuzzy control, characterized in that: It includes an environment and state space module, a reinforcement learning module, a deep learning and fuzzy reasoning integration module, a fuzzy rule and membership function optimization module, and a feedback and real-time control module; the environment and state space module is used to define the state space that interacts with the simulated actual environment, collect temperature and vibration sensor data, and evaluate the current system state based on these data; The reinforcement learning module is trained using a reinforcement learning algorithm to learn how to select an adjustment strategy for the lubrication cycle according to different states; the deep learning and fuzzy reasoning integration module is used to automatically extract fuzzy rules and membership functions, improve the accuracy of fuzzy reasoning, and provide more accurate input information for fuzzy control; the fuzzy rule and membership function optimization module is used to optimize the fuzzy rules and membership function shapes to ensure that the adjustment of the lubrication cycle is more accurate and effective, and to further refine the fuzzy control; the feedback and real-time control module is used to dynamically adjust the lubrication cycle based on the collected temperature, vibration sensor data and system status, so that the lubrication cycle is highly consistent with the actual operating environment.

2. A method for extending the life of a motor based on fuzzy control according to claim 1, comprising the following steps: S01, collect the temperature and vibration values ​​of the motor lubrication points through sensors, and upload them to the cloud through the 4G module, establish the state space, action space and reward function in the cloud, and define the optimization goal as the optimal lubrication cycle; S02, using the Q-learning algorithm, evaluate the quality of the reward function variable through the Q function, learn the optimal strategy and adjust the lubrication cycle by continuously updating the Q function value; the reward function variable is the state-action pair; S03, using a deep Q network to approximate the Q function, and automatically extracting fuzzy rules and membership functions through a neural network; S04, the neural network automatically learns the nonlinear relationship between temperature and vibration difference, and adjusts the fuzzy rules and the shape of the membership function; S05. The system dynamically adjusts the lubrication cycle according to the status information obtained in real time, where the status information includes temperature difference and vibration difference, and sends the cycle instruction to the lubricator. The lubricator performs oil filling according to the received cycle instruction.

3. The method of a system for extending the operating life of a motor based on fuzzy control according to claim 2, characterized in that: In step S01, the state space is the current information of the system, which is used to describe the current status of the system. The state space is defined as the following set: S={TempDiff,VibChange,PrevAction} (1), Where S is the state space; TempDiff is the temperature difference, which indicates the difference between the current temperature and the target temperature; VibChange is the vibration change, which indicates the change of the current vibration signal; PrevAction is the control output at the previous moment, which indicates the adjustment of the previous lubrication cycle; The action space defines the control behaviors that the system can choose. The action space is defined as the following set: A = {increase lubrication cycle, reduce lubrication cycle, remain unchanged} (2), Among them, A is the action space; the control behaviors include increasing the lubrication cycle, reducing the lubrication cycle, and keeping it unchanged; The reward function is used to measure the effect of each control behavior. The reward function is defined as: Among them, R() is the reward function; s t Indicates the current state, which includes temperature value and vibration value; a t Indicates that the system is in s t Action taken in the state; R()=+10 indicates high temperature and high vibration adjustment, R()=-10 indicates low temperature and low vibration adjustment, and R()=0 indicates invalid adjustment.

4. The method of a system for extending the operating life of a motor based on fuzzy control according to claim 2, characterized in that: In step S02, the Q function formula is expressed as: Among them, Q(s t ,a t ) means in state s t Take action a t The expected return after the system is completed; s0 represents the initial state; a0 represents the initial action taken by the system in the s0 state; r t represents the reward signal, which measures the immediate return after taking an action; γ represents the discount factor, which determines the weight of future rewards.

5. The method of a system for extending the operating life of a motor based on fuzzy control according to claim 2, characterized in that: In step S03, the deep Q network approximates the Q function in the Q-learning algorithm through a neural network: Q(s t ,a t ;θ)≈Q * (s t ,a t )(5), Among them, θ is the neural network parameter; Q() represents the Q function value output by the deep Q network; Q * () represents the optimal Q function value output by the deep Q network; The deep Q network makes the current Q function value close to the true target Q function value and minimizes the error of the Q function: Where L(θ) represents the loss function, which is used to measure the gap between the Q function value output by the current deep Q network and the target Q function value; E[] represents the expected value; is the optimal action Q function value in the next state; a′ represents the next action in the next state; r t+1 is the current state t Execute action a t The reward after, Q(s t ,a t ; θ) is the Q function value of the current strategy, θ - are the target neural network parameters.

6. The method of a system for extending the operating life of a motor based on fuzzy control according to claim 5, characterized in that: The training process of a deep Q-network consists of the following steps: S1, randomly initialize the neural network parameters θ of the deep Q network, and the target neural network parameters θ - ; S2. Through the experience replay mechanism of the deep Q network, the system performs actions in the environment and stores the experience in the experience pool, and then randomly samples a batch of data from the experience pool for training; the experience includes the state s t 、Action a t , Reward t+1 , the next state s t+1 ; S3. Each time a batch of data is sampled from the experience pool, the Q function value of the deep Q network is updated according to the Q function value output by the current deep Q network and the target Q function value calculated by the target neural network; the loss is calculated by back propagation and the neural network parameter θ is optimized to approximate the true Q function value. The update formula of the Q function value is: Among them, a is the learning rate, which controls the update step size; S4, using the target neural network, every certain number of steps, the target neural network parameter θ - It will be updated to the neural network parameter θ of the current deep Q network, thereby ensuring that the deep Q network will not oscillate during training due to unstable target values.

7. The method of a system for extending the life of a motor based on fuzzy control according to claim 2, characterized in that: In step S04, the system will first automatically generate a fuzzy rule, and dynamically optimize the fuzzy rule, membership function and controller parameters by combining reinforcement learning and deep learning. The membership function is used to map each element in the fuzzy set to a membership degree of [0,1], indicating the degree to which a certain input is in a certain fuzzy set; Among them, the fuzzy set is the set formed by fuzzy rules; In the automatic optimization process, the nonlinear mapping between input values ​​and membership functions is learned through neural networks. Here, the input values ​​include temperature difference and vibration difference, and the shape of the membership function is automatically adjusted.

Citation Information

Patent Citations

  • Expert experience reasoning system and method based on fuzzy reasoning

    CN108847282B

  • Distributed intelligent lubricating system

    CN118009023A

  • UUV real-time collision avoidance planning method based on deep double-Q network reinforcement learning

    CN110716575A

  • Fuzzy control 'cloud green-making' intelligent algorithm based on reinforcement learning Q-Learning

    CN113485104A

  • Mobile robot path planning algorithm combining fuzzy control and reinforcement learning

    CN115826581A

Cited By

  • Remote task issuing management method and system

    CN120711014A