A system and method for prolonging the operating life of a motor based on fuzzy control

CN120090526BActive Publication Date: 2026-09-29EAST CHINA JIAOTONG UNIVERSITY
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510084778.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-20
Publication Date
2026-09-29
Estimated Expiration
2045-01-20

AI Technical Summary

Technical Problem

该系统确实在一定程度上缓解了疾病判断准确率低和分诊效率低下的问题,但过分依赖专家经验也带来了局限性,面对动态变化的系统,可能难以及时更新规则或调整推理结果,从而引发延迟

Benefits of technology

[0045]通过引入强化学习,如Q-learning、深度Q网络,通过强化学习和深度学习技术,不断提高系统的自适应能力和精度,在这一过程中,强化学习用于优化控制策略,而深度学习用于优化系统的模型和模糊规则,使得系统能够自动适应不同的环境条件,提供更加精准和高效的润滑周期控制。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120090526B_ABST
    Figure CN120090526B_ABST
Patent Text Reader

Abstract

The application discloses a system and method for prolonging the running life of a motor based on fuzzy control, which comprises the following steps: establishing a state space, an action space and a reward function, and defining an optimization target as an optimal lubrication period; using a Q-learning algorithm to learn an optimal strategy by interacting with the environment, and adjusting the lubrication period; using a deep Q network to approximate a Q function value, and automatically extracting fuzzy control rules and membership functions through a neural network; automatically learning a nonlinear relationship between temperature and vibration difference through the neural network, and adjusting the shapes of the fuzzy rules and the membership functions; dynamically adjusting the lubrication period according to real-time states, and realizing adaptive control. Through the collection and processing of relevant information of motor lubrication points, the required lubricating grease during the operation of the motor is inferred, the corresponding lubricating grease is added to the lubrication points, accurate lubrication of the motor is realized, frictional resistance is reduced, and the running life of the motor is prolonged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of intelligent control technology, and specifically relates to a system and method for extending the service life of a motor based on fuzzy control. Background Technology

[0002] During operation, friction occurs between the moving parts of an electric motor. This friction not only leads to energy loss but also accelerates wear and tear on the components. Adding grease can significantly reduce the coefficient of friction. Traditional grease application is manual, which cannot achieve continuous and efficient lubrication. Insufficient grease can result in poor lubrication, shortening the motor's lifespan, while excessive grease increases internal bearing resistance, reducing operating efficiency and potentially causing motor malfunctions. Compared to manual lubrication, intelligent lubrication can indeed provide more precise and efficient grease application to some extent, but its lubrication cycle setting still relies on manual operation. It cannot achieve real-time monitoring of the motor's lubrication status or on-demand grease application, therefore, lubrication accuracy still needs improvement.

[0003] Chinese patent CN118009023A discloses a "distributed lubrication system." This system uploads data such as the amount of grease and battery power from the lubricator to the cloud via an NB-lot module, and the lubrication cycle can be remotely modified by issuing modification commands from the cloud. However, this system can only be modified using a pre-set cycle. If the lubrication accuracy is low, precise lubrication cannot be achieved, and inaccurate cycle settings may affect the lifespan of the motor.

[0004] Chinese patent CN108847282B discloses an "Expert Experience Reasoning System and Method Based on Fuzzy Reasoning." This system uses fuzzy reasoning to calculate the disease information corresponding to a user's symptom information based on rules in an expert experience rule base. While this system does alleviate the problems of low accuracy in disease diagnosis and inefficient triage to some extent, its over-reliance on expert experience also introduces limitations. In the face of a dynamically changing system, it may be difficult to update rules or adjust reasoning results in a timely manner, leading to delays.

[0005] To cope with the dynamically changing operating conditions of motors, improve lubrication accuracy, and extend motor service life, it is necessary to introduce new technologies and methods. The urgent problem to be solved is to adjust the lubrication cycle in real time according to the motor's operating conditions so as to add grease more accurately.

[0006] The information disclosed in this background section is intended only to enhance the understanding of the overall background of the invention and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention

[0007] The purpose of this invention is to provide a system and method for extending the service life of a motor based on fuzzy control, thereby overcoming the defects in the prior art.

[0008] To achieve the above objectives, this invention provides a system for extending the service life of a motor based on fuzzy control, comprising an environment and state space module, a reinforcement learning module, a deep learning and fuzzy inference integration module, a fuzzy rule and membership function optimization module, and a feedback and real-time control module. The environment and state space module defines a state space that interacts with the simulated actual environment, collects temperature and vibration sensor data, and evaluates the current system state based on this data. The reinforcement learning module uses a reinforcement learning algorithm for training, learning to select adjustment strategies for the lubrication cycle based on different states. The deep learning and fuzzy inference integration module automatically extracts fuzzy rules and membership functions, improving the accuracy of fuzzy inference and providing more precise input information for fuzzy control. The fuzzy rule and membership function optimization module optimizes the shape of fuzzy rules and membership functions, ensuring more precise and effective adjustment of the lubrication cycle, and further refining the fuzzy control. The feedback and real-time control module dynamically adjusts the lubrication cycle based on collected temperature and vibration sensor data and system state, ensuring that the lubrication cycle is highly consistent with the actual operating environment.

[0009] A method for extending the service life of a motor based on fuzzy control, comprising the following steps:

[0010] S01. Collect temperature and vibration values ​​at the motor lubrication points using sensors and upload them to the cloud via a 4G module. Establish a state space, action space, and reward function in the cloud and define the optimization target as the optimal lubrication cycle.

[0011] S02. Using the Q-learning algorithm, the quality of the reward function variable is evaluated through the Q function. By continuously updating the Q function value, the optimal policy is learned and the lubrication cycle is adjusted. The reward function variable is a state-action pair.

[0012] S03. Approximate the Q function using a deep Q-network and automatically extract fuzzy rules and membership functions through a neural network;

[0013] S04. The neural network automatically learns the nonlinear relationship between temperature and vibration difference, and adjusts the shape of fuzzy rules and membership functions.

[0014] S05. The system dynamically adjusts the lubrication cycle based on the real-time status information, including temperature difference and vibration difference, and sends the cycle command to the lubricator. The lubricator then injects oil according to the received cycle command.

[0015] Preferably, in the technical solution, in step S01, the state space is the current information of the system, used to describe the current state of the system, and the state space is defined as the following set:

[0016] S={TempDiff,VibChange,PrevAction} (1),

[0017] Where S is the state space; TempDiff is the temperature difference, representing the difference between the current temperature and the target temperature; VibChange is the vibration change, representing the change in the current vibration signal; PrevAction is the control output of the previous moment, representing the adjustment of the previous lubrication cycle;

[0018] The action space defines the control behaviors that the system can choose. The action space is defined as the following set:

[0019] A = {Increase lubrication cycle, decrease lubrication cycle, keep it unchanged} (2).

[0020] Where A represents the action space; the control behaviors include increasing the lubrication cycle, decreasing the lubrication cycle, and keeping it unchanged;

[0021] The reward function is used to measure the effect of each control action. The reward function is defined as follows:

[0022] (3),

[0023] Where R() is the reward function; This indicates the current state, which includes temperature and vibration values. Indicates that the system is in Actions taken under the given conditions; R() = +10 indicates high temperature and high vibration adjustment, R() = -10 indicates low temperature and low vibration adjustment, and R() = 0 indicates ineffective adjustment.

[0024] Preferably, in the technical solution, in step S02, the Q function formula is expressed as:

[0025] (4),

[0026] in, Indicates the state Take action below The expected reward; s0 represents the initial state; a0 represents the initial action taken by the system in state s0; It indicates a reward signal, measuring the immediate reward after taking an action; This represents the discount factor, which determines the weight of future rewards.

[0027] Preferably, in the technical solution, in step S03, the deep Q-network approximates the Q-function in the Q-learning algorithm using a deep neural network:

[0028] (5),

[0029] in, Here are the parameters of the neural network; Q() represents the Q-function value output by the deep Q-network; Q * ( ) represents the optimal Q-function value output by the deep Q-network;

[0030] Deep Q-networks make the current Q-function value closer to the true target Q-function value, minimizing the error of the Q-function:

[0031] (6),

[0032] Where L(θ) represents the loss function, which measures the difference between the Q-function value output by the current deep Q-network and the target Q-function value; E[] represents the expected value; It is the optimal action Q-function value in the next state; a' represents the next action in the next state; Current state Execute action The reward afterwards The Q-function value of the current policy. These are the parameters of the target neural network.

[0033] Preferably, in the technical solution, the training process of the deep Q-network includes the following steps:

[0034] S1. Randomly initialize the neural network parameters of the deep Q-network. and target neural network parameters ;

[0035] S2. Through the experience replay mechanism of deep Q-networks, the system performs actions in the environment and stores the experience in the experience pool. Then, a batch of data is randomly sampled from the experience pool for training. Experience includes states. ,action ,award Next state ;

[0036] S3. After sampling a batch of data from the experience pool each time, update the Q-function value of the deep Q-network based on the Q-function value output by the current deep Q-network and the target Q-function value calculated by the target neural network; calculate the loss and optimize the neural network parameters through backpropagation. To approximate the true Q-function value, the update formula for the Q-function value is:

[0037] (7),

[0038] Where α is the learning rate, which controls the update step size;

[0039] S4. Using a target neural network, the parameters of the target neural network are adjusted at regular intervals. It will be updated to the neural network parameters of the current deep Q-network. This ensures that the deep Q-network will not oscillate during training due to unstable target values.

[0040] Preferably, in the technical solution, in step S04, the system first automatically generates a fuzzy rule, and then dynamically optimizes the fuzzy rule, membership function, and controller parameters by combining reinforcement learning and deep learning;

[0041] Membership functions are used to map each element in a fuzzy set to a membership degree of [0, 1], representing the degree to which an input belongs to a certain fuzzy set; where a fuzzy set is a set formed by fuzzy rules.

[0042] During the automatic optimization process, the neural network learns the nonlinear mapping between the input values ​​and the membership function, where the input values ​​include temperature difference and vibration difference. The shape of the membership function is automatically adjusted so that the system can handle complex inputs more accurately.

[0043] The system interacts with the environment—specifically, the operating status and lubrication effect of the equipment—and adjusts controller parameters using the Q-learning algorithm. When the system detects that "increasing the lubrication cycle when the temperature difference is large" reduces the incidence of equipment failure, the reinforcement learning model adjusts the weights of fuzzy rules based on this feedback signal, making the system more likely to perform the same actions when encountering similar situations in the future. Through the combination of deep learning and reinforcement learning, the system can automatically extract fuzzy rules from temperature and vibration sensor data, optimize membership functions, and continuously adjust the weights of fuzzy rules and the shape of membership functions.

[0044] Compared with the prior art, the present invention has the following beneficial effects:

[0045] By introducing reinforcement learning, such as Q-learning and deep Q-networks, the system's adaptability and accuracy can be continuously improved through reinforcement learning and deep learning techniques. In this process, reinforcement learning is used to optimize control strategies, while deep learning is used to optimize the system's model and fuzzy rules, enabling the system to automatically adapt to different environmental conditions and provide more accurate and efficient lubrication cycle control. Attached Figure Description

[0046] Figure 1This is a system principle block diagram of the present invention for extending the service life of a motor based on fuzzy control;

[0047] Figure 2 This is a flowchart of a method for extending the service life of a motor based on fuzzy control according to the present invention. Detailed Implementation

[0048] The specific embodiments of the present invention will be described in detail below, but it should be understood that the scope of protection of the present invention is not limited to the specific embodiments.

[0049] Unless otherwise expressly stated, throughout the specification and claims, the term "comprising" or its variations such as "including" or "comprises" shall be understood to include the stated elements or components without excluding other elements or other components.

[0050] like Figure 1 As shown, a system for extending the service life of a motor based on fuzzy control includes an environment and state space module, a reinforcement learning module, a deep learning and fuzzy inference integration module, a fuzzy rule and membership function optimization module, and a feedback and real-time control module. The environment and state space module defines a state space that interacts with the simulated actual environment, collects temperature and vibration sensor data, and evaluates the current system state based on this data. The reinforcement learning module uses reinforcement learning algorithms for training, learning to select adjustment strategies for the lubrication cycle based on different states. The deep learning and fuzzy inference integration module automatically extracts fuzzy rules and membership functions, improving the accuracy of fuzzy inference and providing more precise input information for fuzzy control. The fuzzy rule and membership function optimization module optimizes the shape of fuzzy rules and membership functions, ensuring more precise and effective adjustment of the lubrication cycle, and further refining the fuzzy control. The feedback and real-time control module dynamically adjusts the lubrication cycle based on collected temperature and vibration sensor data and system state, ensuring that the lubrication cycle is highly consistent with the actual operating environment.

[0051] like Figure 2 As shown, a method for extending the service life of a motor based on fuzzy control includes the following steps:

[0052] S01. Collect temperature and vibration values ​​at the motor lubrication points using sensors and upload them to the cloud via a 4G module. Establish a state space, action space, and reward function in the cloud and define the optimization target as the optimal lubrication cycle.

[0053] The state space is a collection of information about the current state of the system, used to describe its current condition. The state space is defined as the following set:

[0054] S={TempDiff,VibChange,PrevAction} (1),

[0055] Where S is the state space; TempDiff is the temperature difference, representing the difference between the current temperature and the target temperature; VibChange is the vibration change, representing the change in the current vibration signal; PrevAction is the control output of the previous moment, representing the adjustment of the previous lubrication cycle;

[0056] The action space defines the control behaviors that the system can choose. The action space is defined as the following set:

[0057] A = {Increase lubrication cycle, decrease lubrication cycle, keep it unchanged} (2).

[0058] Where A represents the action space; the control behaviors include increasing the lubrication cycle, decreasing the lubrication cycle, and keeping it unchanged;

[0059] The reward function is used to measure the effect of each control action. The reward function is defined as follows:

[0060] (3),

[0061] Where R() is the reward function; This indicates the current state, which includes temperature and vibration values. Indicates that the system is in The reward function is defined as follows: if increasing the lubrication cycle reduces the temperature and vibration differences of the equipment, making the equipment more stable, the reward function is +10; if decreasing the lubrication cycle increases the temperature and vibration differences of the equipment, making the equipment unstable, the reward function is -10; when the cycle remains unchanged, the reward function is defined as 0.

[0062] S02. Using the Q-learning algorithm, the quality of the reward function variable is evaluated through the Q function. By continuously updating the Q function value, the optimal policy is learned and the lubrication cycle is adjusted. The reward function variable is a state-action pair.

[0063] The formula for the Q function is expressed as follows:

[0064] (4),

[0065] in, Indicates the state Take action below The expected reward; s0 represents the initial state; a0 represents the initial action taken by the system in state s0; It indicates a reward signal, measuring the immediate reward after taking an action; This represents the discount factor, which determines the weight of future rewards;

[0066] A table called the Q-table is used to store all possible state-action pairs. The Q-table is initialized with all values ​​set to 0, indicating that we do not yet have any knowledge about the individual state-action pairs. The system obtains state information by interacting with the environment and selects an action, which may be to increase lubrication, decrease lubrication, or not adjust. The environment provides feedback after each action.

[0067] exist In this state, the system selects "increase lubrication cycle" and the system feedback reward is +10. The Q function is updated according to the Q function update formula. Assuming that the temperature difference and vibration difference have changed at this time, the system selects an action according to the new state and receives the corresponding reward. After multiple rounds of training, the system learns to increase or decrease the lubrication cycle under specific temperature and vibration differences in order to achieve optimal equipment performance.

[0068] S03. Approximate the Q function using a deep Q-network and automatically extract fuzzy rules and membership functions through a neural network;

[0069] Deep Q-networks approximate the Q-function in the Q-learning algorithm using deep neural networks:

[0070] (5),

[0071] in, Here are the parameters of the neural network; Q() represents the Q-function value output by the deep Q-network; Q * ( ) represents the optimal Q-function value output by the deep Q-network;

[0072] Deep Q-networks make the current Q-function value closer to the true target Q-function value, minimizing the error of the Q-function:

[0073] (6),

[0074] Where L(θ) represents the loss function, which measures the difference between the Q-function value output by the current deep Q-network and the target Q-function value; E[] represents the expected value; It is the optimal action Q-function value in the next state; a' represents the next action in the next state; Current state Execute action The reward afterwards The Q-function value of the current policy. For the target neural network parameters;

[0075] The training process for a deep Q-network includes the following steps:

[0076] S1. Randomly initialize the neural network parameters of the deep Q-network. and target neural network parameters ;

[0077] S2. Through the experience replay mechanism of deep Q-networks, the system performs actions in the environment and stores the experience in the experience pool. Then, a batch of data is randomly sampled from the experience pool for training. Experience includes states. ,action ,award Next state ;

[0078] S3. After sampling a batch of data from the experience pool each time, update the Q-function value of the deep Q-network based on the Q-function value output by the current deep Q-network and the target Q-function value calculated by the target neural network; calculate the loss and optimize the neural network parameters through backpropagation. To approximate the true Q-function value, the update formula for the Q-function value is:

[0079] (7),

[0080] Where α is the learning rate, which controls the update step size;

[0081] S4. Using a target neural network, the parameters of the target neural network are adjusted at regular intervals. It will be updated to the neural network parameters of the current deep Q-network. This ensures that the deep Q-network will not oscillate during training due to unstable target values;

[0082] S04. The neural network automatically learns the nonlinear relationship between temperature and vibration difference, and adjusts the shape of fuzzy rules and membership functions.

[0083] In the initial state, the system first automatically generates a fuzzy rule, and then dynamically optimizes the fuzzy rule, membership function, and controller parameters by combining reinforcement learning and deep learning.

[0084] Assume the generated fuzzy rules are as follows: Rule 1: If the temperature difference is small and the vibration difference is small, maintain the current lubrication cycle; Rule 2: If the temperature difference is large and the vibration difference is small, reduce the lubrication cycle; Rule 3: If the temperature difference is small and the vibration difference is large, increase the lubrication cycle.

[0085] By continuously adjusting the weights of fuzzy rules and optimizing the selection of existing fuzzy rules, we can gradually learn more reasonable fuzzy rules, reducing the difficulty of manually setting fuzzy rules. If a certain fuzzy rule reduces the possibility of equipment failure, then the weight of that fuzzy rule will be increased, and vice versa.

[0086] Membership functions are used to map each element in a fuzzy set to a membership degree of [0, 1], representing the degree to which an input belongs to a certain fuzzy set; where a fuzzy set is a set formed by fuzzy rules.

[0087] Initialize membership functions, defining a triangle as the initial membership function. Assume the temperature difference varies within the range [0, 100], and the vibration difference varies within the range [0, 10]. Low temperature is defined as a triangle varying within the range [0, 0, 50], medium temperature as a triangle within the range [40, 50, 60], and high temperature as a triangle within the range [50, 100, 100]. Through training of a deep neural network, the system adjusts the shape of these membership functions. When the temperature difference is large, the peak value of the membership function automatically shifts to the "high temperature" range; when the vibration difference is large, the peak value of the membership function shifts towards the "high vibration" range. Through the reward signal in reinforcement learning, the neural network parameters are optimized through the backpropagation algorithm. Here, the neural network parameters include the shape and peak position of the membership function.

[0088] Define controller parameters. Each fuzzy rule has a weight, which represents the importance of the fuzzy rule in the decision-making process. The weight of rule 1 is 0.3, the weight of rule 2 is 0.5, and the weight of rule 3 is 0.2. The system adjusts the weight of each rule according to the reward signal. For example, if rule 2 improves the operating efficiency of the equipment or reduces wear, the system will increase the weight of the rule, and vice versa.

[0089] Suppose that the system performs the action "reduce lubrication cycle" under the state "temperature difference = 50, vibration difference = 5". The system evaluates the result and receives a reward of +10. The equipment wear is reduced and the operation is more stable. Through Q function value update, the system will adjust the weight of the corresponding rule, such as rule 2, so that the system will be more inclined to choose the rule in similar states.

[0090] S05. The system dynamically adjusts the lubrication cycle based on the real-time status information, including temperature difference and vibration difference, and sends the cycle command to the lubricator. The lubricator then injects oil according to the received cycle command.

[0091] The foregoing description of specific exemplary embodiments of the invention is for illustrative and explanatory purposes. These descriptions are not intended to limit the invention to the precise forms disclosed, and it will be apparent that many changes and variations can be made in accordance with the foregoing teachings. The exemplary embodiments were chosen and described in order to explain the specific principles of the invention and its practical application, thereby enabling those skilled in the art to implement and utilize various different exemplary embodiments of the invention, as well as various different choices and variations. The scope of the invention is intended to be defined by the claims and their equivalents.

Claims

1. A method for extending the service life of a motor based on fuzzy control, characterized in that, The system includes an environment and state space module, a reinforcement learning module, a deep learning and fuzzy inference integration module, a fuzzy rule and membership function optimization module, and a feedback and real-time control module; its steps are as follows: S01. Collect temperature and vibration values ​​at the motor lubrication points using sensors and upload them to the cloud via a 4G module. Establish a state space, action space, and reward function in the cloud and define the optimization target as the optimal lubrication cycle. The state space is a collection of information about the current state of the system, used to describe its current condition. The state space is defined as the following set: S={TempDiff,VibChange,PrevAction} (1), Where S is the state space; TempDiff is the temperature difference, representing the difference between the current temperature and the target temperature; VibChange is the vibration change, representing the change in the current vibration signal; PrevAction is the control output of the previous moment, representing the adjustment of the previous lubrication cycle; The action space defines the control behaviors that the system can choose. The action space is defined as the following set: A = {Increase lubrication cycle, decrease lubrication cycle, keep it unchanged} (2). Where A represents the action space; the control behaviors include increasing the lubrication cycle, decreasing the lubrication cycle, and keeping it unchanged; The reward function is used to measure the effect of each control action. The reward function is defined as follows: (3), Where R() is the reward function; This indicates the current state, which includes temperature and vibration values. Indicates that the system is in Actions taken under the given conditions; R() = +10 indicates high temperature and high vibration adjustment, R() = -10 indicates low temperature and low vibration adjustment, and R() = 0 indicates ineffective adjustment; S02. Using the Q-learning algorithm, the quality of the reward function variable is evaluated through the Q function. By continuously updating the Q function value, the optimal policy is learned and the lubrication cycle is adjusted. The reward function variable is a state-action pair. S03. Approximate the Q function using a deep Q-network and automatically extract fuzzy rules and membership functions through a neural network; S04. The neural network automatically learns the nonlinear relationship between temperature and vibration difference, and adjusts the shape of fuzzy rules and membership functions. S05. The system dynamically adjusts the lubrication cycle based on the real-time status information, including temperature difference and vibration difference, and sends the cycle command to the lubricator. The lubricator then injects oil according to the received cycle command.

2. The method for extending the service life of a motor based on fuzzy control according to claim 1, characterized in that: In step S02, the formula for the Q function is expressed as: (4), in, Indicates the state Take action below The expected reward; s0 represents the initial state; a0 represents the initial action taken by the system in state s0; It indicates a reward signal, measuring the immediate reward after taking an action; This represents the discount factor, which determines the weight of future rewards.

3. The method for extending the service life of a motor based on fuzzy control according to claim 1, characterized in that: In step S03, the deep Q-network approximates the Q-function in the Q-learning algorithm using a neural network: (5), in, Here are the parameters of the neural network; Q() represents the Q-function value output by the deep Q-network; Q * ( ) represents the optimal Q-function value output by the deep Q-network; Deep Q-networks make the current Q-function value closer to the true target Q-function value, minimizing the error of the Q-function: (6), Where L(θ) represents the loss function, which measures the difference between the Q-function value output by the current deep Q-network and the target Q-function value; E[] represents the expected value; It is the optimal action Q-function value in the next state; a' represents the next action in the next state; Current state Execute action The reward afterwards The Q-function value of the current policy. These are the parameters of the target neural network.

4. The method for extending the service life of a motor based on fuzzy control according to claim 3, characterized in that: The training process for a deep Q-network includes the following steps: S1. Randomly initialize the neural network parameters of the deep Q-network. and target neural network parameters ; S2. Through the experience replay mechanism of deep Q-networks, the system performs actions in the environment and stores the experience in the experience pool. Then, a batch of data is randomly sampled from the experience pool for training. Experience includes states. ,action ,award Next state ; S3. After sampling a batch of data from the experience pool each time, update the Q-function value of the deep Q-network based on the Q-function value output by the current deep Q-network and the target Q-function value calculated by the target neural network; calculate the loss and optimize the neural network parameters through backpropagation. To approximate the true Q-function value, the update formula for the Q-function value is: (7), Where α is the learning rate, which controls the update step size; S4. Using a target neural network, the parameters of the target neural network are adjusted at regular intervals. It will be updated to the neural network parameters of the current deep Q-network. This ensures that the deep Q-network will not oscillate during training due to unstable target values.

5. The method for extending the service life of a motor based on fuzzy control according to claim 1, characterized in that: In step S04, the system first automatically generates a fuzzy rule, and then dynamically optimizes the fuzzy rule, membership function, and controller parameters by combining reinforcement learning and deep learning. Membership functions are used to map each element in a fuzzy set to a membership degree of [0, 1], representing the degree to which an input belongs to a certain fuzzy set; The fuzzy set is the set formed by fuzzy rules; During the automatic optimization process, a neural network learns the nonlinear mapping between input values ​​and membership functions, where input values ​​include temperature difference and vibration difference, and automatically adjusts the shape of the membership functions.

Citation Information

Patent Citations

  • Expert experience reasoning system and method based on fuzzy reasoning

    CN108847282B

  • Distributed intelligent lubricating system

    CN118009023A

  • UUV real-time collision avoidance planning method based on deep double-Q network reinforcement learning

    CN110716575A

  • Fuzzy control 'cloud green-making' intelligent algorithm based on reinforcement learning Q-Learning

    CN113485104A