Gearbox Unbalance Fault Diagnosis Method
The multi-scale deep attention reinforcement learning network addresses the challenges of imbalanced data and varying conditions in gear box fault diagnosis by enhancing feature extraction and model robustness, achieving high accuracy and reliability.
Patent Information
- Application Number
- CN202310927800.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-27
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-07-27
AI Technical Summary
The prior art faces data imbalance and difficulty in gearbox fault diagnosis under multiple operating conditions, resulting in insufficient diagnostic accuracy and universality. It is difficult for traditional methods to adapt to the state differences and frequency changes of fault characteristic under multiple operating conditions.
A multi-scale deep attention reinforcement learning network is adopted, and by constructing an unbalanced classification Markov decision-making process and reward strategy, combining multi-scale convolution, channel attention mechanism and residual network, an optimal diagnostic strategy for autonomous learning of the agent is established to achieve fault diagnosis under multiple operating conditions.
It improves the reliability and versatility of gearbox fault diagnosis, enhances the richness of feature extraction and diagnosis accuracy, and adapts to fault identification under multiple operating conditions.
Smart Images

Figure CN116975563B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure belongs to the field of gearbox fault diagnosis, and particularly relates to a gearbox unbalance fault diagnosis method based on a multi-scale deep attention reinforcement learning network. Background Art
[0002] Most actual gearboxes operate under multiple working conditions, and the amount of fault-class data is often less than that of normal-state data. Under such circumstances, gearbox fault diagnosis still faces great challenges. Under multiple working conditions, the differences between and within different states of the gearbox change significantly. Especially when the rotational speed of the gearbox changes, the fault characteristic frequencies of components will change accordingly. Methods based on vibration signal analysis need to rely on professional knowledge to extract and analyze the fault characteristic frequencies under each working condition, making it difficult to achieve large-scale deployment and automated diagnosis. Traditional shallow machine learning methods rely heavily on the construction of feature index sets. However, multiple types of working conditions will increase the differences between samples within the same class state and reduce the differences between samples in different class states, resulting in serious aliasing of the manually constructed fault feature set in the high-dimensional feature space, making it difficult to distinguish different states of the gearbox. At the same time, facing the diversity of operating conditions, it is difficult to ensure the comprehensiveness and stability of manual feature index extraction to meet the diagnostic requirements under all working conditions, affecting the effectiveness of traditional shallow models in gearbox fault diagnosis. Traditional convolutional neural networks usually use single-scale models to extract feature information at a fixed scale. Facing the huge differences between samples in different states under multiple working conditions, rich and stable discriminative feature information also needs to be extracted to enhance the accuracy of gearbox fault diagnosis under multiple working conditions. In addition, since the amount of fault-class data is less than that of normal-state data, and most existing methods are only applicable to fault diagnosis under balanced class distributions, the accuracy and generality of gearbox fault diagnosis under multi-condition imbalance still need to be improved.
[0003] The deep reinforcement learning method can not only adaptively extract discriminative feature information from the original vibration signal, well avoiding the defects of manual feature extraction, but also has the decision-making ability to interact with the environment, which can well solve the generality and effectiveness problems of gearbox fault diagnosis. However, facing the differences in vibration signals within the same class state and the similarities in vibration signals between different class states under multi-condition imbalance of the gearbox, it is still necessary to design a reasonable CNN structure to adaptively extract a deep and rich discriminative feature set to meet the fault diagnosis requirements under multiple working conditions. At the same time, without changing the original data space structure, how to make full use of the decision-making advantages of deep reinforcement learning also requires designing a reasonable environment simulation to automatically realize the effective learning of fault diagnosis strategies under class imbalance.
[0004] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present invention, and thus may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention
[0005] Aiming at the deficiencies existing in the prior art, the purpose of the present disclosure is to propose a gearbox imbalance fault diagnosis method based on a multi-scale deep attention reinforcement learning network. Based on the deep reinforcement learning algorithm, an environment simulation is constructed with an imbalance classification Markov decision process, and an environment reward strategy is defined by defining the class deviation degree to realize the learning of the fault diagnosis strategy under class imbalance. In order to enhance the feature perception ability of the model under multiple working conditions, a network structure is established by using multi-scale convolution, channel attention mechanism and residual network, so as to improve the reliability and universality of gearbox imbalance fault diagnosis.
[0006] To achieve the above object, the present disclosure provides the following technical solutions:
[0007] A gearbox imbalance fault diagnosis method based on a multi-scale deep attention reinforcement learning network includes the following steps:
[0008] Step S1: Signal acquisition, obtaining vibration signals of the gearbox in different health conditions under multiple working conditions, and constructing a training set and a test set based on the vibration signals, where the number of normal state samples is more than that of fault state samples;
[0009] Step S2: Environment simulation construction, establishing a Markov decision process for imbalance classification, and designing a reward function to establish a data environment simulation required for fault diagnosis under class imbalance;
[0010] Step S3: Establishing a multi-working-condition imbalance deep reinforcement learning network, based on the training set, the intelligent agent and the environment continuously interact to train the intelligent agent to autonomously learn the optimal diagnosis strategy, and the intelligent agent includes at least two multi-scale feature deep attention networks with the same structure;
[0011] Step S4: Fault identification, inputting the samples in the test set into the trained intelligent agent one by one, and identifying the gearbox fault type according to the diagnosis strategy and analyzing the diagnosis result.
[0012] In the gearbox imbalance fault diagnosis method, in step S1, the vibration signal is sample-segmented, and a training sample signal and a test sample signal are constructed. Each sample length contains 2048 data points, and the amplitude is normalized to the range of [-1, 1].
[0013] In the gearbox imbalance fault diagnosis method described above, in step S2, the Markov decision process includes a state space S, an action space A, and a reward function R; the state space S consists of all training samples in the training set, and each environmental state s corresponds to a training sample; the action space A is K actions corresponding to the health status of the gearbox, A = {0, 1,..., K - 1}, where the total number of fault types is K; the reward function R is: under the current diagnostic query, corresponding to the current environmental state s t ∈S, the agent executes the action a t , and compares whether the action a t is consistent with the label class under the current diagnostic query. If it is consistent, the environment returns a positive reward to the agent, otherwise it returns a negative reward, forming a reward strategy to enable the agent to pay more attention to the learning of minority class samples under the unbalanced distribution of class data volume, and to realize the effective learning of the diagnostic strategy under class imbalance. Among them,
[0014] training set D train ={D1; D1;...; D k ;...}, where D k represents the training subset of the k-th class, which is expressed as: where (x i , l i ) represents the corresponding sample and label information, and n k represents the total number of samples in the k-th class subset. The reward function R is set by the following formula:
[0015]
[0016] In the formula, R(s t , a t , l t ) represents the environmental feedback reward value obtained by executing the action a t at the state s t , which is simplified and replaced by r t ; log2(·) is the logarithm function with base 2, which is used to evaluate the difference between classes; ρ k represents the class deviation degree of the k-th class, which is the imbalance degree of the sample volume of the k-th class subset in the training set D train relative to the sample volume of the least class in D train , and is obtained by the following formula:
[0017]
[0018] In the formula, D min represents the sample subset of the least class in the training set D train , and |·| represents taking the modulus, that is, it represents the sample volume in the sample subset.
[0019] In the gearbox unbalance fault diagnosis method described above, in step S3, the scale feature depth attention network consists of a multi-scale feature layer, a channel attention layer, an Inception multi-scale layer, a max pooling layer, two residual modules, a global average pooling layer, and an output layer.
[0020] In the gearbox unbalance fault diagnosis method described above, the agent includes at least two multi-scale feature depth attention networks of the same-structured current Q network Eval-Net and another multi-scale feature depth attention network of the target Q network Target-Net.
[0021] In the gearbox unbalance fault diagnosis method described above, step S3 includes the following steps:
[0022] Step S3.1: Set the maximum number of autonomous learning episodes Episode. The maximum number of autonomous learning episodes Episode refers to the transition trajectory of the environment from the initial state s1 to the final state s T , Episode = {s1, a1, r1, s2, a2, r2, …, s T , a T , r T}, where T represents the termination time step, and the current training episode Episode ends, and the next episode starts until the set number of interaction rounds is reached; each round of autonomous learning contains T diagnostic inquiries, and each diagnostic inquiry corresponds to an environmental state;
[0023] Step S3.2: Randomly initiate a diagnostic inquiry to obtain the current state s t ∈s, input the corresponding training sample into the current Q network Eval-Net, and the agent selects the current action a t ∈A according to the linear annealing ∈-greedy algorithm. The environment returns the current reward r t according to the reward function R, and randomly initiates the next diagnostic inquiry. The state s t is converted to the next state s t+1 , and the generated experience data e = {s t , a t , r t , s t+1} is saved to the experience pool ;
[0024] Step S3.3: Repeat step S3.2 until the end of this round of T diagnostic inquiries, and output the total reward obtained from this round of diagnostic inquiries. The total reward is j represents the jth diagnostic inquiry in this round of diagnostic inquiries, and r j represents the reward corresponding to the jth diagnostic inquiry;
[0025] Step S3.4: Randomly sample a predetermined batch of experience data e = {s from the experience pool t , a t , r t , s t+1}, and based on the experience data e, use the gradient descent method to train the current Q-network Eval-Net, update the model parameters of Eval-Net, and perform a delayed update on the target Q-network Target-Net. The delayed update is to directly copy the model parameters of Eval-Net to the target Q-network Target-Net every C learning rounds;
[0026] Step S3.5: Start the next round of autonomous learning process, and repeat steps S3.2 to S3.4 until the maximum number of autonomous learning rounds is reached, and the autonomous learning process ends;
[0027] Step S3.6: Save the model parameters with a higher total reward obtained in each round of diagnostic query during the autonomous learning process as the optimal diagnostic strategy learned by the agent.
[0028] In the gearbox unbalance fault diagnosis method described above, step S3.2 includes the following steps:
[0029] Step S3.2.1: Preset the initial value of ∈ as ∈ = 1, the change rate of ∈ as Δ∈ = 1 / 100000, and the minimum value of ∈ min as ∈ = 0.01;
[0030] Step S3.2.2: Before each execution of the action a t , randomly generate a random number between [0, 1]. If the random number belongs to the interval [0, ∈], randomly select an execution action a t from the action space A; if the random number belongs to the interval (∈, 1], use the action corresponding to the maximum output of the current Q-network Eval-Net as the execution action a t ;
[0031] Step S3.2.3: After each round of diagnostic query, dynamically update the value of ∈. If ∈ is less than ∈ min , ∈ = ∈ min , otherwise ∈ = ∈ - Δ∈.
[0032] In the gearbox unbalance fault diagnosis method described above, step S3.4 includes the following steps:
[0033] Step S3.4.1: According to the experience data e = {s sampled from the experience pool t , a t , r t , s t+1}, the current Q-network Eval-net outputs the current Q-value as Q(s t , a t ; θ), and the target Q-network Target-net calculates the target Q-value as where θ and θ- are the network parameters of Eval-net and Target-net respectively; γ is the discount factor, γ ∈ [0, 1];
[0034] Step S3.4.2: Calculate the mean squared error (MSE) of the current Q-value and the target Q-value:
[0035] L(θ) = E[(Q′ - Q(s t , a t ; θ)) 2 (3)
[0036] Step S3.4.3: Calculate the gradient of the mean squared error L(θ) with respect to the network parameter θ:
[0037]
[0038] where represents the operation of taking the gradient with respect to the parameter θ; represents that the empirical data (s t , a t , r t , s t+1 ) is from represents randomly and uniformly sampling from the experience pool ;
[0039] Step S3.4.4: Repeat Steps S3.4.1 to S3.4.3, and update the model parameters of the current Q-network Eval-net according to the gradient descent method.
[0040] In the gearbox imbalance fault diagnosis method described above, the fault states include gear fault states and bearing fault states.
[0041] In the gearbox imbalance fault diagnosis method described above, the gear fault states include cut tooth fault CTF, missing tooth fault MTF, root crack fault RCF, and tooth surface wear fault SWF, and the bearing fault states include rolling element fault BWF, compound fault SWF, inner ring fault IRF, and outer ring fault ORF.
[0042] Compared with the prior art, the beneficial effects brought by the present disclosure are:
[0043] Multi-scale convolution operations are adopted to automatically extract features of different scales in each state, enhancing the richness of feature extraction and the accuracy of diagnosis. Then, aiming at the redundancy of feature information in different channels, a channel attention mechanism is used to adaptively weight the feature information of different channels, recalibrate the multi-scale feature information, highlight the more important effective features in each state, and enhance the generalization of fault diagnosis. In addition, to facilitate the training of the deep model and avoid the problem of model degradation that may be caused by the deepening of the network, a residual network is introduced to extract deeper abstract features and enhance the diagnostic effect of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] By reading the following detailed description of the preferred embodiments, various other advantages and benefits of the present disclosure will become clear to those of ordinary skill in the art. The accompanying drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present disclosure. Obviously, the following described drawings are only some embodiments of the present disclosure, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings. Moreover, throughout the drawings, the same reference numerals are used to represent the same components.
[0045] In the drawings:
[0046] Figure 1 is a flowchart of the gearbox unbalance fault diagnosis method based on the multi-scale deep attention reinforcement learning network proposed by the present invention;
[0047] Figures 2(a) to 2(i) is the vibration signal of the measured planetary gearbox under different health conditions;
[0048] Figure 3 is the structure diagram of the deep convolutional neural network;
[0049] Figure 4 is the flowchart of the agent interaction process;
[0050] Figures 5(a) to 5(i) is a schematic diagram of the multi-scale feature learning of nine states of the gearbox under the condition of 1200 rpm - 0 N.m.
[0051] The present invention will be further explained below with reference to the drawings and embodiments. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] The specific embodiments of the present disclosure will be described in more detail below with reference to the drawings. Although the specific embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0053] It should be noted that in the specification and claims, certain terms are used to refer to specific components. Those skilled in the art should understand that technicians may use different nouns to refer to the same component. The specification and claims do not use the difference in nouns as a way to distinguish components, but use the difference in the functions of components as the criterion for distinction. For example, the term "comprising" or "including" mentioned throughout the specification and claims is an open-ended term, so it should be interpreted as "including but not limited to". The subsequent description in the specification is the preferred embodiment for implementing the present disclosure, but the description is for the purpose of the general principles of the specification and is not used to limit the scope of the present disclosure. The protection scope of the present disclosure shall be subject to that defined by the appended claims.
[0054] For the convenience of understanding the embodiments of the present disclosure, the following will further explain with specific embodiments in conjunction with the accompanying drawings, and each accompanying drawing does not constitute a limitation on the embodiments of the present invention.
[0055] For better understanding, as Figures 1 to 5(i) shown, a gearbox unbalance fault diagnosis method based on a multi-scale deep attention reinforcement learning network includes the following steps:
[0056] S1: Signal acquisition, obtaining vibration signals of the gearbox in different health conditions under various working conditions, and constructing a training set and a test set based on the vibration signals, where the number of normal state samples is much larger than that of fault state samples;
[0057] S2: Environment simulation construction, according to the characteristics of the uneven data distribution of the health conditions of the gearbox, establishing an unbalanced classification Markov decision process, designing a reward function, and establishing a data environment simulation required for fault diagnosis under class imbalance;
[0058] S3: Establishing a multi-condition unbalanced deep reinforcement learning network, based on the training set, the agent and the environment continuously interact to train the agent to autonomously learn the optimal diagnosis strategy, and the agent includes at least two multi-scale feature deep attention networks with the same structure;
[0059] S4: Fault identification, inputting the samples in the test set into the trained agent one by one, and identifying the fault types of the gearbox according to the diagnosis strategy and analyzing the diagnosis results.
[0060] Further, in the S1, the vibration signal is sample-segmented, and a training sample signal and a test sample signal are constructed. Each sample length contains 2048 data points, and the amplitude is normalized to the range of [-1, 1].
[0061] Further, in S2, the Markov decision process includes a state space S, an action space A, and a reward function R. The state space S consists of all training samples in the training set, and each environmental state s corresponds to a training sample. The action space A is K actions corresponding to the health condition of the gearbox, A = {0, 1, …, K - 1}, where K represents the total number of fault types. The reward function R is: under the current diagnostic query, for the current environmental state s t ∈S, the agent executes the action a t , and compares whether the action a t is consistent with the label class under the current diagnostic query. If it is consistent, the environment returns a positive reward to the agent; otherwise, it returns a negative reward. A new reward strategy is defined to enable the agent to pay more attention to the learning of minority class samples under the unbalanced distribution of class data, and to achieve effective learning of the diagnostic strategy under class imbalance.
[0062] Suppose the training set D train = {D1; D1;...; D k ;...}, where D k represents the training subset of the k-th class, which can be expressed as: where (x i , l i ) represents the corresponding sample and label information, and n k represents the total number of samples in the k-th class subset. The reward function R is set by the following formula:
[0063]
[0064] In the formula, R(s t , a t , l t ) represents the environmental feedback reward value obtained by executing the action a t at the state s t , which is simplified and replaced by r t ; log2(·) is the logarithm function with base 2, which is used to evaluate the difference between classes; ρ k represents the class deviation degree of the k-th class, which is defined as the imbalance degree of the sample size of the k-th class subset in the training set D train relative to the sample size of the least class in D train , and can be obtained by the following formula:
[0065]
[0066] In the formula, D min represents the sample subset of the least class in D train , usually the sample set belonging to a certain fault class, and |·| represents taking the modulus, which here represents the sample size in the sample subset.
[0067] Further, in S3, a deep reinforcement learning network is established, and the agent interacts with the environment continuously to train the agent to autonomously learn the optimal diagnosis strategy, which specifically includes the following steps:
[0068] S3.1: Set the maximum number of autonomous learning rounds Episode. Episode refers to the transfer trajectory of the environment from the initial state s1 to the final state s T , Episode = {s1, a1, r1, s2, a2, r2,..., s T , a T , r T}, where T represents the termination time step, and the current training episode Episode is ended, and the next episode is started until the set interaction round is reached; Each round of autonomous learning contains T diagnostic inquiries, and each diagnostic inquiry corresponds to an environmental state;
[0069] S3.2: Randomly initiate a diagnostic inquiry to obtain the current state s t ∈S, input the corresponding training sample into the current Q network Eval-Net, and the agent selects the current action a t ∈A according to the linear annealing ∈-greedy algorithm. The environment returns the current reward r t according to the reward function R, and randomly initiates the next diagnostic inquiry. The state s t is converted to the next state s t+1 , and the generated experience data e = {s t , a t , r t , s t+1} is saved to the experience pool ;
[0070] S3.3: Repeat S3.2 until the T diagnostic inquiries in this round are completed, and output the total reward obtained from the diagnostic inquiries in this round. The total reward is k represents the kth diagnostic inquiry in this round of diagnostic inquiries, and r k represents the reward corresponding to the kth diagnostic inquiry;
[0071] S3.4: Randomly sample a predetermined batch of experience data e = {s from the experience pool t , a t , r t , s t+1}, based on the empirical data e, the gradient descent method is used to train the current Q-network Eval-Net, update the model parameters of Eval-Net, and perform a delayed update on the target Q-network Target-Net. The delayed update is to directly copy the model parameters of Eval-Net to the target Q-network Target-Net every C learning rounds;
[0072] S3.5: Start the next round of autonomous learning process, repeat S3.2 to S3.4 until the maximum number of autonomous learning rounds is reached, and the autonomous learning process ends;
[0073] S3.6: Save the model parameters with a higher total reward obtained in each round of diagnostic query during the autonomous learning process as the optimal diagnostic strategy learned by the agent.
[0074] Further, in S3.2, the action a t ∈A is selected according to the dynamic ∈-greedy algorithm, which includes the following sub-steps:
[0075] S3.2.1: Preset the initial value of ∈ as ∈ = 1, the change rate of ∈ as Δ∈ = 1 / 100000, and the minimum value of ∈ min = 0.01;
[0076] S3.2.2: Before each execution of the action a t a random number between [0, 1] is randomly generated. If the random number belongs to [0, ∈], a random action a is selected from the action space A to execute t ; if the random number belongs to (∈, 1], the action corresponding to the maximum output of the current Q-network Eval-Net is used as the execution action a t ;
[0077] S3.2.3: After each round of diagnostic query ends, the value of ∈ is dynamically updated. If ∈ is less than ∈ min , ∈ = ∈ min , otherwise ∈ = ∈ - Δ∈;
[0078] Further, the process of updating the model parameters of Eval-Net in S3.4 specifically includes the following steps:
[0079] S3.4.1: According to the empirical data e = {s sampled from the experience pool t , a t , r t , s t+1}, the current Q-network Eval-net outputs the current Q value as Q(s t , a t ; θ), and the target Q-network Target-net calculates the target Q value as Among them, θ and θ - are the network parameters of Eval-net and Target-net respectively; γ is the discount factor, γ ∈ [0, 1];
[0080] S3.4.2: Calculate the mean square error (MSE) of the current Q value and the target Q value:
[0081] L(θ) = E[(Q′ - Q(s t , a t ; θ)) 2 (3)
[0082] S3.4.3: Calculate the gradient of the mean square loss L(θ) with respect to the network parameter θ:
[0083]
[0084] Among them, represents the gradient operation with respect to the parameter θ; represents the empirical data (s t , a t , r t , s t+1 ) is sourced from represents randomly and uniformly sampling from the experience pool;
[0085] S3.4.4: Repeat S3.4.1 to S3.4.3, and update the model parameters of the current Q network Eval-net according to the gradient descent method.
[0086] The present invention automatically realizes feature extraction of various state scales through multi-scale convolution operations, enhancing the richness of feature extraction and the accuracy of diagnosis. Then, aiming at the redundancy of feature information under different channels, a channel attention mechanism is adopted to adaptively weight the feature information of different channels, recalibrate the multi-scale feature information, highlight the more important effective features in each state, so as to enhance the generalization of fault diagnosis. In addition, in order to facilitate the training of the deep model and avoid the problem of model degradation that may be caused by the deepening of the network, a residual network is introduced to extract deeper abstract features and enhance the diagnostic effect of the model.
[0087] In one embodiment, the gearbox is a planetary gearbox.
[0088] To further understand the present invention, in one embodiment, as Figure 1 shown, a gearbox unbalance fault diagnosis method based on a multi-scale deep attention reinforcement learning network mainly includes the following steps:
[0089] S1: Signal acquisition, obtaining vibration signals of the gearbox in different health conditions under various working conditions, constructing a training set and a test set based on the vibration signals, where the number of normal state samples is much larger than that of fault state samples;
[0090] Furthermore, the experiment simulated 9 different types of health conditions of the gearbox, including the healthy state (HEA), four types of gear fault states (cut tooth fault CTF, missing tooth fault MTF, root crack RCF, and surface wear SWF), and four types of bearing fault states (ball fault BWF, compound fault SWF, inner ring fault IRF, and outer ring fault ORF). The vibration signals collected under different health conditions were segmented into training vibration signals and test vibration signals, and then sample segmentation was performed. Each sample length contained 2048 data points, and the amplitude was normalized to the range of [-1, 1]. Training sample signals and test sample signals were constructed, making the number of normal state samples much larger than that of fault state samples. An example of a set of sample signals under different health conditions is shown as Figures 2(a) to 2(i) shown, representing partial vibration signals of the gearbox under multiple working conditions: Figure 2(a) healthy state HEA; Figure 2(b) cut tooth fault CTF; Figure 2(c) missing tooth fault MTF; Figure 2(d) root crack RCF; Figure 2(e) surface wear SWF; Figure 2(f) ball fault BF; Figure 2(g) compound fault CWF; Figure 2(h) inner ring fault IRF; Figure 2(i) outer ring fault ORF.
[0091] S2: Environment simulation construction, according to the characteristics of uneven data distribution of the health conditions of the gearbox, establishing an unbalanced classification Markov decision process, designing a reward function, and establishing a data environment simulation required for fault diagnosis under class imbalance;
[0092] Furthermore, in the above S2, the Markov decision process includes a state space S, an action space A, and a reward function R; the state space S consists of all training samples in the training set, and each environmental state s corresponds to a training sample; the action space A is K actions corresponding to the health conditions of the gearbox, A = {0, 1,..., K - 1}, where K represents the total number of fault types; the reward function R is: under the current diagnostic query, for the current environmental state s t ∈S, the intelligent agent executes the action a t , comparing whether the action a t is consistent with the label class under the current diagnostic query. If they are consistent, the environment returns a positive reward to the intelligent agent, otherwise it returns a negative reward. A new reward strategy is defined to enable the intelligent agent to pay more attention to the learning of minority class samples under the uneven distribution of class data volumes, and to achieve effective learning of the diagnostic strategy under class imbalance.
[0093] Assume the training set D train={D1; D1;...; D k ;...}, where D k represents the training subset of the k-th class, which can be expressed as: where (x i , l i ) represents the corresponding sample and label information, n k represents the total number of samples in the k-th class subset. The reward function R is set as follows:
[0094]
[0095] In the formula, R(s t , a t , l t ) represents the environmental feedback reward value obtained by executing the action a t at the state s t . Using r t for simplification and substitution; log2(·) is the logarithm function with base 2, which is used to evaluate the difference between classes; ρ k represents the class deviation degree of the k-th class, which is defined as the imbalance degree of the sample size of the k-th class subset in the training set D train relative to the sample size of the least class in D train , and can be obtained by the following formula:
[0096]
[0097] In the formula, D min represents the sample subset of the least class in D train , usually the sample set belonging to a certain fault class, and |·| represents taking the modulus, which here represents the sample size in the sample subset.
[0098] In this embodiment, the gearbox includes nine states, namely the healthy state, four gear fault states, and four bearing fault states. Therefore, the said K = 9, and the action space A = {0, 1,..., 8}.
[0099] S3: Establish a multi-condition unbalanced deep reinforcement learning network. Based on the training set, the agent and the environment interact continuously to train the agent to autonomously learn the optimal diagnosis strategy. The agent includes at least two multi-scale feature deep attention networks with the same structure. The scale feature deep attention network consists of a multi-scale feature layer, a channel attention layer, an Inception multi-scale layer, a max pooling layer, two residual modules, a global average pooling layer, and an output layer. One multi-scale feature deep attention network is the current Q network Eval-Net, and the other multi-scale feature deep attention network is the target Q network Target-Net, as Figure 3 shown, and its specific structural parameters are shown in Table 1;
[0100] Table 1 Structural detailed parameters of multi-scale feature depth attention
[0101]
[0102] Furthermore, in S3, a deep reinforcement learning network is established, and the agent interacts with the environment continuously to train the agent to autonomously learn the optimal diagnosis strategy, which specifically includes the following steps:
[0103] S3.1: Set the maximum number of autonomous learning rounds Episode. Episode refers to the transfer trajectory of the environment from the initial state s1 to the final state s T , Episode = {s1, a1, r1, s2, a2, r2, …, s T , a T , r T}, where T represents the termination time step, and the current training episode Episode ends, and the next episode starts until the set interaction rounds are reached; each round of autonomous learning contains T diagnostic inquiries, and each diagnostic inquiry corresponds to an environmental state;
[0104] S3.2: Randomly initiate a diagnostic inquiry to obtain the current state s t ∈S, input the corresponding training sample into the current Q network Eval-Net, and the agent selects the current action a t ∈A according to the linear annealing ∈-greedy algorithm. The environment returns the current reward r t according to the reward function R, and randomly initiates the next diagnostic inquiry. The state s t is converted to the next state s t+1 , and the generated experience data e = {s t , a t , rt, s t+1} is saved to the experience pool ;
[0105] S3.3: Repeat S3.2 until the T diagnostic inquiries in this round end, and output the total reward obtained from the diagnostic inquiries in this round. The total reward is k represents the kth diagnostic inquiry in this round of diagnostic inquiries, and r k represents the reward corresponding to the kth diagnostic inquiry;
[0106] S3.4: Randomly sample a predetermined batch of experience data e = {s from the experience pool t , a t , r t , s t+1}, based on the empirical data e, the gradient descent method is used to train the current Q-network Eval-Net, update the model parameters of Eval-Net, and perform a delayed update on the target Q-network Target-Net. The delayed update is to directly copy the model parameters of Eval-Net to the target Q-network Target-Net every C learning rounds;
[0107] S3.5: Start the next round of autonomous learning process, and repeat S3.2 to S3.4 until the maximum number of autonomous learning rounds is reached, and the autonomous learning process ends;
[0108] S3.6: Save the model parameters with a higher total reward obtained in each round of diagnostic query during the autonomous learning process as the optimal diagnostic strategy learned by the agent.
[0109] An example of the process flow diagram of the agent interaction is as Figure 4 shown.
[0110] Furthermore, in S3.2, the action a t ∈A is selected according to the dynamic ∈-greedy algorithm, which includes the following sub-steps:
[0111] S3.2.1: Preset the initial value of ∈ as ∈ = 1, the change rate of ∈ as Δ∈ = 1 / 100000, and the minimum value of ∈ min = 0.01;
[0112] S3.2.2: Before each execution of the action a t , a random number between [0, 1] is randomly generated. If the random number belongs to [0, ∈], a random action a t is selected from the action space A for execution; if the random number belongs to (∈, 1], the action corresponding to the maximum output of the current Q-network Eval-Net is used as the execution action a t ;
[0113] S3.2.3: After each round of diagnostic query ends, the value of ∈ is dynamically updated. If ∈ is less than ∈ min , ∈ = ∈ min , otherwise ∈ = ∈ - Δ∈;
[0114] S3.4: Randomly sample a certain batch of Batch_size = 128 empirical data e = {s from the experience pool, t , a t , r t , s t+1}, using the empirical data e, the current Q-network Eval-Net is trained by the gradient descent method, the model parameters of Eval-Net are updated, and the target Q-network Target-Net is updated with a delay. The delayed update is to directly copy the model parameters of Eval-Net to the target Q-network Target-Net every C = 100 learning rounds.
[0115] Further, the process of updating the model parameters of Eval-Net in S3.4 specifically includes the following steps:
[0116] S3.4.1: According to the empirical data e = {s , a t , r t , s t} sampled from the experience pool t+1 , the current Q-value Q(s t , a t ; θ) is output by the current Q-network Eval-net, and the target Q-value is calculated by the target Q-network Target-net as where θ and θ - are the network parameters of Eval-net and Target-net respectively; γ is the discount factor, γ ∈ [0, 1];
[0117] S3.4.2: Calculate the mean square error (MSE) of the current Q-value and the target Q-value:
[0118] L(θ) = E[(Q′ - Q(s t , a t ; θ)) 2 (3)
[0119] S3.4.3: Calculate the gradient of the mean square loss L(θ) with respect to the network parameter θ:
[0120]
[0121] where represents the gradient operation with respect to the parameter θ; represents that the empirical data (s t , a t , r t , s t+1 ) is from represents randomly and uniformly sampling from the experience pool ;
[0122] S3.4.4: Repeat S3.4.1 to S3.4.3, and update the model parameters of the current Q-network Eval-net according to the gradient descent method.
[0123] S4: Fault identification. Input the test set samples into the agent trained in S3 one by one. According to the diagnostic strategy learned by the agent, identify the gearbox fault type and analyze the diagnostic results.
[0124] In this embodiment, the method of the present invention is implemented under six different working conditions of the planetary gearbox: 1200 rpm - 0 N.m, 1800 rpm - 0 N.m, 1800 rpm - 3.66 N.m, 1800 rpm - 10.98 N.m, 2400 rpm - 0 N.m, and 3000 rpm - 0 N.m. Among them, the maximum number of healthy state samples is 500, and the number of the remaining fault samples is much less than 500. The specific state descriptions are shown in Table 2.
[0125] Table 2 Nine state descriptions of the gearbox in the imbalanced dataset
[0126]
[0127] Under each of the above working conditions, there are 500 test samples in the test set. The samples are input into the trained strategy for diagnostic testing in sequence, and compared with some methods. Three indicators in imbalanced classification, namely balanced accuracy bAcc, macro F1-score macro-F1, and G-mean value, are calculated as the comprehensive evaluation indicators of the diagnostic performance of the diagnostic model. The comparison of the diagnostic results is shown in Table 3.
[0128] Table 3 shows that: the method of the present invention has high balanced accuracy, macro F1-score, and G-mean value for gearbox imbalance fault diagnosis under different working conditions. The fault diagnosis accuracy exceeds 99.4%. It has certain advantages compared with some traditional methods and deep convolutional neural networks, and is more suitable for gearbox imbalance fault diagnosis, which also proves the significance and effectiveness of the present invention.
[0129] Comparison of the average diagnostic results of the methods in Table 3 (%)
[0130]
[0131] In addition, in order to visually observe the learning effect of the multi-scale feature layer in the MFDARL model of this chapter, Figures 5(a) to 5(i)It shows the multi-scale feature learning of nine states of the gearbox under the condition of 1200 rpm - 0 N·m, which are: Figure 5(a) healthy state HEA; Figure 5(b) chipped tooth fault CTF; Figure 5(c) missing tooth fault MTF; Figure 5(d) tooth root crack RCF; Figure 5(e) tooth surface wear SWF; Figure 5(f) rolling element fault BF; Figure 5(g) compound fault CWF; Figure 5(h) inner race fault IRF; Figure 5(i) outer race fault ORF. Given that convolutional kernels of different sizes can extract features in different frequency bands, it can be seen that its feature map is similar to the usual time-scale characterization mapping diagram. From Figures 5(a) to 5(i) It can be known that convolutional operations of different scales can capture different feature information, and the model can adaptively extract effective feature information at different scales according to vibration signals of different states, improving the fault diagnosis effect. For example, for the fault state MTF (Figure 5(a)), scale 3 can extract more feature information. For the fault state MTF (Figure 5(a)), scales 2 and 3 can extract more feature information. For the fault state MTF (Figure 5(a)), scales 1 to 3 can all extract more feature information. At the same time, there is a certain redundancy in different channel information.
[0132] To sum up, by means of the above technical solutions of the present invention, multi-scale convolutional operations are used to automatically achieve multi-scale feature extraction for each state, and combined with a deep reinforcement learning model, the powerful perception of deep learning is used for automatic feature extraction and the autonomous decision-making ability of reinforcement learning to autonomously learn the best diagnosis strategy, and still has high generalization under the condition of relatively limited training data.
[0133] Although the implementation schemes of the present disclosure have been described above in conjunction with the accompanying drawings, the present disclosure is not limited to the above specific implementation schemes and application fields. The above specific implementation schemes are merely illustrative and guiding, rather than restrictive. Those of ordinary skill in the art can also make many forms under the inspiration of this specification and without departing from the scope protected by the claims of the present disclosure, and all of these belong to the scope of protection of the present invention.
Claims
1. A gearbox unbalance fault diagnosis method based on a multi-scale depth attention reinforcement learning network, characterized in that It includes the following steps: Step S1: Signal acquisition. Vibration signals of the gearbox in different health conditions under various working conditions are obtained, and a training set and a test set are constructed based on the vibration signals, where the number of normal state samples is more than that of fault state samples; Step S2: Environment simulation construction. A Markov decision process for imbalance classification is established, a reward function is designed, and a data environment simulation required for fault diagnosis under class imbalance is established; Step S3: Establish a multi-condition imbalance deep reinforcement learning network. Based on the training set, the agent continuously interacts with the environment to train the agent to autonomously learn the optimal diagnosis strategy. The agent includes at least two multi-scale feature deep attention networks with the same structure; Step S4: Fault identification. The samples in the test set are input into the trained agent one by one, and the gearbox fault types are identified according to the diagnosis strategy, and the diagnosis results are analyzed; Among them, in step S1, the fault states include gear fault states and bearing fault states; In step S2, the Markov decision process includes a state space S, an action space A, and a reward function R; the state space S consists of all training samples in the training set, and each environmental state s corresponds to a training sample; the action space A is K actions corresponding to the health condition of the gearbox, , where the total number of failure types is ; the reward function R is: under the current diagnostic query, corresponding to the current environmental state s t ∈S, the agent executes the action a t , compare whether the action a t is consistent with the label class under the current diagnostic query. If it is consistent, the environment returns a positive reward to the agent, otherwise it returns a negative reward, forming a reward strategy to enable the agent to pay more attention to the learning of minority class samples under the unbalanced distribution of class data, and to achieve effective learning of the diagnostic strategy under class imbalance; Among them, the training set , where represents the training subset of the k-th class, which is expressed as: , where represents the corresponding sample and label information, represents the total number of samples in the k-th class subset, and the reward function R is set by the following formula: ; wherein, represents the state perform an action when the obtained environmental feedback reward value, in order to simplify and substitute; is the logarithm function with base 2, used to evaluate the difference between classes; represents the class deviation degree of the k-th class, which is the imbalance degree of the sample size of the k-th class subset in the training set relative to the sample size of the least class in obtained by the following formula: ; In the formula, represents the sample subset of the least class in the training set , and represents taking the modulus, that is, it represents the sample size in the sample subset; In step S3, the multi-scale feature deep attention network is composed of a multi-scale feature layer, a channel attention layer, an Inception multi-scale layer, a max pooling layer, two residual modules, a global average pooling layer, and an output layer.
2. The gearbox unbalance fault diagnosis method according to claim 1, wherein In step S1, the vibration signal is segmented into samples, and training sample signals and test sample signals are constructed. Each sample length contains 2048 data points, and the amplitude is normalized to the range of [-1, 1].
3. The gearbox unbalance fault diagnosis method according to claim 1, wherein The agent includes at least two multi-scale feature deep attention networks of the current Q network Eval-Net with the same structure and the multi-scale feature deep attention network of another target Q network Target-Net.
4. The gearbox unbalance fault diagnosis method according to claim 3, wherein Step S3 includes the following steps: Step S3.1: Set the maximum number of self-learning episodes Episode. The maximum number of self-learning episodes Episode refers to the transition trajectory of the environment from the initial state to the final state . Among them, represents the termination time step, and then the current training episode is ended , and the next episode starts until the set number of interaction rounds is reached; each round of self-learning includes T diagnostic inquiries, and each diagnostic inquiry corresponds to an environmental state; Step S3.2: Randomly initiate a diagnostic query to obtain the current state , input the corresponding training sample into the current Q-network Eval-Net, and the agent selects the current action according to the linear annealing greedy algorithm , the environment returns the current reward according to the reward function R , and randomly initiate the next diagnostic query, and the state transitions to the next state , save the generated experience data to the experience pool ; Step S3.3: Repeat Step S3.2 until the T - time diagnostic inquiry of this round ends, and output the total reward obtained from this round of diagnostic inquiry. The total reward is , indicating the th diagnostic inquiry in this round of diagnostic inquiry, indicating the th diagnostic inquiry - corresponding reward; Step S3.4: From the experience pool Randomly sample predetermined batches of empirical data , based on empirical data , use the gradient descent method to train the current Q network Eval-Net, update the Eval-Net model parameters, and delay the update of the target Q network Target-Net. The delayed update is to directly copy the model parameters of Eval-Net to the target Q network Target-Net after every C learning rounds; Step S3.5: Start the next round of autonomous learning process, and repeat steps S3.2 to S3.4 until the maximum number of autonomous learning rounds is reached, and the autonomous learning process ends; Step S3.6: Save the model parameters with a higher total reward obtained in each round of diagnostic queries during the autonomous learning process as the optimal diagnosis strategy learned by the agent.
5. The gearbox imbalance fault diagnosis method according to claim 4, wherein step S3.2 includes the following steps: Step S3.2.1: Preset in advance Initial value , Rate of change , Minimum value ; Step S3.2.2: Each time before performing the action , randomly generate a random number. If the random number belongs to , randomly select an action from the action space to perform the action ; if the random number belongs to , use the action corresponding to the maximum output of the current Q-network Eval-Net as the action to be performed ; Step S3.2.3: After each round of diagnostic inquiry, dynamically update the value. If is less than , otherwise .
6. The method for diagnosing the imbalance fault of a gearbox according to claim 4, characterized in that, Step S3.4 includes the following steps: Step S3.4.1: According to the experience data sampled from the experience pool , , the current Q-network Eval-net outputs the current Q-value as , and the target Q-network Target-net calculates the target Q-value as , where and are the network parameters of Eval-net and Target-net respectively; is the discount factor, ∈ [0, 1]; Step S3.4.2: Calculate the mean square error (MSE) of the current Q value and the target Q value; ; Step S3.4.3: Calculate the gradient of the mean square loss L(θ) with respect to the network parameter θ; ; Among them, represents the gradient operation on the parameter ; represents the empirical data derived from U( ); U( ) means randomly and uniformly sampling from the empirical pool ; Step S3.4.4: Repeat steps S3.4.1 to S3.4.3, and update the model parameters of the current Q network Eval-net according to the gradient descent method.
7. The gearbox unbalance fault diagnosis method according to claim 1, characterized in that, The gear fault states include cut tooth fault (CTF), missing tooth fault (MTF), root crack fault (RCF), and tooth surface wear fault (SWF). The bearing fault states include rolling element fault (BWF), compound fault (SWF), inner ring fault (IRF), and outer ring fault (ORF).
Citation Information
Patent Citations
Planetary gear box fault diagnosis method based on deep reinforcement learning model
CN112633245A
Fault diagnosis model self-learning method based on asynchronous parallel reinforcement learning
CN112801272A