Bearing Fault Diagnosis Method Based on Neural Networks and Multi-Criterion Preference Consensus

By employing neural networks and multi-criteria preference consensus methods, combined with physical marginal value functions and reinforcement learning consensus-reaching processes, the adaptability problem of bearing fault diagnosis under complex working conditions is solved. This achieves high-precision, low-round group fault diagnosis, improving the accuracy and interpretability of bearing fault diagnosis.

CN121834522BActive Publication Date: 2026-05-26TAIYUAN NORMAL UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
TAIYUAN NORMAL UNIV
Filing Date
2026-03-11
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing bearing fault diagnosis methods are not adaptable enough to complex working conditions, making it difficult to accurately capture early and subtle faults. Furthermore, the lack of dynamic adaptive adjustment mechanism in multi-model fusion results in a lack of reliability and robustness of the diagnostic results.

Method used

A method based on neural networks and multi-criteria preference consensus is adopted, combining physical marginal value function and neural network attention mechanism. Through a consensus-reaching process driven by reinforcement learning, the adaptive optimization of multi-expert fault diagnosis results is achieved. The initial evaluation probability is generated by a neural network multivariate decision-making auxiliary model and dynamically adjusted through the reinforcement learning consensus-reaching process, ultimately generating the collective probability of bearing state.

Benefits of technology

It improves the accuracy, robustness, and interpretability of bearing fault diagnosis, and is suitable for efficient and stable diagnosis under complex working conditions, enabling accurate early warning of faults and low-cost predictive maintenance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121834522B_ABST
    Figure CN121834522B_ABST
Patent Text Reader

Abstract

This invention relates to the field of bearing fault diagnosis technology, specifically to a bearing fault diagnosis method based on neural networks and multi-criteria preference consensus. The method includes: data acquisition and feature extraction, data preprocessing, matching a fault diagnosis model, preliminary diagnosis and probability generation, final diagnosis and probability generation, and bearing state decision-making. This invention provides a multivariate decision-making auxiliary model that integrates physical marginal value functions and neural network attention mechanisms, combined with a reinforcement learning-driven consensus-building process, to achieve adaptive optimization of multi-expert fault diagnosis results. This not only solves the problems of existing bearing fault diagnosis methods being unable to handle complex working conditions, lacking physical interpretability, relying heavily on single decision-making patterns, and using static weight allocation, but also improves the accuracy, robustness, and interpretability of bearing fault diagnosis, achieving high-precision, low-round group fault diagnosis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of bearing fault diagnosis technology, and specifically to a bearing fault diagnosis method based on neural networks and multi-criteria preference consensus. Background Technology

[0002] In industrial rotating machinery, bearings, as core transmission components, directly determine the safety and reliability of the equipment. Statistics show that approximately 30% of rotating machinery failures originate from bearing damage. However, because early, subtle bearing fault characteristics are easily masked by strong noise, traditional bearing fault diagnosis methods face significant challenges.

[0003] Currently, most bearing fault diagnosis methods are intelligent diagnostic methods based on vibration signal analysis, such as spectrum analysis, wavelet packet decomposition, empirical mode decomposition, and end-to-end fault identification methods based on convolutional neural networks, recurrent neural networks, and deep autoencoders. These data-driven deep learning models can achieve high classification accuracy in experimental environments with sufficient labeled data and relatively simple working conditions, but they still have obvious limitations in actual industrial scenarios.

[0004] On the one hand, existing diagnostic methods mostly rely on a single model or a fixed combination strategy, lacking a dynamic adaptation mechanism to cope with complex failure modes, resulting in a significant lack of adaptability of the model in diverse failure scenarios; moreover, when faced with early weak faults or complex faults, the feature representation ability of a single model is limited, making it difficult to accurately capture key information about fault evolution, thus resulting in a lack of reliability and robustness of diagnostic results.

[0005] On the other hand, while existing diagnostic methods also employ multi-model collaborative decision-making mechanisms, current multi-model fusion methods mostly adopt static weight allocation, lacking a dynamic adaptive adjustment mechanism based on fault characteristics and model confidence. This results in poor matching between fused weights and actual fault states. Moreover, when discrepancies arise between different models or feature dimensions, the lack of an intelligent consensus optimization mechanism to coordinate decisions reduces the stability and reliability of diagnostic results. For example, the weight setting of traditional multi-criteria decision-making relies on manual prior knowledge and lacks the ability to autonomously learn and optimize from historical diagnostic results. Its static parameters cannot adapt to the dynamic changes in bearing operating conditions, leading to a decrease in decision reliability under new faults or variable operating conditions. Summary of the Invention

[0006] To address the shortcomings of the existing technologies, this invention provides a bearing fault diagnosis method based on neural networks and multi-criteria preference consensus, which solves the problems of existing bearing fault diagnosis methods being unable to cope with complex working conditions, lacking physical interpretability, relying on a single decision-making mode, and using static weight allocation.

[0007] This invention provides a multivariate decision-making assistance model that integrates physical marginal value functions and neural network attention mechanisms, and combines a consensus-reaching process driven by reinforcement learning to achieve adaptive optimization of multi-expert fault diagnosis results, thereby improving the accuracy, robustness, and interpretability of bearing fault diagnosis.

[0008] This invention provides a bearing fault diagnosis method based on neural networks and multi-criteria preference consensus, comprising the following steps:

[0009] S1. Data Acquisition and Feature Extraction: Acquire the model information and vibration signal of the target bearing, and extract the multi-domain features of the target bearing from the vibration signal;

[0010] S2. Data Preprocessing: Standardize the multi-domain feature data of the target bearing and output the standardized feature vector of the target bearing. ,in For the i-th eigenvalue, For feature dimensions;

[0011] S3. Matching fault diagnosis model: Based on the model information of the target bearing, select the bearing condition diagnosis model that matches the target bearing model from the bearing condition diagnosis model database. The bearing condition diagnosis model includes the trained neural network multivariate decision aid NN-MCDA model and the trained reinforcement learning consensus achievement process RL-CRP model.

[0012] S4. Preliminary diagnosis and probability generation: Input the standardized feature vector of the target bearing into the trained neural network multivariate decision-making auxiliary NN-MCDA model that matches the target bearing model, and output multiple expert initial evaluation probabilities of the target bearing under different states. The trained neural network multivariate decision-making auxiliary NN-MCDA model is a hybrid model composed of neural network and multivariate decision analysis.

[0013] S5. Final Diagnosis and Probability Generation: The training reinforcement learning consensus reaching process RL-CRP model is used to input the initial evaluation probabilities of multiple experts for the target bearing under different states into the target bearing model. The output is the final evaluation probability matrix of multiple experts for the target bearing under different states and the weights of each expert in the last iteration. The training reinforcement learning consensus reaching process RL-CRP model is a hybrid model composed of reinforcement learning and consensus reaching process.

[0014] S6. Bearing State Decision: Based on the probability matrix of multiple experts' final evaluations of the target bearing under different states and the weights of each expert in the last iteration, the collective probability of the target bearing under different states is generated, and the state with the highest collective probability is selected as the final state of the target bearing. The expression for the collective probability of the target bearing under different states is as follows:

[0015]

[0016] In the formula, Let N be the collective probability of all experts for the m-th state of the k-th sample, where N is the number of experts. The normalized weights of the nth expert in the final iteration of the RL-CRP model during the post-training reinforcement learning consensus-reaching process satisfy nonnegativity. With normality ; Let be the probability of the nth expert's evaluation of the mth state of the kth sample.

[0017] Preferably, the multi-domain features of the target bearing include three types of features: time domain, frequency domain, and time-frequency domain. The time domain features include kurtosis, peak value, root mean square value, variance, waveform factor, impulse index, margin index, peak factor, skewness, amplitude factor, zero-crossing rate, and energy operator. The frequency domain features include frequency amplitude, centroid frequency, frequency variance, spectral kurtosis, harmonic component energy, sideband energy ratio, spectral peak density, frequency band energy, spectral entropy, and spectral standard deviation. The time-frequency domain features include wavelet packet energy, IMF energy, time-frequency matrix singular values, instantaneous frequency variance, marginal spectral energy, wavelet entropy, Hilbert spectral entropy, and time-frequency clustering.

[0018] In addition, in step S2, Z-score is used to standardize the multi-domain feature data of the target bearing; the bearing status includes four status categories: normal, inner ring fault, outer ring fault, and rolling element fault.

[0019] Preferably, the trained neural network multivariate decision-making auxiliary NN-MCDA model is obtained by training the neural network multivariate decision-making auxiliary NN-MCDA model, wherein the neural network multivariate decision-making auxiliary NN-MCDA model includes several parallel neural network multivariate decision-making auxiliary NN-MCDA sub-models; each neural network multivariate decision-making auxiliary NN-MCDA sub-model includes an input layer, a marginal value mapping layer, a multilayer perceptron module, a utility fusion layer, and an output layer, specifically as follows:

[0020] The input layer takes the standardized feature vector of the target bearing as input.

[0021] The marginal value mapping layer receives the standardized feature vector of the target bearing and transforms each feature value in the standardized feature vector into a contribution to each state through the marginal value function, obtaining the linear contribution of each feature to each state. The marginal value function is shown in the following equation:

[0022]

[0023] In the formula, m is the state index symbol, and d is the polynomial degree index symbol. Let be the linear contribution of the i-th feature of the k-th sample to the m-th state. Let be the polynomial degree of the i-th feature with respect to the m-th state. The fitting coefficients of the d-th term for the i-th feature to the m-th state are... It is the i-th feature of the k-th sample;

[0024] The multilayer perceptron module receives the standardized feature vector of the target bearing, captures the high-order interactions and nonlinear relationships between features, and outputs the nonlinear contribution vector of the target bearing to each state.

[0025] The utility fusion layer receives the outputs of the marginal value mapping layer and the multilayer perceptron module. It integrates these outputs using a comprehensive utility function to obtain the comprehensive utility value of the target bearing for each state. The comprehensive utility function is shown in the following equation:

[0026]

[0027] In the formula, Let be the combined utility value of the k-th sample with respect to the m-th state; Let be the tradeoff coefficient for the m-th state. ; For feature dimension, Let i be the feature weight of the i-th feature in the m-th state. This represents the nonlinear contribution of the k-th sample to the m-th state.

[0028] The output layer receives the comprehensive utility value of the target bearing for each state and converts it into the initial evaluation probability of the target bearing for each state using the Softmax function. The conversion formula is shown below:

[0029]

[0030] In the formula, Let M be the initial evaluation probability of the k-th sample in the m-th state, and M be the number of states.

[0031] Preferably, the multilayer perceptron module includes an input layer, a hidden layer, and an output layer. A CBAM submodule is embedded between the hidden layer and the output layer. The CBAM submodule consists of channel attention units and spatial attention units, specifically:

[0032] The input layer receives the standardized feature vector of the target bearing;

[0033] The hidden layer performs a linear transformation using the weight matrix and bias vector, projecting the input of the input layer into a high-dimensional space. Then, the ReLU activation function is applied to introduce non-linearity, capturing feature interactions and generating preliminary high-dimensional features. The linear transformation is shown in the following equation:

[0034]

[0035] The nonlinear activation is shown in the following equation:

[0036]

[0037] In the formula, This represents the combined value of the k-th sample in the hidden layer. Let be the weight matrix of the i-th feature in the hidden layer. The bias vector of the hidden layer. This is the output after activating the k-th sample;

[0038] The CBAM submodule reshapes the initial high-dimensional features output from the hidden layer into feature maps and inputs these feature maps into the channel attention unit. The channel attention unit weights the feature maps along the channel dimension based on the global average pooling and global max pooling results. The channel-weighted feature maps are then input into the spatial attention unit, which further enhances the channel-weighted feature maps based on the average pooling and max pooling results at each spatial location. Finally, the initial high-dimensional features are multiplied by the channel attention weights to obtain the enhanced channel features. These enhanced channel features are then multiplied by the spatial attention weights to output the weighted global high-dimensional features, as shown in the expression below:

[0039]

[0040]

[0041]

[0042]

[0043] In the formula, F is the input feature map. This indicates that global average pooling is performed on each channel of F. This indicates that global max pooling is performed on each channel of F, and MLP is a shared multilayer perceptron. It is the Sigmoid activation function. For channel attention weights, Spatial attention weights, Indicates to and Perform the splicing operation; Indicates the kernel size as Convolution operations are used to capture local relationships in spatial locations; To enhance channel features, The weighted global high-dimensional features;

[0044] The output layer receives the weighted global high-dimensional features output by the CBAM submodule, configures independent weight parameters for each state to be diagnosed, and calculates a feature pattern matching weight vector for each state through a linear transformation between the weight vector and the weighted global high-dimensional features. This achieves the correlation measurement between the target bearing and each state, and outputs the nonlinear contribution vector of the target bearing to each state, i.e., [ The formula for calculating the linear transformation between the weight vector and the weighted global high-dimensional features is as follows:

[0045]

[0046] In the formula, Let be the nonlinear contribution of the k-th sample to the m-th state. Let m be the weight matrix for the m-th state. The weighted global high-dimensional feature of the k-th sample. This is the bias vector.

[0047] Preferably, the channel attention unit includes:

[0048] The global average pooling subunit calculates the average value for each channel of the feature map to obtain the global characteristics of that channel.

[0049] The global max pooling subunit maximizes each channel of the feature map to capture the response at significant activation locations;

[0050] The shared multilayer perceptron subunit receives the outputs of the global average pooling subunit and the global max pooling subunit, and learns the relationship between pooling features.

[0051] The element-sum subunit adds the outputs of the global average pooling subunit and the global max pooling subunit in the shared multilayer perceptron subunit;

[0052] The output sub-units are processed by using the Sigmoid activation function to compress the summation results of the element-sum sub-units to the range of (0, 1), thus obtaining the channel attention weights.

[0053] Spatial attention units include:

[0054] The global average pooling sub-unit calculates the mean of all channels at each location in space;

[0055] The global max-pooling sub-unit calculates the maximum value of all channels at each location in the space.

[0056] The splicing sub-unit concatenates the outputs of the global average pooling sub-unit and the global max pooling sub-unit to obtain a two-channel feature map.

[0057] The convolutional computation subunit uses a convolutional kernel to convolve the two-channel feature maps to obtain a spatial weight map.

[0058] The output sub-units are compressed into the range (0, 1) using the Sigmoid activation function to obtain the spatial attention weights.

[0059] Preferably, the trained neural network multivariate decision-making auxiliary NN-MCDA model for each bearing model consists of multiple parallel neural network multivariate decision-making auxiliary NN-MCDA sub-models. To enable each neural network multivariate decision-making auxiliary NN-MCDA sub-model to form differentiated diagnostic characteristics, each neural network multivariate decision-making auxiliary NN-MCDA sub-model uses the same network structure and training process for independent parameter optimization, but they differ in training sample sampling strategy, feature constraints, and loss function parameters. The specific training steps for each trained neural network multivariate decision-making auxiliary NN-MCDA model are as follows:

[0060] S41. Obtain multiple sets of historical bearing vibration samples of this model under different conditions, and extract the multi-domain features of each set of historical bearing vibration samples under different conditions, wherein the number of historical bearing vibration samples under each condition is equal.

[0061] S42. Standardize the multi-domain features of each group of historical bearing vibration samples under different states to obtain the standardized feature vector of each group of historical bearing vibration samples under different states.

[0062] S43. Divide multiple sets of historical bearing vibration samples under the same state into a training set and a validation set according to the proportion. Each set of historical bearing vibration samples in the training set / validation set carries its corresponding standardized feature vector.

[0063] S44. Initialize multiple neural network multivariate decision-making auxiliary NN-MCDA sub-models, and set differentiated training strategies for each neural network multivariate decision-making auxiliary NN-MCDA sub-model from three levels. Specifically: First, set differentiated sample sampling weights or resampling probabilities to weight or resample the samples in the training set, thereby prompting each neural network multivariate decision-making auxiliary NN-MCDA sub-model to focus on different data distributions; Second, set feature subset masks for each neural network multivariate decision-making auxiliary NN-MCDA sub-model to constrain each neural network multivariate decision-making auxiliary NN-MCDA sub-model to focus on specific feature types; Finally, set differentiated class weights in the loss function to adjust the misclassification penalty for samples in different states, thereby enhancing the recognition ability of each neural network multivariate decision-making auxiliary NN-MCDA sub-model on the preset emphasis state.

[0064] S45. Input all the training sets corresponding to each state into each neural network multivariate decision aid NN-MCDA sub-model for independent training. During the training process, the samples in the training set are weighted and / or resampled according to the sample sampling weights set by the neural network multivariate decision aid NN-MCDA sub-model. This enables each neural network multivariate decision aid NN-MCDA sub-model to learn the complete mapping relationship from the input standardized feature vector to all states. After training, each neural network multivariate decision aid NN-MCDA sub-model becomes an "expert" that can work independently and output an evaluation probability vector covering all states for each input sample.

[0065] S46. Use the validation set corresponding to each state to evaluate each trained neural network multivariate decision-making auxiliary NN-MCDA sub-model. Based on the evaluation results, select the parameters with the best overall performance on the validation set for each neural network multivariate decision-making auxiliary NN-MCDA sub-model, and solidify it into a trained neural network multivariate decision-making auxiliary NN-MCDA sub-model with specific expertise. During the evaluation, comprehensively consider the recall rate of the neural network multivariate decision-making auxiliary NN-MCDA sub-model in its emphasized state, as well as the false alarm rate and / or specificity in its non-emphasized state, to ensure the effectiveness and reliability of its expertise.

[0066] S47. Integrate multiple trained neural network multivariate decision-making auxiliary NN-MCDA sub-models into a complete trained neural network multivariate decision-making auxiliary NN-MCDA model.

[0067] Preferably, the post-trained reinforcement learning consensus reaching process (RL-CRP) model is obtained by training the reinforcement learning consensus reaching process (RL-CRP) model, which includes a four-layer structure, specifically:

[0068] The state construction layer constructs an environmental state vector based on the initial evaluation probabilities of multiple experts under different states of the target bearing. The environmental state vector includes the consensus index of each expert, the group harmony degree, and the iteration round.

[0069] The decision-making strategy layer inputs the environmental state vector into the reinforcement learning architecture for iterative optimization, generating multiple expert evaluation probability matrices for the target bearing under different states. The reinforcement learning architecture includes two independent intelligent agents, namely... Agent and Agent, among which The agent is used to generate feedback parameters to control the adjustment range of expert opinions, so that the evaluation results of each expert converge to the group average. The agent is used to generate expert weight vectors and normalization is ensured through Softmax activation.

[0070] The reward evaluation layer calculates the reward function for each intelligent agent based on group harmony, generation round, and changes in expert weights, and updates the policy networks of the two intelligent agents respectively. The agent's reward function focuses on three main objectives: encouraging high harmony, low iteration rounds, and smooth consensus changes. The agent's reward function encourages high-consensus experts to have higher weights, while constraining the weight changes to be smooth.

[0071] Iterative optimization layer, when all expert consensus indices are higher than the consensus threshold The optimization may terminate when the number of iterations exceeds the upper limit, and the output will be the probability matrix of the final evaluation of the target bearing by multiple experts under different states, as well as the weight of each expert in the last iteration.

[0072] Preferably, the step of constructing an environmental state vector based on multiple initial expert evaluation probabilities of the target bearing under different states is as follows:

[0073] S511. Treat each vibration signal of the target bearing as a sample, and transform the multiple initial expert evaluation probabilities of each sample under different states into a decision matrix. The decision matrix for each sample is shown below:

[0074]

[0075] In the formula, Let be the decision matrix for the k-th sample. Let be the probability of the nth expert's evaluation of the mth state;

[0076] S512. Based on the decision matrix of all samples, construct an environmental state vector that includes consensus index, harmony degree, and iteration round. ,in Let be the relational consensus index of the nth expert on all samples in the t-th iteration. Let be the harmony degree at iteration t; the steps for calculating the relation-level consensus index in each iteration are as follows:

[0077] S5121. Calculate the element-level consensus index. The calculation formula is as follows:

[0078]

[0079]

[0080] In the formula, Let be the element-level consensus index of the nth expert on the mth state of the kth sample. Let be the probability of the nth expert's evaluation of the mth state of the kth sample. Let R be the group-weighted average evaluation probability of all experts for the m-th state of the k-th sample, R be the set of evaluation probabilities of all experts for all states of all samples, and N be the number of experts. For the normalized weight of the nth expert, and ;

[0081] S5122. Calculate the scheme-level consensus index. The calculation formula is as follows:

[0082]

[0083] In the formula, Let M be the scheme-level consensus index of the nth expert on all states of the kth sample, where M is the number of states;

[0084] S5123. Calculate the relation-level consensus index. The calculation formula is as follows:

[0085]

[0086] In the formula, Let K be the relational consensus index of the nth expert on all samples, and K be the number of samples.

[0087] The formula for calculating harmony in each iteration is shown below:

[0088]

[0089] In the formula, HD represents the harmony degree, and t is the iteration round. Let be the probability of the nth expert's evaluation of the mth state of the kth sample in the tth iteration.

[0090] Preferred, The agent's action space is defined as Its value is generated through the Actor network, as shown below:

[0091]

[0092] In the formula, In the t-th iteration The actions of the agent In the t-th iteration Feedback parameters generated by the agent for Agent's strategy network, The environment state vector is the input policy network. for The set of learnable parameters of the agent in the policy network; for The output layer weight matrix of the proxy, where ReLU is the modified linear unit activation function. for The hidden layer weight matrix of the proxy, Indicates will The state feature mapping function is converted into a feature vector that can be processed by the network. for The hidden layer bias vector of the proxy. for The output layer bias vector of the proxy uses the Sigmoid activation function to ensure... ;

[0093] The agent's reward function is formed by a weighted combination of harmony, iteration rounds, and changes in the consensus index, as shown in the following formula:

[0094]

[0095] In the formula, In the t-th iteration The agent's reward function, , , These are the weighting coefficients;

[0096] The agent's action space is defined as Its value is generated through the Actor network, as shown below:

[0097]

[0098] In the formula, In the t-th iteration The actions of an agent; In the t-th iteration The expert weight vector generated by the agent, In the t-th iteration The normalized weights of the nth expert generated by the agent; for Agent's strategy network, The environment state vector is the input policy network. for The set of learnable parameters of the agent in the policy network; for The output layer weight matrix of the proxy, where ReLU is the modified linear unit activation function. for The hidden layer weight matrix of the proxy, Indicates will The state feature mapping function is converted into a feature vector that can be processed by the network. for The hidden layer bias vector of the proxy. for The proxy's output layer bias vector, using Activation function, ensure and ;

[0099] The agent's reward function is formed by the weighted sum of expert weights and consensus index, minus the squared penalty for changes in weights, as shown in the following formula:

[0100]

[0101] In the formula, In the t-th iteration The agent's reward function; In the t-th iteration The normalized weights of the nth expert generated by the agent. and ; This is an evaluation item for the immediate effect of the strategy, used to encourage the allocation of high weights to high-consensus experts in order to directly improve the level of collective consensus. A coefficient used to control the intensity of punishment; As a constraint, the overall change range of the weight strategy is quantified through L2 norm, and drastic changes in the weight vector are penalized to ensure the continuity and robustness of weight allocation; The change in weight represents the weight of the first... Round iteration and the first Overall changes in weight allocation strategy during rounds of iteration ; It is a quantitative indicator of the "overall strategic change range";

[0102] In addition, in each iteration, Agent and After obtaining the feedback parameters and expert weight vector based on the same environmental state vector, the agent first calculates the group-weighted average evaluation probability of all experts for the m-th state of the k-th sample in the current iteration based on the expert weight vector. Then, the consensus index of each expert in the current iteration is calculated, and for those experts whose consensus index is ≤ consensus threshold, the consensus index is then calculated. Experts use feedback parameters to update the evaluation probability in the next iteration, and finally calculate the environment state vector in the next iteration based on the updated evaluation probability. In each iteration, for consensus index ≤ consensus threshold... The experts update their evaluation probabilities for the next iteration using the following formula:

[0103]

[0104] In the formula, Let be the probability of the nth expert's evaluation of the mth state of the kth sample in the t-th iteration. Let be the group-weighted average evaluation probability of all experts for the m-th state of the k-th sample in the t-th iteration. In the t-th iteration Feedback parameters generated by the agent.

[0105] Preferably, the training steps of the reinforcement learning consensus-reaching process (RL-CRP) model corresponding to each type of bearing are as follows:

[0106] S521. Obtain multiple sets of historical bearing vibration samples of this model under different conditions, and extract the multi-domain features of each set of historical bearing vibration samples, wherein the number of historical bearing vibration samples under each condition is equal.

[0107] S522. Standardize the multi-domain features of each group of historical bearing vibration samples to obtain the standardized feature vector of each group of historical bearing vibration samples.

[0108] S523. Input the standardized feature vectors of all historical bearing vibration samples into the trained neural network multivariate decision-making auxiliary NN-MCDA model respectively, and obtain multiple initial expert evaluation probabilities for each group of historical bearing vibration samples under different states.

[0109] S524. Divide all historical bearing vibration samples into training set and validation set according to the proportion. The number of historical bearing vibration samples in each state in the training set and validation set is equal. Each group of historical bearing vibration samples in the training set and validation set carries multiple initial expert evaluation probabilities in different states.

[0110] S525. Input the initial expert evaluation probabilities of each group of historical bearing vibration samples in the training set under different states into the reinforcement learning consensus reaching process (RL-CRP) model for training, and obtain the initial reinforcement learning consensus reaching process (RL-CRP) model.

[0111] S526. Input the initial expert evaluation probabilities of each group of historical bearing vibration samples in different states in the validation set into the initial reinforcement learning consensus reaching process RL-CRP model. Evaluate the consensus degree, classification accuracy and convergence rounds of the initial reinforcement learning consensus reaching process RL-CRP model on the validation set. When the evaluation results meet the preset performance indicators, solidify the initial reinforcement learning consensus reaching process RL-CRP model into a trained reinforcement learning consensus reaching process RL-CRP model. If the evaluation results do not meet the requirements, the reward function weight, learning rate or training rounds can be adjusted to retrain the reinforcement learning consensus reaching process RL-CRP model.

[0112] Compared with the prior art, the present invention has the following beneficial effects:

[0113] 1. This invention first strengthens the collaborative optimization of the marginal value function and the multilayer perceptron in the neural network multivariate decision-making auxiliary NN-MCDA model, and generates multiple initial expert evaluation probabilities by the neural network multivariate decision-making auxiliary NN-MCDA model. Then, a dual-agent dynamic adjustment mechanism is introduced in the reinforcement learning consensus reaching process (RL-CRP) model, and the dual agents are bridged through the state space. The dual agents of the reinforcement learning consensus reaching process (RL-CRP) model iteratively optimize the multiple initial expert evaluation probabilities generated by the neural network multivariate decision-making auxiliary NN-MCDA model, and generate multiple final expert evaluation probability matrices for bearing state decision-making, thereby achieving high-precision, low-round group fault diagnosis.

[0114] 2. This invention combines data-driven approaches with physical mechanisms by synergizing the neural network multivariate decision-making assistance (NN-MCDA) model with the reinforcement learning consensus achievement process (RL-CRP) model, thereby improving the interpretability of bearing fault diagnosis.

[0115] 3. This invention utilizes reinforcement learning to adaptively optimize group decision-making, effectively solving the problem of disagreement among multiple experts and significantly improving consensus efficiency. At the same time, by modeling the consensus process as a learnable environment and utilizing the mechanism of automatic optimization of the consensus path through reinforcement learning, efficient and stable diagnostic consensus can be achieved in complex working conditions and dynamic environments.

[0116] 4. The bearing fault diagnosis method provided by this invention, which combines the neural network multivariate decision-making auxiliary NN-MCDA model with the reinforcement learning consensus achievement process RL-CRP model, is applicable to complex working conditions such as wind power generation and rail transit, and is conducive to achieving accurate early fault warning and low-cost predictive maintenance.

[0117] 5. This invention provides a multivariate decision-making assistance model that integrates physical marginal value functions and neural network attention mechanisms, and combines a consensus-reaching process driven by reinforcement learning to achieve adaptive optimization of multi-expert fault diagnosis results. This not only solves the problems of existing bearing fault diagnosis methods being unable to cope with complex working conditions, insufficient physical interpretability, reliance on a single decision mode, and the use of static weight allocation, but also improves the accuracy, robustness, and interpretability of bearing fault diagnosis. Attached Figure Description

[0118] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0119] Figure 1 This is a flowchart of a bearing fault diagnosis method based on neural networks and multi-criteria preference consensus in an embodiment of the present invention. Detailed Implementation

[0120] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0121] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. The terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, unless otherwise explicitly specified and limited, the terms "installed," "connected," and "linked" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0122] like Figure 1 As shown, this invention provides a bearing fault diagnosis method based on neural networks and multi-criteria preference consensus, comprising the following steps:

[0123] S1. Data Acquisition and Feature Extraction: Obtain the model information and vibration signal of the target bearing, and extract the multi-domain features of the target bearing from the vibration signal.

[0124] In this embodiment, a triaxial accelerometer is fixed to the radial position of the bearing housing, and the sampling rate is set to 2560Hz to obtain the vibration signal of the bearing.

[0125] In this embodiment, the triaxial accelerometer is model PCB352C33.

[0126] In this application, the multi-domain characteristics of the target bearing include three types: time domain, frequency domain, and time-frequency domain.

[0127] It should be noted that time-domain, frequency-domain, and time-frequency-domain analysis are the three major methodologies of signal processing, each with its own unique advantages and limitations. Time-domain analysis directly focuses on the change of signal amplitude over time, reflecting the strength and periodicity of fault impacts. It is characterized by its simple calculations, high real-time performance, and sensitivity to impact-related faults (such as bearing pitting), but it is susceptible to noise interference and cannot capture frequency information. Frequency-domain analysis reveals the energy distribution of frequency components through Fourier transforms, excelling at identifying fault characteristic frequencies (such as inner or outer race fault frequencies) and harmonics, and exhibiting strong noise resistance. However, it relies on the assumption of signal stationarity and completely loses temporal dynamics. Time-frequency-domain analysis combines the time and frequency dimensions, making it suitable for non-stationary, time-varying signals (such as vibrations under varying operating conditions). It can pinpoint the time of fault occurrence and track its evolution, but its computational complexity is high. The core difference between the three is that the time domain emphasizes "waveform statistics," the frequency domain emphasizes "spectral energy," and the time-frequency-domain emphasizes "dynamic correlation." Combining these three approaches forms a complementary analytical system, jointly supporting high-precision fault diagnosis. Therefore, this application uses the time domain, frequency domain, and time-frequency domain characteristic data of the target bearing as a basis to determine the state of the target bearing.

[0128] In this embodiment, the time-domain features are obtained directly from the original waveform, the frequency-domain features are obtained through Fourier transform, and the time-frequency-domain features are obtained through wavelet packet decomposition and Empirical Mode Decomposition (EMD) algorithms.

[0129] In this embodiment, the time-domain characteristics of the target bearing include kurtosis, peak value, root mean square value, variance, waveform factor, impulse index, margin index, peak factor, skewness, amplitude factor, zero-crossing rate, and energy operator; the frequency-domain characteristics of the target bearing include frequency amplitude, centroid frequency, frequency variance, spectral kurtosis, harmonic component energy, sideband energy ratio, spectral peak density, frequency band energy, spectral entropy, and spectral standard deviation; the time-frequency domain characteristics of the target bearing include wavelet packet energy, IMF energy, time-frequency matrix singular values, instantaneous frequency variance, marginal spectral energy, wavelet entropy, Hilbert spectral entropy, and time-frequency clustering.

[0130] S2. Data Preprocessing: Standardize the multi-domain feature data of the target bearing and output the standardized feature vector of the target bearing. ,in For the i-th eigenvalue, For feature dimensions.

[0131] In this embodiment, Z-score is used to standardize the multi-domain feature data of the target bearing to eliminate the influence of dimensions.

[0132] S3. Matching Fault Diagnosis Models: Based on the model information of the target bearing, select bearing condition diagnosis models that match the target bearing model from the bearing condition diagnosis model database. The bearing condition diagnosis models include the trained neural network multivariate decision aid NN-MCDA model and the trained reinforcement learning consensus achievement process RL-CRP model.

[0133] In this application, the bearing condition diagnosis model database is pre-configured with bearing condition diagnosis models adapted to various bearing models. Each bearing condition diagnosis model adapted to a different bearing model includes a trained neural network multivariate decision aid NN-MCDA model and a trained reinforcement learning consensus achievement process RL-CRP model.

[0134] S4. Preliminary diagnosis and probability generation: Input the standardized feature vector of the target bearing into the trained neural network multivariate decision-making auxiliary NN-MCDA model that matches the target bearing model, and output multiple expert initial evaluation probabilities of the target bearing under different states. The trained neural network multivariate decision-making auxiliary NN-MCDA model is a hybrid model composed of neural network and multivariate decision analysis.

[0135] In this application embodiment, the bearing status includes four categories: normal, inner ring fault, outer ring fault, and rolling element fault. It should be noted that the states described in this application include, but are not limited to, these four status categories.

[0136] In this application, the trained neural network multivariate decision-making auxiliary NN-MCDA model is obtained by training the neural network multivariate decision-making auxiliary NN-MCDA model, wherein the neural network multivariate decision-making auxiliary NN-MCDA model includes several parallel neural network multivariate decision-making auxiliary NN-MCDA sub-models; each neural network multivariate decision-making auxiliary NN-MCDA sub-model includes an input layer, a marginal value mapping layer, a multilayer perceptron module, a utility fusion layer, and an output layer, specifically as follows:

[0137] The input layer takes the standardized feature vector of the target bearing as input.

[0138] The marginal value mapping layer receives the standardized feature vector of the target bearing and transforms each feature value in the standardized feature vector into a contribution to each state through the marginal value function, obtaining the linear contribution of each feature to each state. The marginal value function is shown in the following equation:

[0139]

[0140] In the formula, m is the state index symbol, and d is the polynomial degree index symbol. Let be the linear contribution of the i-th feature of the k-th sample to the m-th state. Let be the polynomial degree of the i-th feature with respect to the m-th state. The fitting coefficients of the d-th term for the i-th feature to the m-th state are... It is the i-th feature of the k-th sample;

[0141] The multilayer perceptron module receives the standardized feature vector of the target bearing, captures the high-order interactions and nonlinear relationships between features, and outputs the nonlinear contribution vector of the target bearing to each state.

[0142] The utility fusion layer receives the outputs of the marginal value mapping layer and the multilayer perceptron module. It integrates these outputs using a comprehensive utility function to obtain the comprehensive utility value of the target bearing for each state. The comprehensive utility function is shown in the following equation:

[0143]

[0144] In the formula, Let be the combined utility value of the k-th sample with respect to the m-th state; Let be the tradeoff coefficient for the m-th state. ; For feature dimension, Let i be the feature weight of the i-th feature in the m-th state. This represents the nonlinear contribution of the k-th sample to the m-th state.

[0145] The output layer receives the comprehensive utility value of the target bearing for each state and converts it into the initial evaluation probability of the target bearing for each state using the Softmax function. The conversion formula is shown below:

[0146]

[0147] In the formula, Let M be the initial evaluation probability of the k-th sample in the m-th state, and M be the number of states.

[0148] It should be noted that each vibration signal represents a sample. In actual bearing fault diagnosis, only one vibration signal from the target bearing is input. Therefore, the trained neural network multivariate decision-making auxiliary NN-MCDA model has only one sample for preliminary diagnosis and probability generation. However, when training the neural network multivariate decision-making auxiliary NN-MCDA model, multiple samples are input for training.

[0149] In this application, the multilayer perceptron module includes an input layer, a hidden layer, and an output layer. A CBAM submodule is embedded between the hidden layer and the output layer. The CBAM submodule consists of channel attention units and spatial attention units, specifically:

[0150] The input layer receives the standardized feature vector of the target bearing;

[0151] The hidden layer performs a linear transformation using the weight matrix and bias vector, projecting the input of the input layer into a high-dimensional space. Then, the ReLU activation function is applied to introduce non-linearity, capturing feature interactions and generating preliminary high-dimensional features. The linear transformation is shown in the following equation:

[0152]

[0153] The nonlinear activation is shown in the following equation:

[0154]

[0155] In the formula, This represents the combined value of the k-th sample in the hidden layer. Let be the weight matrix of the i-th feature in the hidden layer. The bias vector of the hidden layer. This is the output after activating the k-th sample;

[0156] The CBAM submodule reshapes the initial high-dimensional features output from the hidden layer into feature maps and inputs these feature maps into the channel attention unit. The channel attention unit weights the feature maps along the channel dimension based on the global average pooling and global max pooling results. The channel-weighted feature maps are then input into the spatial attention unit, which further enhances the channel-weighted feature maps based on the average pooling and max pooling results at each spatial location. Finally, the initial high-dimensional features are multiplied by the channel attention weights to obtain the enhanced channel features. These enhanced channel features are then multiplied by the spatial attention weights to output the weighted global high-dimensional features, as shown in the expression below:

[0157]

[0158]

[0159]

[0160]

[0161] In the formula, F is the input feature map. This indicates that global average pooling is performed on each channel of F. This indicates that global max pooling is performed on each channel of F, and MLP is a shared multilayer perceptron. It is the Sigmoid activation function. For channel attention weights, Spatial attention weights, Indicates to and Perform the splicing operation; Indicates the kernel size as Convolution operations are used to capture local relationships in spatial locations; To enhance channel features, The weighted global high-dimensional features;

[0162] The output layer receives the weighted global high-dimensional features output by the CBAM submodule, configures independent weight parameters for each state to be diagnosed, and calculates a feature pattern matching weight vector for each state through a linear transformation between the weight vector and the weighted global high-dimensional features. This achieves the correlation measurement between the target bearing and each state, and outputs the nonlinear contribution vector of the target bearing to each state, i.e., [ The formula for calculating the linear transformation between the weight vector and the weighted global high-dimensional features is as follows:

[0163]

[0164] In the formula, Let be the nonlinear contribution of the k-th sample to the m-th state. Let m be the weight matrix for the m-th state. The weighted global high-dimensional feature of the k-th sample. This is the bias vector.

[0165] It's important to note that the weighted global high-dimensional feature output by the CBAM submodule is a comprehensive, global feature representation optimized by the attention mechanism. It is not segmented or copied into multiple feature representations for different states, but rather shared by all states. In other words, the weighted global high-dimensional feature output by the CBAM submodule has the same feature representation for each state. For example, if the weighted global high-dimensional feature output by the CBAM submodule has a feature representation of 1.5 for the normal state, then its feature representations for inner race faults, outer race faults, and rolling element faults will all be 1.5.

[0166] In this application, the linear transformation process of the output layer in the multilayer perceptron module essentially constructs a mapping relationship between the weighted global high-dimensional features and each state. The resulting output represents the nonlinear contribution of the current sample to each state, and its magnitude is positively correlated with the feature-state matching degree. Furthermore, in the multilayer perceptron module, the nonlinear contribution vector of the target bearing to each state output by the output layer, i.e., [ This fully characterizes the matching degree distribution of the current sample with all states, and can provide a solid theoretical basis for the RL-CRP model of utility fusion layer and post-training reinforcement learning consensus achievement process.

[0167] In this embodiment, the number of hidden layer units in the multilayer perceptron module is 64.

[0168] In this application, the CBAM submodule first utilizes the channel attention unit to deeply analyze and accurately identify key feature types, laying a solid foundation for subsequent analysis. Subsequently, relying on the spatial attention unit, it precisely locates the key distribution areas of features, thereby achieving dynamic enhancement of feature representation, making feature information richer and more targeted. Finally, by embedding the CBAM submodule into the multilayer perceptron module of the neural network multivariate decision-aided NN-MCDA model, this application can focus on physical features closely related to the fault with extremely high accuracy during bearing fault diagnosis, such as the impact component in vibration signals. This innovative fusion method not only significantly improves the robustness of bearing fault diagnosis, enabling it to maintain stable and reliable diagnostic performance even under complex and changing actual working conditions, but also greatly enhances the interpretability of diagnostic results, providing strong evidence for in-depth analysis and accurate judgment of fault causes.

[0169] In this application, the channel attention unit includes:

[0170] The global average pooling subunit calculates the average value for each channel of the feature map to obtain the global characteristics of that channel.

[0171] The global max pooling subunit maximizes each channel of the feature map to capture the response at significant activation locations;

[0172] The shared multilayer perceptron subunit receives the outputs of the global average pooling subunit and the global max pooling subunit, and learns the relationship between pooling features.

[0173] The element-sum subunit adds the outputs of the global average pooling subunit and the global max pooling subunit in the shared multilayer perceptron subunit;

[0174] The output sub-units are then processed using the Sigmoid activation function to compress the summation results of the element-sum sub-units to the range of (0, 1), thus obtaining the channel attention weights.

[0175] In this embodiment, the shared multilayer sensing subunit consists of two sensing layers: one for compressing the channel and the other for restoring the channel.

[0176] In this application, the channel attention unit adopts a dual-path information fusion strategy, which organically combines the global context information captured by the average pooling operation with the salient feature information extracted by the max pooling operation. Through this dual-path collaborative approach, the channel attention unit can dynamically learn the importance weights of different feature channels, greatly improving the ability and sensitivity of the neural network multivariate decision-making aid NN-MCDA sub-model to capture key signals. This enables the neural network multivariate decision-making aid NN-MCDA sub-model to quickly focus on the most diagnostically valuable information when faced with complex signals generated by bearing faults, laying a solid foundation for accurate and efficient diagnosis of bearing faults.

[0177] In this application, the spatial attention unit includes:

[0178] The global average pooling sub-unit calculates the mean of all channels at each location in space;

[0179] The global max-pooling sub-unit calculates the maximum value of all channels at each location in the space.

[0180] The splicing sub-unit concatenates the outputs of the global average pooling sub-unit and the global max pooling sub-unit to obtain a two-channel feature map.

[0181] The convolutional computation subunit uses a convolutional kernel to convolve the two-channel feature maps to obtain a spatial weight map.

[0182] The output sub-units are compressed into the range (0, 1) using the Sigmoid activation function to obtain the spatial attention weights.

[0183] In this application, each neural network multivariate decision-making auxiliary NN-MCDA model includes multiple parallel neural network multivariate decision-making auxiliary NN-MCDA sub-models. Each neural network multivariate decision-making auxiliary NN-MCDA sub-model corresponds to a virtual expert. Their marginal value function coefficients, feature weights, and neural network parameters are not shared with each other. They are trained independently using different parameter initialization, training samples, or feature selection strategies, thereby forming different contribution assessments for the same vibration feature, realizing the diversity of multi-expert diagnostic opinions, and providing differentiated inputs for the reinforcement learning consensus-building process RL-CRP model after training.

[0184] Furthermore, during the training process, different loss weights and / or regularization coefficients can be set for different states in several sub-models of each neural network multivariate decision-making auxiliary NN-MCDA model. This allows each sub-model to form different class preferences and complexity constraints when optimizing the marginal value function coefficients and feature weights, thereby further improving the diversity of expert evaluation results and the effectiveness of consensus optimization.

[0185] In this embodiment, the trained neural network multivariate decision-making auxiliary NN-MCDA model for each bearing model is composed of multiple parallel neural network multivariate decision-making auxiliary NN-MCDA sub-models. To enable each neural network multivariate decision-making auxiliary NN-MCDA sub-model to form differentiated diagnostic characteristics, each neural network multivariate decision-making auxiliary NN-MCDA sub-model uses the same network structure and training process for independent parameter optimization, but they differ in training sample sampling strategy, feature constraints, and loss function parameters. The specific training steps for each trained neural network multivariate decision-making auxiliary NN-MCDA model are as follows:

[0186] S41. Obtain multiple sets of historical bearing vibration samples of this model under different conditions, and extract the multi-domain features of each set of historical bearing vibration samples under different conditions, wherein the number of historical bearing vibration samples under each condition is equal.

[0187] S42. Standardize the multi-domain features of each group of historical bearing vibration samples under different states to obtain the standardized feature vector of each group of historical bearing vibration samples under different states.

[0188] S43. Divide multiple sets of historical bearing vibration samples under the same state into a training set and a validation set according to the proportion. Each set of historical bearing vibration samples in the training set / validation set carries its corresponding standardized feature vector.

[0189] S44. Initialize multiple neural network multivariate decision-making auxiliary NN-MCDA sub-models, and set differentiated training strategies for each neural network multivariate decision-making auxiliary NN-MCDA sub-model from three levels. Specifically: First, set differentiated sample sampling weights or resampling probabilities to weight or resample the samples in the training set, thereby prompting each neural network multivariate decision-making auxiliary NN-MCDA sub-model to focus on different data distributions; Second, set feature subset masks for each neural network multivariate decision-making auxiliary NN-MCDA sub-model to constrain each neural network multivariate decision-making auxiliary NN-MCDA sub-model to focus on specific feature types; Finally, set differentiated class weights in the loss function to adjust the misclassification penalty for samples in different states, thereby enhancing the recognition ability of each neural network multivariate decision-making auxiliary NN-MCDA sub-model on the preset emphasis state.

[0190] S45. Input all the training sets corresponding to each state into each neural network multivariate decision aid NN-MCDA sub-model for independent training. During the training process, the samples in the training set are weighted and / or resampled according to the sample sampling weights set by the neural network multivariate decision aid NN-MCDA sub-model. This enables each neural network multivariate decision aid NN-MCDA sub-model to learn the complete mapping relationship from the input standardized feature vector to all states. After training, each neural network multivariate decision aid NN-MCDA sub-model becomes an "expert" that can work independently and output an evaluation probability vector covering all states for each input sample.

[0191] S46. Use the validation set corresponding to each state to evaluate each trained neural network multivariate decision-making auxiliary NN-MCDA sub-model. Based on the evaluation results, select the parameters with the best overall performance on the validation set for each neural network multivariate decision-making auxiliary NN-MCDA sub-model, and solidify it into a trained neural network multivariate decision-making auxiliary NN-MCDA sub-model with specific expertise. During the evaluation, comprehensively consider the recall rate of the neural network multivariate decision-making auxiliary NN-MCDA sub-model in its emphasized state, as well as the false alarm rate and / or specificity in its non-emphasized state, to ensure the effectiveness and reliability of its expertise.

[0192] S47. Integrate multiple trained neural network multivariate decision-making auxiliary NN-MCDA sub-models into a complete trained neural network multivariate decision-making auxiliary NN-MCDA model.

[0193] It should be noted that step S45, "learning the mapping relationship from the input standardized feature vector to all states," is a data-driven parameter optimization mechanism executed automatically by the machine. The essence of this mechanism is to iteratively adjust millions of parameters within the NN-MCDA sub-model using the gradient descent algorithm until it can automatically fit the complex nonlinear relationship between input features and output states. Specifically, this automated process involves a loop of three key steps: First, the NN-MCDA sub-model performs forward propagation on the input standardized feature vector to generate an initial state probability prediction; then, it quantifies the error between this prediction and the true label using a loss function (such as cross-entropy loss); finally, it calculates the error gradient based on the backpropagation algorithm, and the optimizer fine-tunes all weights and biases of the NN-MCDA sub-model according to the gradient descent rule. By repeatedly performing this "prediction-evaluation-correction" iteration across the entire training set, the parameters of the NN-MCDA sub-model are continuously optimized, ultimately internalizing the accurate mapping relationship into its parameters.

[0194] S5. Final Diagnosis and Probability Generation: The initial evaluation probabilities of multiple experts for the target bearing under different states are input into the training reinforcement learning consensus reaching process RL-CRP model that matches the target bearing model. The output is the final evaluation probability matrix of multiple experts for the target bearing under different states and the weights of each expert in the last iteration. The training reinforcement learning consensus reaching process RL-CRP model is a hybrid model composed of reinforcement learning and consensus reaching process.

[0195] In this application, the Post-Training Reinforcement Learning Consensus Reaching Process (RL-CRP) Model is obtained by training the Reinforcement Learning Consensus Reaching Process (RL-CRP) Model, which includes a four-layer structure, specifically:

[0196] The state construction layer constructs an environmental state vector based on the initial evaluation probabilities of multiple experts under different states of the target bearing. The environmental state vector includes the consensus index of each expert, the group harmony degree, and the iteration round.

[0197] The decision-making strategy layer inputs the environmental state vector into the reinforcement learning architecture for iterative optimization, generating multiple expert evaluation probability matrices for the target bearing under different states. The reinforcement learning architecture includes two independent intelligent agents, namely... Agent and Agent, among which The agent is used to generate feedback parameters to control the adjustment range of expert opinions, so that the evaluation results of each expert converge to the group average. The agent is used to generate expert weight vectors and normalization is ensured through Softmax activation.

[0198] The reward evaluation layer calculates the reward function for each intelligent agent based on group harmony, generation round, and changes in expert weights, and updates the policy networks of the two intelligent agents respectively. The agent's reward function focuses on three main objectives: encouraging high harmony, low iteration rounds, and smooth consensus changes. The agent's reward function encourages high-consensus experts to have higher weights, while constraining the weight changes to be smooth.

[0199] Iterative optimization layer, when all expert consensus indices are higher than the consensus threshold The optimization may terminate when the number of iterations exceeds the upper limit, and the output will be the probability matrix of the final evaluation of the target bearing by multiple experts under different states, as well as the weight of each expert in the last iteration.

[0200] In this application, the specific steps for constructing an environmental state vector based on multiple initial expert evaluation probabilities of the target bearing under different states are as follows:

[0201] S511. Treat each vibration signal of the target bearing as a sample, and transform the multiple initial expert evaluation probabilities of each sample under different states into a decision matrix. The decision matrix for each sample is shown below:

[0202]

[0203] In the formula, Let be the decision matrix for the k-th sample. Let be the probability of the nth expert's evaluation of the mth state.

[0204] It should be noted that the probability matrices of multiple expert evaluations of the target bearing under different states and the probability matrices of multiple final expert evaluations of the target bearing under different states in this application are expressed in the same form as the decision matrix in step S511.

[0205] It should be noted that in actual bearing fault diagnosis, only one vibration signal from the target bearing is input. Therefore, when constructing the environmental state vector based on multiple initial expert evaluation probabilities of the target bearing under different states in the trained reinforcement learning consensus-reaching process (RL-CRP) model, only one sample is used, resulting in only one decision matrix. However, when training the RL-CRP model, multiple samples are input for training.

[0206] S512. Based on the decision matrix of all samples, construct an environmental state vector that includes consensus index, harmony degree, and iteration round. ,in Let be the relational consensus index of the nth expert on all samples in the t-th iteration. Let be the harmony degree at the t-th iteration.

[0207] In this application, the consensus index includes three levels: element-level, scheme-level, and relation-level. The consensus index in the environment state vector is the relation-level consensus index.

[0208] In this application, the element-level consensus index is used to calculate the consistency of a single evaluation point, and the calculation formula is as follows:

[0209]

[0210]

[0211] In the formula, Let be the element-level consensus index of the nth expert on the mth state of the kth sample. Let be the probability of the nth expert's evaluation of the mth state of the kth sample. Let R be the group-weighted average evaluation probability of all experts for the m-th state of the k-th sample, R be the set of evaluation probabilities of all experts for all states of all samples, and N be the number of experts. For the normalized weight of the nth expert, and .

[0212] In this application, the scheme-level consensus index is used to aggregate the consensus of an expert on all states of a certain sample, and the calculation formula is as follows:

[0213]

[0214] In the formula, Let M be the scheme-level consensus index of the nth expert on all states of the kth sample, where M is the number of states.

[0215] In this application, the relational consensus index is used to aggregate an expert's consensus on all samples to obtain their overall individual consensus score. The calculation formula is as follows:

[0216]

[0217] In the formula, Let K be the consensus index of the nth expert on the relationship level of all samples, and K be the number of samples.

[0218] In this application, harmony is used to quantify the overall consistency of the group. The formula for calculating harmony in each iteration is as follows:

[0219]

[0220] In the formula, HD represents the harmony degree, and t is the iteration round. Let be the probability of the nth expert's evaluation of the mth state of the kth sample in the tth iteration.

[0221] In this application, the individual relational consensus index can quantify the degree of deviation between each expert's current assessment and the group opinion. ∈[0,1], the lower the value, the more the expert's opinion deviates from the consensus; the harmony degree characterizes the overall consistency level of the group, calculated by aggregating the CI values ​​of individuals, and t records the iteration rounds to control process convergence. This application encodes the group state into structured information through a multi-granularity consensus perception mechanism, enabling the reinforcement learning agent to accurately identify sources of disagreement (such as abnormally low CI values ​​of specific experts), assess the progress of global consensus (through the harmony degree HD value), and dynamically adjust optimization strategies (such as imposing time penalties based on t). This state representation method, while ensuring information completeness, conforms to the Markov property required by Markov Decision Processes (MDPs), and can provide a solid theoretical basis for subsequent intelligent decision-making based on policy gradients.

[0222] In this application, at t=0, the environmental state vector middle , Obtained based on the initial probability matrix.

[0223] In this application, the action spaces of the two intelligent agents are both continuous spaces. Physical constraints are implemented through an Actor network combined with a specific activation function to ensure that the actions conform to the actual meaning.

[0224] In this application, The agent's action space is defined as Its value is generated through the Actor network, as shown below:

[0225]

[0226] In the formula, In the t-th iteration The actions of the agent In the t-th iteration Feedback parameters generated by the agent for Agent's strategy network, The environment state vector is the input policy network. for The set of learnable parameters of the agent in the policy network; for The output layer weight matrix of the proxy, where ReLU is the modified linear unit activation function. for The hidden layer weight matrix of the proxy, Indicates will The state feature mapping function is converted into a feature vector that can be processed by the network. for The hidden layer bias vector of the proxy. for The output layer bias vector of the proxy uses the Sigmoid activation function to ensure... .

[0227] In this application, The agent's reward function is formed by a weighted combination of harmony, iteration rounds, and changes in the consensus index, as shown in the following formula:

[0228]

[0229] In the formula, In the t-th iteration The agent's reward function, , , These are the weighting coefficients.

[0230] In the embodiments of this application, The agent's reward function focuses on three main objectives: encouraging high harmony, low iteration rounds, and smooth consensus changes. Among these, encouraging high harmony is... The core objective of the agent is to encourage the agent to continuously optimize its behavior during interactions by directly rewarding improvements in consensus levels; penalizing excessively long iteration cycles is... The agent's second objective is to incentivize agents to actively explore and learn more efficient strategies by penalizing excessively long iteration cycles, thereby accelerating the consensus process, ensuring a consistent decision is reached within the shortest possible number of cycles, and improving overall decision-making efficiency. Penalizing drastic fluctuations in the consensus index is... The third objective of the agent is to encourage agents to follow a smooth consensus change path by penalizing drastic fluctuations in the consensus index, thereby preventing agents from adopting aggressive strategies such as "slamming on the gas and slamming on the brakes," thus maintaining the stability and consistency of the decision-making process and ensuring that the final consensus has high reliability and stability.

[0231] It should be noted that in this application The core function of the proxy is to act as an adaptive regulator of consensus convergence, with the core objective being to dynamically output feedback parameters. This is used to quantify the magnitude by which experts whose opinions deviate from the group mean converge towards the group opinion in each round of consensus iteration. It needs to achieve an optimal balance between two conflicting objectives: if... While an excessively large consensus size can force experts to quickly converge on the group's opinion, thus improving convergence efficiency, it can easily overlook the original opinions of highly reliable experts, leading to distorted consensus. While a small threshold can fully respect individual expert judgments and ensure the robustness of opinions, it can also lead to slow consensus convergence and even result in intractable local disagreements. Therefore, in this application... The ultimate goal of the proxy is to achieve a balance between efficient convergence of the consensus process and robustness of opinions.

[0232] In this application, The agent's action space is defined as Its value is generated through the Actor network, as shown below:

[0233]

[0234] In the formula, In the t-th iteration The actions of an agent; In the t-th iteration The expert weight vector generated by the agent, In the t-th iteration The normalized weights of the nth expert generated by the agent; for Agent's strategy network, The environment state vector is the input policy network. for The set of learnable parameters of the agent in the policy network; for The output layer weight matrix of the proxy, where ReLU is the modified linear unit activation function. for The hidden layer weight matrix of the proxy, Indicates will The state feature mapping function is converted into a feature vector that can be processed by the network. for The hidden layer bias vector of the proxy. for The proxy's output layer bias vector, using Activation function, ensure and .

[0235] In this application, The agent's reward function is formed by the weighted sum of expert weights and consensus index, minus the squared penalty for changes in weights, as shown in the following formula:

[0236]

[0237] In the formula, In the t-th iteration The agent's reward function, In the t-th iteration The normalized weights of the nth expert generated by the agent. and ; This is an evaluation item for the immediate effect of the strategy, used to encourage the allocation of high weights to high-consensus experts in order to directly improve the level of collective consensus. A coefficient used to control the intensity of punishment; As a constraint, the overall change range of the weight strategy is quantified through L2 norm, and drastic changes in the weight vector are penalized to ensure the continuity and robustness of weight allocation; The change in weight represents the weight of the first... Round iteration and the first Overall changes in weight allocation strategy during rounds of iteration ; It is a quantitative indicator of the "overall strategic change range".

[0238] It should be noted that in this application The actual effect of this term is to impose smoothness constraints and penalize drastic changes in the weight vector between adjacent rounds. When When the value is large, it indicates a sudden change in the expert weight allocation strategy, resulting in increased penalties; when... When the value is relatively small, it indicates that the weight change is stable and the penalty is reduced. In this application, through... It can ensure The weight allocation strategy learned by the agent is stable and evolves continuously, rather than oscillating violently between rounds.

[0239] It should be noted that in this application The core function of an agent is to act as a dynamic allocator of expert discourse power, with its core objective being to generate responses that satisfy constraints. and The expert weight vector is optimized based on the current consensus state, accurately identifying two types of experts and assigning them differentiated weights: one type consists of highly reliable experts with high accuracy and low opinion fluctuation in historical diagnostic tasks, and the other type consists of experts with high consensus whose opinions deviate little from the group opinion in the current iteration. By assigning higher weights to these two types of experts, the group opinion is guided to converge towards a more credible direction, thereby improving the accuracy of bearing fault diagnosis.

[0240] In this application, in each iteration, Agent and After obtaining the feedback parameters and expert weight vector based on the same environmental state vector, the agent first calculates the group-weighted average evaluation probability of all experts for the m-th state of the k-th sample in the current iteration based on the expert weight vector. Then, the consensus index of each expert in the current iteration is calculated, and for those experts whose consensus index is ≤ consensus threshold, the consensus index is then calculated. Experts use feedback parameters to update the evaluation probability in the next iteration, and finally calculate the environment state vector in the next iteration based on the updated evaluation probability.

[0241] In this application, in each iteration, for consensus exponents ≤ consensus thresholds... The experts update their evaluation probabilities for the next iteration using the following formula:

[0242]

[0243] In the formula, Let be the probability of the nth expert's evaluation of the mth state of the kth sample in the t-th iteration. Let be the group-weighted average evaluation probability of all experts for the m-th state of the k-th sample in the t-th iteration. In the t-th iteration Feedback parameters generated by the agent.

[0244] In this application, the reinforcement learning consensus-reaching process (RL-CRP) model outputs not only the probability matrix of the final expert evaluation of the target bearing under different states and the weights of each expert in the last iteration, but also a set of key process evaluation parameters, including the final consensus index, harmony degree, total number of iterations, and termination reason. These parameters together constitute a quantitative evaluation system for the consensus-reaching process, which can provide important basis for verifying the credibility of diagnostic results, analyzing decision efficiency, and optimizing system debugging.

[0245] In this application, the specific training steps of the RL-CRP model for the post-training reinforcement learning consensus reaching process corresponding to each type of bearing are as follows:

[0246] S521. Obtain multiple sets of historical bearing vibration samples of this model under different conditions, and extract multi-domain features of each set of historical bearing vibration samples, wherein the number of historical bearing vibration samples under each condition is equal.

[0247] S522. Standardize the multi-domain features of each group of historical bearing vibration samples to obtain the standardized feature vector of each group of historical bearing vibration samples.

[0248] S523. Input the standardized feature vectors of all historical bearing vibration samples into the trained neural network multivariate decision-making auxiliary NN-MCDA model. For each group of historical bearing vibration samples, obtain multiple initial expert evaluation probabilities under different states.

[0249] S524. Divide all historical bearing vibration samples into a training set and a validation set according to the proportion. The number of historical bearing vibration samples in each state is equal in the training set and the validation set. Each group of historical bearing vibration samples in the training set and the validation set carries multiple initial expert evaluation probabilities for each state.

[0250] For example, in this embodiment of the application, the historical vibration sample data of a certain type of bearing under the four states of normal, inner ring failure, outer ring failure and rolling element failure are all 50. If the training set and the validation set are divided in an 8:2 ratio in this embodiment of the application, then the training set in step S524 will have 40 historical vibration samples under the four states of normal, inner ring failure, outer ring failure and rolling element failure.

[0251] S525. Input the initial expert evaluation probabilities of each group of historical bearing vibration samples in different states into the reinforcement learning consensus reaching process (RL-CRP) model for training, and obtain the initial reinforcement learning consensus reaching process (RL-CRP) model.

[0252] In this embodiment of the application, step S525 specifically involves: taking the initial expert evaluation probabilities of each group of historical bearing vibration samples in the training set under different states as the initial environment states, and inputting them into the reinforcement learning consensus-reaching process (RL-CRP) model. Agency and The agent repeatedly interacts with the environment in the state space, generating feedback parameters and expert weight vectors, iteratively updating the expert evaluation matrix, and calculating a multi-objective reward function based on consensus index, harmony degree, and iteration rounds; the Deep Deterministic Policy Gradient (DDPG) algorithm is used to... Agent and The agent's policy network and value network parameters are updated until the consensus index and convergence efficiency on the training set meet the preset requirements, thus obtaining the initial reinforcement learning consensus achievement process RL-CRP model.

[0253] In this embodiment of the application, the Deep Deterministic Strategy Gradient (DDPG) algorithm is used to... Agent and When updating the agent's policy network and value network parameters, the hyperparameters set include the Actor learning rate, Critic learning rate, discount factor, soft update coefficient, and batch sampling number.

[0254] It should be noted that the Deep Deterministic Policy Gradient (DDPG) algorithm used in this application includes Critic network evaluation, Actor network policy optimization, experience replay, and target network soft update. The Critic network acts as a value function approximator, aiming to evaluate the long-term expected return of the "state-action" relationship by updating parameters by minimizing temporal difference errors. The Actor network acts as a policy function, aiming to find the action policy that maximizes long-term returns. Based on the deterministic policy gradient theorem, it utilizes the value gradient provided by its online Critic network to guide the update direction. Experience replay involves storing an experience tuple (current state, agent action, reward, and next state) in an experience replay buffer after each execution of the state-action-transition-reward interaction. Small batches of experience are randomly sampled from the buffer to update the Actor and Critic networks, breaking sample correlation and improving training stability. Target network soft update involves fusing the parameters of the online network into the parameters of the target network at a certain ratio, allowing the target network's parameters to slowly track the online network's parameters.

[0255] In the embodiments of this application, Agent and Each agent has its own independent network, experience buffer, and loss calculation logic, and each agent performs its own parameter optimization synchronously.

[0256] It should be noted that both intelligent agents in this application are trained using the Deep Deterministic Policy Gradient (DDPG) algorithm. The core reason for this adaptability is that the action space for consensus optimization is a continuous space, and the DDPG algorithm, through its Actor-Critic architecture and experience replay mechanism, can efficiently handle offline reinforcement learning tasks with continuous actions.

[0257] S526. Input the initial expert evaluation probabilities of each group of historical bearing vibration samples in different states in the validation set into the initial reinforcement learning consensus reaching process RL-CRP model. Evaluate the consensus degree, classification accuracy and convergence rounds of the initial reinforcement learning consensus reaching process RL-CRP model on the validation set. When the evaluation results meet the preset performance indicators, solidify the initial reinforcement learning consensus reaching process RL-CRP model into a trained reinforcement learning consensus reaching process RL-CRP model. If the evaluation results do not meet the requirements, the reward function weight, learning rate or training rounds can be adjusted to retrain the reinforcement learning consensus reaching process RL-CRP model.

[0258] S6. Bearing State Decision: Based on the probability matrix of multiple experts' final evaluation of the target bearing under different states and the weights of each expert in the last iteration, generate the collective probability of the target bearing under different states, and select the state with the highest collective probability as the final state of the target bearing.

[0259] Preferably, when the collective probability of two or more states is equal and all are at their highest values, the system will automatically trigger the "composite fault suspected state" judgment mechanism: First, output the state uncertainty indicator and activate the alarm, while generating a detailed diagnostic report; Second, the system will perform subsequent operations according to the preset strategy prompts, specifically: In scenarios that support automatic precision detection, automatically increase the sampling rate for high-frequency reanalysis; In scenarios that rely on manual decision-making, push alarm information and detailed data to the expert interface to request manual review, thereby forming a complete and configurable boundary case handling closed loop.

[0260] In this application, based on the probability matrix of the final evaluation by multiple experts for the target bearing under different states and the weights of each expert in the last iteration, the expression for the collective probability of the target bearing under different states is generated as follows:

[0261]

[0262] In the formula, Let be the collective probability of all experts for the m-th state of the k-th sample. The normalized weights of the nth expert in the final iteration of the RL-CRP model during the post-training reinforcement learning consensus-reaching process satisfy nonnegativity. With normality .

[0263] In this application, the collective probability corresponding to the final state of the target bearing will also be used as the confidence level output to assess the reliability and quality of group decision-making. When the confidence level is greater than or equal to a preset threshold, it indicates that the consensus of the expert group is high, the diagnostic conclusion is reliable, and it can be directly adopted; when the confidence level is less than the preset threshold, it indicates that there is a significant disagreement among experts, the diagnostic conclusion is highly uncertain, and professional personnel need to conduct manual review. In this way, intelligent self-checking and hierarchical processing of the reliability of diagnostic results can be achieved.

[0264] It is important to emphasize that the training samples in this embodiment are mainly derived from bearing vibration data under specific speed and load conditions. These samples are used to verify the performance of the proposed method in bearing fault diagnosis under these specific conditions, ensuring its effectiveness and reliability in practical applications. In practical applications, to improve the model's adaptability to different operating conditions, in addition to the basic information of bearing model, operating condition data such as speed, load level, and operating condition code can be added as basic information. Sample data from different operating conditions are used to train the neural network multivariate decision-making auxiliary (NN-MCDA) model and the reinforcement learning consensus-reaching process (RL-CRP) model. This allows the NN-MCDA and RL-CRP models to automatically learn the modulation effect of operating condition variables on the relationship between various vibration characteristics and states during training. Therefore, this application can also fine-tune and train the neural network multivariate decision-making assistance (NN-MCDA) model and the reinforcement learning consensus achievement process (RL-CRP) model for different speed ranges and load conditions based on a certain type of bearing. A bearing condition diagnosis model database is constructed using the bearing model and operating condition parameters as indexes. The corresponding bearing condition diagnosis model is selected from the bearing condition diagnosis model database according to the bearing model and operating condition for diagnosis, thereby realizing bearing fault diagnosis across operating conditions.

[0265] In summary, this invention provides a multivariate decision-making assistance model that integrates physical marginal value functions and neural network attention mechanisms, and combines a consensus-reaching process driven by reinforcement learning. This not only achieves adaptive optimization of multi-expert fault diagnosis results, but also improves the accuracy, robustness, and interpretability of bearing fault diagnosis.

[0266] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A bearing fault diagnosis method based on neural network and multi-criteria preference consensus, characterized in that, Includes the following steps: S1. Data Acquisition and Feature Extraction: Acquire the model information and vibration signal of the target bearing, and extract the multi-domain features of the target bearing from the vibration signal; S2. Data Preprocessing: Standardize the multi-domain feature data of the target bearing and output the standardized feature vector of the target bearing. ,in For the i-th eigenvalue, For feature dimensions; S3. Matching fault diagnosis model: Based on the model information of the target bearing, select the bearing condition diagnosis model that matches the target bearing model from the bearing condition diagnosis model database. The bearing condition diagnosis model includes the trained neural network multivariate decision aid NN-MCDA model and the trained reinforcement learning consensus achievement process RL-CRP model. S4. Preliminary Diagnosis and Probability Generation: The standardized feature vector of the target bearing is input into a trained neural network multivariate decision-making auxiliary NN-MCDA model that matches the target bearing model. The model outputs multiple initial expert evaluation probabilities for the target bearing under different states. The trained neural network multivariate decision-making auxiliary NN-MCDA model is a hybrid model composed of neural networks and multivariate decision analysis. This model is obtained by training the neural network multivariate decision-making auxiliary NN-MCDA model, which includes multiple parallel neural network multivariate decision-making auxiliary NN-MCDA sub-models. Each sub-model is an expert capable of independent operation. Each sub-model includes an input layer, a marginal value mapping layer, a multilayer perceptron module, a utility fusion layer, and an output layer, specifically: The input layer takes the standardized feature vector of the target bearing as input. The marginal value mapping layer receives the standardized feature vector of the target bearing and transforms each feature value in the standardized feature vector into a contribution to each state through the marginal value function, obtaining the linear contribution of each feature to each state. The marginal value function is shown in the following equation: , In the formula, m is the state index symbol, and d is the polynomial degree index symbol. Let be the linear contribution of the i-th feature of the k-th sample to the m-th state. Let be the polynomial degree of the i-th feature with respect to the m-th state. The fitting coefficients of the d-th term for the i-th feature to the m-th state are... It is the i-th feature of the k-th sample; The multilayer perceptron module receives the standardized feature vector of the target bearing, captures the high-order interactions and nonlinear relationships between features, and outputs the nonlinear contribution vector of the target bearing to each state. The utility fusion layer receives the outputs of the marginal value mapping layer and the multilayer perceptron module. It integrates these outputs using a comprehensive utility function to obtain the comprehensive utility value of the target bearing for each state. The comprehensive utility function is shown in the following equation: , In the formula, Let be the combined utility value of the k-th sample with respect to the m-th state; Let be the tradeoff coefficient for the m-th state. ; For feature dimension, Let i be the feature weight of the i-th feature in the m-th state. This represents the nonlinear contribution of the k-th sample to the m-th state. The output layer receives the comprehensive utility value of the target bearing for each state and converts it into the initial evaluation probability of the target bearing for each state using the Softmax function. The conversion formula is shown below: , In the formula, Let M be the initial evaluation probability of the k-th sample in the m-th state, where M is the number of states; S5. Final Diagnosis and Probability Generation: The initial expert evaluation probabilities of the target bearing under different states are input into a post-trained reinforcement learning consensus-reaching process (RL-CRP) model that matches the target bearing model. The output is a final expert evaluation probability matrix for the target bearing under different states, along with the expert weights in the final iteration. The RL-CRP model is a hybrid model composed of reinforcement learning and consensus-reaching processes. It is obtained by training the reinforcement learning consensus-reaching process (RL-CRP) model and includes a four-layer structure: The state construction layer constructs an environmental state vector based on the initial evaluation probabilities of multiple experts under different states of the target bearing. The environmental state vector includes the consensus index of each expert, the group harmony degree, and the iteration round. The decision-making strategy layer inputs the environmental state vector into the reinforcement learning architecture for iterative optimization, generating multiple expert evaluation probability matrices for the target bearing under different states. The reinforcement learning architecture includes two independent intelligent agents, namely... Agent and Agent, among which The agent is used to generate feedback parameters to control the adjustment range of expert opinions, so that the evaluation results of each expert converge to the group average. The agent is used to generate expert weight vectors and normalization is ensured through Softmax activation. The reward evaluation layer calculates the reward function for each intelligent agent based on group harmony, generation round, and changes in expert weights, and updates the policy networks of the two intelligent agents respectively. The agent's reward function focuses on three main objectives: encouraging high harmony, low iteration rounds, and smooth consensus changes. The agent's reward function encourages high-consensus experts to have high weights, while constraining the weight changes to be smooth. Iterative optimization layer, when all expert consensus indices are higher than the consensus threshold The optimization is terminated when the number of iterations exceeds the upper limit, and the output is the probability matrix of the final evaluation of the target bearing by multiple experts under different states and the weight of each expert in the last iteration. S6. Bearing State Decision: Based on the probability matrix of multiple experts' final evaluations of the target bearing under different states and the weights of each expert in the last iteration, the collective probability of the target bearing under different states is generated, and the state with the highest collective probability is selected as the final state of the target bearing. The expression for the collective probability of the target bearing under different states is as follows: , In the formula, Let N be the collective probability of all experts for the m-th state of the k-th sample, where N is the number of experts. The normalized weights of the nth expert in the final iteration of the RL-CRP model during the post-training reinforcement learning consensus-reaching process satisfy nonnegativity. With normality ; Let be the probability of the nth expert's evaluation of the mth state of the kth sample.

2. The bearing fault diagnosis method based on neural networks and multi-criteria preference consensus as described in claim 1, characterized in that, The multi-domain characteristics of the target bearing include three types: time domain, frequency domain, and time-frequency domain. The time domain characteristics include kurtosis, peak value, root mean square value, variance, waveform factor, impulse index, margin index, peak factor, skewness, amplitude factor, zero-crossing rate, and energy operator. The frequency domain characteristics include frequency amplitude, centroid frequency, frequency variance, spectral kurtosis, harmonic component energy, sideband energy ratio, spectral peak density, frequency band energy, spectral entropy, and spectral standard deviation. The time-frequency domain features include wavelet packet energy, IMF energy, singular values ​​of the time-frequency matrix, instantaneous frequency variance, marginal spectrum energy, wavelet entropy, Hilbert spectral entropy, and time-frequency clustering. In step S2, Z-score is used to standardize the multi-domain feature data of the target bearing. The bearing status includes four categories: normal, inner ring fault, outer ring fault, and rolling element fault.

3. The bearing fault diagnosis method based on neural networks and multi-criteria preference consensus as described in claim 1, characterized in that, The multilayer perceptron module includes an input layer, a hidden layer, and an output layer. A CBAM submodule is embedded between the hidden layer and the output layer. The CBAM submodule consists of channel attention units and spatial attention units, specifically: The input layer receives the standardized feature vector of the target bearing; The hidden layer performs a linear transformation using the weight matrix and bias vector, projecting the input of the input layer into a high-dimensional space. Then, the ReLU activation function is applied to introduce non-linearity, capturing feature interactions and generating preliminary high-dimensional features. The linear transformation is shown in the following equation: , The nonlinear activation is shown in the following equation: , In the formula, This represents the combined value of the k-th sample in the hidden layer. Let be the weight matrix of the i-th feature in the hidden layer. The bias vector of the hidden layer. This is the output after activating the k-th sample; The CBAM submodule reshapes the initial high-dimensional features output from the hidden layer into feature maps and inputs these feature maps into the channel attention unit. The channel attention unit weights the feature maps along the channel dimension based on the global average pooling and global max pooling results. The channel-weighted feature maps are then input into the spatial attention unit, which further enhances the channel-weighted feature maps based on the average pooling and max pooling results at each spatial location. Finally, the initial high-dimensional features are multiplied by the channel attention weights to obtain the enhanced channel features. These enhanced channel features are then multiplied by the spatial attention weights to output the weighted global high-dimensional features, as shown in the expression below: , , , , In the formula, F is the input feature map. This indicates that global average pooling is performed on each channel of F. This indicates that global max pooling is performed on each channel of F, and MLP is a shared multilayer perceptron. It is the Sigmoid activation function. For channel attention weights, Spatial attention weights, Indicates to and Perform the splicing operation; Indicates the kernel size as Convolution operations are used to capture local relationships in spatial locations; To enhance channel features, The weighted global high-dimensional features; The output layer receives the weighted global high-dimensional features output by the CBAM submodule, configures independent weight parameters for each state to be diagnosed, and calculates a feature pattern matching weight vector for each state through a linear transformation between the weight vector and the weighted global high-dimensional features. This achieves the correlation measurement between the target bearing and each state, and outputs the nonlinear contribution vector of the target bearing to each state, i.e., [ The formula for calculating the linear transformation between the weight vector and the weighted global high-dimensional features is as follows: , In the formula, Let be the nonlinear contribution of the k-th sample to the m-th state. Let m be the weight matrix for the m-th state. The weighted global high-dimensional feature of the k-th sample. This is the bias vector.

4. The bearing fault diagnosis method based on neural networks and multi-criteria preference consensus as described in claim 3, characterized in that, Channel attention units include: The global average pooling subunit calculates the average value for each channel of the feature map to obtain the global characteristics of that channel. The global max pooling subunit maximizes each channel of the feature map to capture the response at significant activation locations; The shared multilayer perceptron subunit receives the outputs of the global average pooling subunit and the global max pooling subunit, and learns the relationship between pooling features. The element-sum subunit adds the outputs of the global average pooling subunit and the global max pooling subunit in the shared multilayer perceptron subunit; The output sub-units are processed by using the Sigmoid activation function to compress the summation results of the element-sum sub-units to the range of (0, 1), thus obtaining the channel attention weights. Spatial attention units include: The global average pooling sub-unit calculates the mean of all channels at each location in space; The global max-pooling sub-unit calculates the maximum value of all channels at each location in the space. The splicing sub-unit concatenates the outputs of the global average pooling sub-unit and the global max pooling sub-unit to obtain a two-channel feature map. The convolutional computation subunit uses a convolutional kernel to convolve the two-channel feature maps to obtain a spatial weight map. The output sub-units are compressed into the range (0, 1) using the Sigmoid activation function to obtain the spatial attention weights.

5. The bearing fault diagnosis method based on neural networks and multi-criteria preference consensus as described in claim 1, characterized in that, Each bearing model's post-trained neural network multivariate decision-making auxiliary NN-MCDA model consists of multiple parallel neural network multivariate decision-making auxiliary NN-MCDA sub-models. To ensure differentiated diagnostic characteristics among these sub-models, each sub-model employs the same network structure and training process for independent parameter optimization. However, they differ in their training sample sampling strategy, feature constraints, and loss function parameters. The specific training steps for each post-trained neural network multivariate decision-making auxiliary NN-MCDA model are as follows: S41. Obtain multiple sets of historical bearing vibration samples of this model under different conditions, and extract the multi-domain features of each set of historical bearing vibration samples under different conditions, wherein the number of historical bearing vibration samples under each condition is equal. S42. Standardize the multi-domain features of each group of historical bearing vibration samples under different states to obtain the standardized feature vector of each group of historical bearing vibration samples under different states. S43. Divide multiple sets of historical bearing vibration samples under the same state into a training set and a validation set according to the proportion. Each set of historical bearing vibration samples in the training set / validation set carries its corresponding standardized feature vector. S44. Initialize multiple neural network multivariate decision-making auxiliary NN-MCDA sub-models, and set differentiated training strategies for each neural network multivariate decision-making auxiliary NN-MCDA sub-model from three levels. Specifically: First, set differentiated sample sampling weights or resampling probabilities to weight or resample the samples in the training set, thereby prompting each neural network multivariate decision-making auxiliary NN-MCDA sub-model to focus on different data distributions; Second, set feature subset masks for each neural network multivariate decision-making auxiliary NN-MCDA sub-model to constrain each neural network multivariate decision-making auxiliary NN-MCDA sub-model to focus on one feature type; Finally, set differentiated class weights in the loss function to adjust the misclassification penalty for samples in different states, thereby enhancing the recognition ability of each neural network multivariate decision-making auxiliary NN-MCDA sub-model on the preset emphasis state. S45. Input all the training sets corresponding to each state into each neural network multivariate decision aid NN-MCDA sub-model for independent training. During the training process, the samples in the training set are weighted and / or resampled according to the sample sampling weights set by the neural network multivariate decision aid NN-MCDA sub-model. This enables each neural network multivariate decision aid NN-MCDA sub-model to learn the complete mapping relationship from the input standardized feature vector to all states. After training, each neural network multivariate decision aid NN-MCDA sub-model becomes an "expert" that can work independently and outputs an evaluation probability vector covering all states for each input sample. S46. Use the validation set corresponding to each state to evaluate each trained neural network multivariate decision-making auxiliary NN-MCDA sub-model. Based on the evaluation results, select the parameters with the best overall performance on the validation set for each neural network multivariate decision-making auxiliary NN-MCDA sub-model, and solidify them into a trained neural network multivariate decision-making auxiliary NN-MCDA sub-model with expertise. During the evaluation, comprehensively consider the recall rate of the neural network multivariate decision-making auxiliary NN-MCDA sub-model in its emphasized state, as well as the false alarm rate and / or specificity in its non-emphasized state, to ensure the effectiveness and reliability of its expertise. S47. Integrate multiple trained neural network multivariate decision-making auxiliary NN-MCDA sub-models into a complete trained neural network multivariate decision-making auxiliary NN-MCDA model.

6. The bearing fault diagnosis method based on neural networks and multi-criteria preference consensus as described in claim 1, characterized in that, The specific steps for constructing an environmental state vector based on multiple initial expert assessment probabilities of the target bearing under different states are as follows: S511. Treat each vibration signal of the target bearing as a sample, and transform the multiple initial expert evaluation probabilities of each sample under different states into a decision matrix. The decision matrix for each sample is shown below: , In the formula, Let be the decision matrix for the k-th sample. Let be the probability of the nth expert's evaluation of the mth state; S512. Based on the decision matrix of all samples, construct an environmental state vector that includes consensus index, harmony degree, and iteration round. ,in Let be the relational consensus index of the nth expert on all samples in the t-th iteration. Let be the harmony degree at iteration t; the steps for calculating the relation-level consensus index in each iteration are as follows: S5121. Calculate the element-level consensus index. The calculation formula is as follows: , , In the formula, Let be the element-level consensus index of the nth expert on the mth state of the kth sample. Let be the probability of the nth expert's evaluation of the mth state of the kth sample. Let R be the group-weighted average evaluation probability of all experts for the m-th state of the k-th sample, R be the set of evaluation probabilities of all experts for all states of all samples, and N be the number of experts. For the normalized weight of the nth expert, and ; S5122. Calculate the scheme-level consensus index. The calculation formula is as follows: , In the formula, Let M be the scheme-level consensus index of the nth expert on all states of the kth sample, where M is the number of states; S5123. Calculate the relation-level consensus index. The calculation formula is as follows: , In the formula, Let K be the relational consensus index of the nth expert on all samples, and K be the number of samples. The formula for calculating harmony in each iteration is shown below: , In the formula, HD represents the harmony degree, and t is the iteration round. Let be the probability of the nth expert's evaluation of the mth state of the kth sample in the tth iteration.

7. The bearing fault diagnosis method based on neural networks and multi-criteria preference consensus as described in claim 1, characterized in that, The agent's action space is defined as Its value is generated through the Actor network, as shown below: , In the formula, In the t-th iteration The actions of the agent In the t-th iteration Feedback parameters generated by the agent for Agent's strategy network, The environment state vector is the input policy network. for The set of learnable parameters of the agent in the policy network; for The output layer weight matrix of the proxy, where ReLU is the modified linear unit activation function. for The hidden layer weight matrix of the proxy, Indicates will The state feature mapping function is converted into a feature vector that can be processed by the network. for The hidden layer bias vector of the proxy. for The output layer bias vector of the proxy uses the Sigmoid activation function to ensure... ; The agent's reward function is formed by a weighted combination of harmony, iteration rounds, and changes in the consensus index, as shown in the following formula: , In the formula, In the t-th iteration The agent's reward function, , , These are the weighting coefficients; The agent's action space is defined as Its value is generated through the Actor network, as shown below: , In the formula, In the t-th iteration The actions of an agent; In the t-th iteration The expert weight vector generated by the agent, In the t-th iteration The normalized weights of the nth expert generated by the agent; for Agent's strategy network, The environment state vector is the input policy network. for The set of learnable parameters of the agent in the policy network; for The output layer weight matrix of the proxy, where ReLU is the modified linear unit activation function. for The hidden layer weight matrix of the proxy, Indicates will The state feature mapping function is converted into a feature vector that can be processed by the network. for The hidden layer bias vector of the proxy. for The proxy's output layer bias vector, using Activation function, ensure and ; The agent's reward function is formed by the weighted sum of expert weights and consensus index, minus the squared penalty for changes in weights, as shown in the following formula: , In the formula, In the t-th iteration The agent's reward function; In the t-th iteration The normalized weights of the nth expert generated by the agent. and ; This is an evaluation item for the immediate effect of the strategy, used to encourage the allocation of high weights to high-consensus experts in order to directly improve the level of collective consensus. A coefficient used to control the intensity of punishment; As a constraint, the overall change range of the weight strategy is quantified through L2 norm, and drastic changes in the weight vector are penalized to ensure the continuity and robustness of weight allocation; Let be the weight change, representing the difference between the t-th iteration and the t-th iteration. Overall changes in weight allocation strategy during rounds of iteration ; It is a quantitative indicator of the "overall strategic change range"; In addition, in each iteration, Agent and After obtaining the feedback parameters and expert weight vector based on the same environmental state vector, the agent first calculates the group-weighted average evaluation probability of all experts for the m-th state of the k-th sample in the current iteration based on the expert weight vector. Then, the consensus index of each expert in the current iteration is calculated, and for those experts whose consensus index is ≤ consensus threshold, the consensus index is then calculated. Experts use feedback parameters to update the evaluation probability in the next iteration, and finally calculate the environment state vector in the next iteration based on the updated evaluation probability. In each iteration, for consensus index ≤ consensus threshold... The experts update their evaluation probabilities for the next iteration using the following formula: , In the formula, Let be the probability of the nth expert's evaluation of the mth state of the kth sample in the t-th iteration. Let be the group-weighted average evaluation probability of all experts for the m-th state of the k-th sample in the t-th iteration. In the t-th iteration Feedback parameters generated by the agent.

8. The bearing fault diagnosis method based on neural networks and multi-criteria preference consensus as described in claim 1, characterized in that, The specific training steps of the RL-CRP model for achieving consensus after training for each type of bearing are as follows: S521. Obtain multiple sets of historical bearing vibration samples of this model under different conditions, and extract the multi-domain features of each set of historical bearing vibration samples, wherein the number of historical bearing vibration samples under each condition is equal. S522. Standardize the multi-domain features of each group of historical bearing vibration samples to obtain the standardized feature vector of each group of historical bearing vibration samples. S523. Input the standardized feature vectors of all historical bearing vibration samples into the trained neural network multivariate decision-making auxiliary NN-MCDA model respectively, and obtain multiple initial expert evaluation probabilities for each group of historical bearing vibration samples under different states. S524. Divide all historical bearing vibration samples into training set and validation set according to the proportion. The number of historical bearing vibration samples in each state in the training set and validation set is equal. Each group of historical bearing vibration samples in the training set and validation set carries multiple initial expert evaluation probabilities in different states. S525. Input the initial expert evaluation probabilities of each group of historical bearing vibration samples in the training set under different states into the reinforcement learning consensus reaching process (RL-CRP) model for training, and obtain the initial reinforcement learning consensus reaching process (RL-CRP) model. S526. Input the initial expert evaluation probabilities of each group of historical bearing vibration samples in different states in the validation set into the initial reinforcement learning consensus reaching process RL-CRP model. Evaluate the consensus degree, classification accuracy and convergence rounds of the initial reinforcement learning consensus reaching process RL-CRP model on the validation set. When the evaluation results meet the preset performance indicators, solidify the initial reinforcement learning consensus reaching process RL-CRP model into a trained reinforcement learning consensus reaching process RL-CRP model. If the evaluation results do not meet the requirements, adjust the reward function weight, learning rate or training rounds, and retrain the reinforcement learning consensus reaching process RL-CRP model.