Method and device for diagnosing poor lubrication fault of equipment based on artificial intelligence
Through the feature domain perception and data mapping technology based on artificial intelligence, the problem of insufficient nonlinear relationship capture capability in lubrication status monitoring is solved, and the lubrication fault diagnosis with high precision and high generalization capabilities is achieved, which improves the reliability and safety of equipment operation.
Patent Information
- Application Number
- CN202510680454.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2025-07-08
AI Technical Summary
The existing lubrication state monitoring and diagnostic methods have limited ability to capture data nonlinear relationships, making it difficult to fully mine complex patterns in high-dimensional data, resulting in low diagnostic accuracy and unable to effectively improve the accuracy of lubrication state monitoring and the generalization ability of the model.
Using an artificial intelligence-based method, feature domain perception and data mapping are performed by obtaining a variety of real-time data during the operation of the device, and using a feature extraction model to map the data to a low-dimensional space. Combining adaptive divergence regularization and nonlinear activation functions, feature weights and activation functions are dynamically adjusted to achieve efficient diagnosis of lubrication state.
It improves the accuracy of equipment lubrication fault diagnosis, can quantify the physical correlation between expressing sensor data, significantly improves the accuracy of diagnosis and generalization of models, and ensures robustness and reliability in complex environments.
Smart Images

Figure CN120277537A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a method and device for diagnosing equipment lubrication failure based on artificial intelligence. Background Art
[0002] The lubrication state of equipment directly affects the operation efficiency and reliability of industrial systems. During long-term operation, poor lubrication can lead to increased equipment wear, increased energy consumption, and even major failures, seriously affecting production safety and economic benefits. However, existing lubrication state monitoring and diagnosis methods face multiple technical problems. How to improve the accuracy of lubrication state monitoring through intelligent means, enhance the generalization ability of the model, and at the same time meet the data management requirements in the industrial environment has become an urgent problem to be solved. However, the existing technology has limited ability to capture the non-linear relationship of data, making it difficult to fully explore complex patterns in high-dimensional data and reducing the diagnostic accuracy. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide a method and device for diagnosing equipment lubrication failure based on artificial intelligence, which can quantitatively express the physical correlation between various sensor data, fully explore data patterns, effectively capture data relationships, and ensure diagnostic accuracy.
[0004] In a first aspect, an embodiment of the present invention provides a method for diagnosing equipment lubrication failure based on artificial intelligence, the method comprising: obtaining operation monitoring data during the operation of a target device, wherein the operation monitoring data includes various real-time data related to the lubrication state during the operation of the device, and the various real-time data are derived from real-time monitoring data and periodic record data of sensors in the industrial field; based on the data characteristics of the operation monitoring data, performing feature domain perception on the correlation between the real-time data in the operation monitoring data to determine a feature domain perception result corresponding to the operation monitoring data; based on the feature domain perception result, performing data mapping on the operation monitoring data to determine target features in the operation monitoring data; classifying the target features according to the equipment failure probability to determine a lubrication state failure diagnosis result of the target device; and determining the lubrication condition of the target device based on the lubrication state failure diagnosis result.
[0005] Combined with the first aspect, an embodiment of the present invention provides a first implementation manner of the first aspect, wherein the step of classifying the target features according to the equipment failure probability to determine a lubrication state failure diagnosis result of the target device includes: determining a failure category score corresponding to the target feature; based on the failure category score, determining a category probability distribution corresponding to the target feature; and determining the highest probability category in the category probability distribution as the lubrication state failure diagnosis result of the target device.
[0006] In combination with the first aspect, the embodiments of the present invention provide a second implementation manner of the first aspect. Among them, the steps of performing feature domain perception on the correlation between real-time data in the operation monitoring data based on the data characteristics of the operation monitoring data and determining the feature domain perception result corresponding to the operation monitoring data include: calculating the perception weight corresponding to each feature according to the feature distance between each feature of the operation monitoring data; based on the perception weight, performing non-linear activation on the operation monitoring data to determine the feature domain perception result corresponding to the operation monitoring data.
[0007] In combination with the first aspect, the embodiments of the present invention provide a third implementation manner of the first aspect. Among them, the steps of performing data mapping on the operation monitoring data based on the feature domain perception result and determining the target feature in the operation monitoring data include: using the pre-constructed feature extraction model to perform non-linear activation on the operation monitoring data based on the feature domain perception result, mapping the operation monitoring data to a low-dimensional space, and determining the target feature of the operation monitoring data; among them, the construction method of the feature extraction model includes: obtaining the pre-constructed training sample set, inputting the training sample set into the initial neural network, and performing data capture on the training sample set through the neural network to obtain the feature output; among them, the training sample set includes the device operation samples related to the lubrication state during the operation of the device and the sample labels corresponding to the device operation samples; determining the local error corresponding to the feature output and determining the divergence measure of the feature output; based on the divergence measure, calculating the loss function of the initial neural network; according to the loss function and the local error, updating the parameters of the initial neural network; until the initial neural network meets the preset training requirements, constructing the feature extraction model based on the initial neural network. In combination with the first aspect, the embodiments of the present invention provide a fourth implementation manner of the first aspect. Among them, the local error is determined based on the global information fusion of the feature output of each layer of the initial neural network; the steps of updating the parameters of the neural network according to the loss function and the local error include: determining the similarity measure result corresponding to the feature output, and updating the learning rate of the neural network based on the similarity measure result; based on the updated learning rate, calculating the weight update value of the neural network; according to the local error, the weight update value, and the loss function, updating the parameters of the weight of the neural network. In combination with the first aspect, the embodiments of the present invention provide a fifth implementation manner of the first aspect. Among them, the steps of calculating the loss function of the initial neural network based on the divergence measure include: calculating the cross-entropy loss of the initial neural network according to the feature output; based on the divergence measure, calculating the adaptive divergence regularization loss term of the initial neural network; according to the cross-entropy loss, the adaptive divergence regularization loss term, and the preset regularization weight, calculating the loss function of the neural network.
[0008] Combined with the first aspect, the embodiments of the present invention provide a sixth implementation manner of the first aspect. Among them, the method for constructing a training sample set includes: obtaining pre-collected device operation samples, performing data annotation on the device operation samples based on the device lubrication state corresponding to the device operation samples, and constructing an initial sample set; performing preliminary sample expansion on the initial sample set according to the data distribution of the initial sample set to construct a first expanded sample set; determining target samples to be expanded from the first expanded sample set based on the sample dimensions of the first expanded sample set; performing linear interpolation processing on the target samples to be expanded based on a preset interpolation ratio coefficient to construct a second expanded sample set; performing data augmentation on the second expanded sample set to construct a third expanded sample set; and performing data correction on the third expanded sample set to construct a training sample set. Combined with the first aspect, the embodiments of the present invention provide a seventh implementation manner of the first aspect. Among them, the step of performing data augmentation on the second expanded sample set to construct a third expanded sample set includes: performing data perturbation on the second expanded sample set to determine the gradient change of the second expanded sample set; and performing data augmentation on the second expanded sample set based on the gradient change to construct a third expanded sample set. Combined with the first aspect, the embodiments of the present invention provide an eighth implementation manner of the first aspect. Among them, the step of performing preliminary sample expansion on the initial sample set according to the data distribution of the initial sample set to construct a first expanded sample set includes: determining the data gradient of the initial sample set, generating local perturbations of the initial sample set based on the data gradient; and performing data expansion on the initial sample set based on the local perturbations and a preset random noise to construct a first expanded sample set.
[0009] In the second aspect, the embodiments of the present invention provide a device for diagnosing lubrication failure of a device based on artificial intelligence. Among them, the device includes: a data acquisition module for acquiring operation monitoring data during the operation of a target device. Among them, the operation monitoring data includes various real-time data related to the lubrication state during the operation of the device, and the various real-time data are from the real-time monitoring data and periodic record data of sensors in the industrial field; a data processing module for performing feature domain perception on the relevance between the real-time data in the operation monitoring data based on the data characteristics of the operation monitoring data to determine the feature domain perception result corresponding to the operation monitoring data; a feature extraction module for performing data mapping on the operation monitoring data based on the feature domain perception result to determine the target features in the operation monitoring data; a classification module for classifying the probability of device failure of the target features to determine the lubrication state failure diagnosis result of the target device; and an output module for determining the lubrication condition of the target device based on the lubrication state failure diagnosis result.
[0010] The embodiments of the present invention bring the following beneficial effects: A method and device for diagnosing faults of poor equipment lubrication based on artificial intelligence provided by the embodiments of the present invention perform feature domain perception on various real-time data related to the lubrication state during the operation of the equipment, then perform data mapping to determine corresponding target features, explicitly fuse domain knowledge during the data processing process, perceive the physical coupling existing in the sensor data in the equipment lubrication system, identify sample clusters under similar working conditions, and quantify the physical correlation of the data. The fault type can be associated and mapped based on the degree of deviation of the feature from the normal distribution, effectively improving the accuracy of fault diagnosis.
[0011] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are realized and obtained by the structures specifically pointed out in the specification, claims, and drawings.
[0012] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given, and in conjunction with the accompanying drawings, the detailed description is as follows. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0014] Figure 1 It is a flowchart of a method for diagnosing faults of poor equipment lubrication based on artificial intelligence provided by the embodiments of the present invention;
[0015] Figure 2 It is a flowchart of another method for diagnosing faults of poor equipment lubrication based on artificial intelligence provided by the embodiments of the present invention;
[0016] Figure 3 It is a flowchart of a method for constructing a training sample set provided by the embodiments of the present invention;
[0017] Figure 4 It is a schematic diagram comparing the test accuracies of different data augmentation methods provided by the embodiments of the present invention;
[0018] Figure 5 It is a schematic diagram of a training loss convergence curve provided by the embodiments of the present invention;
[0019] Figure 6A schematic diagram for comparing the accuracies of embodiments of the present invention under different training sample sizes provided by embodiments of the present invention;
[0020] Figure 7 A schematic diagram of the experimental results of noise robustness provided by embodiments of the present invention;
[0021] Figure 8 A schematic diagram of the structure of a device lubrication defect fault diagnosis device based on artificial intelligence provided by embodiments of the present invention;
[0022] Figure 9 A schematic diagram of the structure of an electronic device provided by embodiments of the present invention. Detailed implementation manners
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are some, rather than all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0024] The embodiments of the present invention provide a device lubrication defect fault diagnosis method and device based on artificial intelligence, which can quantitatively express the physical connection between various sensor data, fully explore data patterns, and effectively capture data relationships to ensure the diagnosis accuracy. For ease of understanding of this embodiment, a device lubrication defect fault diagnosis method based on artificial intelligence disclosed in the embodiments of the present invention will be introduced in detail first. Refer to Figure 1 The flow schematic diagram of a device lubrication defect fault diagnosis method based on artificial intelligence shown in the figure. The method includes the following steps:
[0025] Step S102, obtain the operation monitoring data of the target device during operation.
[0026] In an embodiment of the present invention, real-time data related to the lubrication state during the operation of the device is obtained to determine the fault condition of the device. Among them, the operation monitoring data includes various real-time data related to the lubrication state during the operation of the device, and the various real-time data is sourced from sensors in the industrial field, such as real-time monitoring data and periodic record data from temperature sensors, pressure sensors, vibration sensors, etc., to ensure the timeliness and continuity of the data. The data storage format is a structured database format. In one embodiment, the attributes of the data include: Ra (running time), reflecting the total running time of the device from startup to the current time; Da (temperature value), the current temperature of the device collected from the temperature sensor; Fa (pressure value), the current pressure of the device obtained from the pressure sensor; La (vibration level), the vibration intensity recorded by the device vibration sensor; Ma (current intensity), the current intensity measured by the current sensor; Na (rotation speed), the real-time rotation speed of the device; Oa (voltage), the measured voltage value; Pa (load), the load condition of the device; Qa (flow rate), the flow rate of liquid or gas; Sa (sound level), the sound level generated during the operation of the device. It should be noted that this embodiment is only to illustrate a data format and type of the present invention. In practical applications, the attributes of the data are usually more than 10, and the number of data attributes may reach dozens or even hundreds. In this embodiment, 5 pieces of data are as shown in Table 1 below:
[0027] Table 1, Example of real-time data related to the lubrication state during the operation of the device:
[0028]
[0029] Step S104, based on the data characteristics of the operation monitoring data, perform feature domain perception on the correlation between the real-time data in the operation monitoring data, and determine the feature domain perception result corresponding to the operation monitoring data.
[0030] Step S106, based on the feature domain perception result, perform data mapping on the operation monitoring data to determine the target features in the operation monitoring data.
[0031] Fault signals of the equipment lubrication state (such as vibration, temperature, pressure, etc.) often exhibit complex non - linear coupling, such as the non - linear correlation between high - frequency vibration and temperature rise. Traditional linear models (such as PCA) or simple neural networks are difficult to capture such patterns. To efficiently extract and represent the non - linear features in the equipment lubrication state, the embodiments of the present invention map the original features to a new feature space that is more suitable for capturing the internal structure of the data. For example, in image recognition, the convolutional layer can capture local features; in time - series analysis, temporal features can be captured through the sliding window technique. Then, non - linear activation of the data is performed to make the features more abundant and diverse, so as to facilitate capturing complex patterns and relationships in the data. In specific implementation, considering the physical relevance of sensor data, for example, features such as temperature, vibration, current, etc. have physical coupling in the equipment lubrication system (such as when lubrication is poor, friction increases → temperature rises → vibration intensifies), domain awareness is carried out, and high - dimensional sensor data is mapped to a low - dimensional subspace with physical meaning to achieve efficient characterization of the equipment lubrication state. In the feature space, samples in the high - dimensional feature space have a local neighborhood structure. The embodiments of the present invention quantify the physical correlation of the data based on the gradient correlation between features by identifying sample clusters under similar working conditions (such as calculating the Euclidean distance between features), so as to maintain the physical rationality of the data distribution during interpolation or perturbation. Among them, the embodiments of the present invention dynamically capture the correlation between key features (such as the co - variation of vibration and temperature) through the above - mentioned domain awareness mechanism, and use multi - layer non - linear transformation to compress the data dimension, so that the captured features can not only separate different fault modes but also conform to the equipment operation mechanism. Compared with traditional methods, the embodiments of the present invention explicitly fuse domain knowledge in the process of data processing to ensure the interpretability of the embedded space (for example, a certain dimension may correspond to friction intensity, and another dimension reflects heat load). Finally, the obtained target features not only retain the key information of the original data but also eliminate redundant noise.
[0032] Step S108, perform equipment fault probability classification on the target features to determine the lubrication state fault diagnosis result of the target equipment.
[0033] Step S110, determine the lubrication condition of the target equipment based on the lubrication state fault diagnosis result.
[0034] In one embodiment, the extracted features can be input into the Softmax layer to be transformed into a fault probability distribution, so as to determine whether there is a fault related to the lubrication state of the current device, and give the specific fault type and severity, completing the multi-level state classification from normal to severely poor lubrication. In one embodiment, the classification categories include: normal, slightly poor lubrication, moderately poor lubrication, and severely poor lubrication. By inputting the output features after feature extraction into the Softmax layer for linear transformation, the non-linear patterns in the feature space are transformed into the original scores of each fault category, and then these scores are converted into probability values through the exponential normalization operation of Softmax, representing the probability distribution of different types of faults occurring in the current device. Finally, the category with the highest probability is selected as the lubrication state fault diagnosis result of the target device. For the above target features, fault probability classification can be performed, which can perform associated mapping on the fault type based on feature changes, quantitatively express the degree of deviation of the features from the normal distribution, and efficiently extract and represent the target features to improve the accuracy probability of fault diagnosis. When the device is operating normally, the data of various sensors will be stable within a specific numerical range, showing a compact distribution pattern; once a lubrication fault occurs, some key features will deviate significantly or form a special combination pattern. For example, when the bearing is under-lubricated, the energy of the high-frequency component of the vibration data will increase significantly, accompanied by a slow rise in temperature but stable pressure, and this feature combination will show a tailing phenomenon towards the high-vibration region in the data distribution diagram. Oil contamination is manifested as increased current fluctuations and a sudden increase in the energy of a specific sound frequency band, forming a discrete outlier distribution in the feature space. Through the above steps of data processing in the embodiments of the present invention, the distribution law of features can be captured. The whole process takes into account both the threshold breakthrough of a single feature and the non-linear relationship of the collaborative changes of multiple features. Further, the embodiments of the present invention can also adopt a data storage and management unit to provide the data storage and management functions required by the system, ensuring the security and availability of historical data and real-time data. The functions of the data storage and management unit include data storage, data retrieval, and data backup. A relational database or a non-relational database can be used to store structured and unstructured data to implement the data storage function, including information such as original sensing data, diagnosis results, and user feedback; the data retrieval function provides an efficient query interface, enabling users and other units to quickly access the required historical data and analysis results; the data backup function regularly backs up the system data to prevent data loss, and at the same time meets the data security and compliance requirements in the industrial environment.
[0035] Further, on the basis of the above embodiments, the embodiments of the present invention also provide another artificial intelligence-based method for diagnosing equipment lubrication failure. Figure 2 The flowchart of another artificial intelligence-based method for diagnosing equipment lubrication failure provided by the embodiments of the present invention is shown. Referring to Figure 2 FIG., the method includes the following steps:
[0036] Step S202: Obtain the operation monitoring data of the target device during operation.
[0037] Step S204: Calculate the perception weight corresponding to each feature according to the feature distance between each feature of the operation monitoring data.
[0038] In the embodiment of the present invention, by perceiving the relationship between each feature and other features, the expression ability of the network is optimized, so as to effectively process multi-dimensional and diverse input features. Among them, in order to ensure that the model always focuses on the most relevant feature patterns, the absolute distance between features is calculated in real time to determine the perception weight of each feature point, so as to realize dynamic weighting based on the distance information between features and dynamically quantify the correlation strength between different features. The perception weight of a feature can reflect the correlation between features. For example, when the lubrication is poor, features such as temperature, vibration, and current will show co-variation (such as an increase in temperature accompanied by an increase in vibration). The perception weight of a feature is calculated by the relative distance square ratio between the feature and other features. If a feature (such as vibration) is significantly different from most features (abnormal vibration), the perception weight of the feature increases, enhancing the contribution of the vibration feature. Specifically, the calculation method of the perception weight is expressed as:
[0039]
[0040] In the formula, is the perception weight of the i-th feature point; is the j-th feature of the output data of the (l - 1)-th layer of the feature space embedding neural network. Among them, The dimension of is [F 2 ; the dimensions of the numerator and denominator are both [F 2 , and the ratio is dimensionless. Therefore, through the normalization operation, that is, the denominator is the sum of the squares of the total distances, the absolute distance is converted into a relative weight, is dimensionless and the dimension is [1].
[0041] Step S206: Based on the perception weight, perform non-linear activation on the operation monitoring data to determine the perception result of the feature domain corresponding to the operation monitoring data.
[0042] In specific implementation, in the embodiment of the present invention, through the combined calculation of non-linear activation functions Reflect the feature distribution. For example, changes in the lubrication state can lead to a non-linear shift in the feature distribution (such as the temperature rising from linear to exponential), the ReLU activation function captures linear features (such as temperature changes during the stable stage), and the Sigmoid activation function fits non-linear mutations (such as vibration spikes during severe lubrication failure). Further, based on the correlation between current features, feature distribution, and feature gradient information, combined with the perception weights, the network can flexibly reconstruct the feature space for different working conditions and fault modes, dynamically select and adjust the mapping methods of each feature, so as to significantly improve the expression ability of non-linear coupling features of the equipment state while maintaining the compactness of the model. Correspondingly, the corresponding feature domain perception result is determined in the following way, and the calculation method is expressed as:
[0043]
[0044] In the formula, is the i-th feature of the output data of the (l - 1)-th layer of the feature space embedding neural network; α p is the adaptive activation coefficient of the feature, representing the emphasis of the feature in the activation function combination; β p is the non-linear adjustment factor of the feature, representing the adjustment strength for the Sigmoid activation function; Re() is the ReLU activation function, and Sig() is the Sigmoid activation function; is the perception weight of the i-th feature, representing the correlation between this feature and other features. Among them, is used as the input feature, and the dimension is [F]; α p , β p are used as adaptive coefficients, dimensionless, that is, the dimension is [1]; is used as the perception weight, which has been normalized, dimensionless, that is, the dimension is [1]; Re(), Sig() are activation functions, and the output dimension is the same as the input, and the dimension is [F]; Therefore, and are both [F], and after weighting and multiplying by the dimensionless it is still [F], and they can be directly added, and the dimensions of all terms are the same.
[0045] The fault characteristics of the lubrication system often manifest as the non-linear coupling of multi-sensor parameters (such as vibration, temperature, current), and the importance of each parameter varies dynamically in different fault stages. Early wear may be dominated by high-frequency vibration, while temperature may become a key indicator during severe faults. In the embodiments of the present invention, the device mechanism knowledge and data characteristics are deeply integrated, the device operation mechanism (such as the physical coupling relationship between vibration and temperature) is encoded as learnable feature weights, and an adaptive activation mechanism is used to dynamically adjust the attention degree to different features, converting physical constraints into differentiable optimization objectives. Traditional neural networks use fixed activation functions and equal feature processing methods, unable to capture this dynamically changing physical association. In the embodiments of the present invention, through the adaptive activation coefficient α of the feature p dynamically balance the ratio of the two, reflecting the feature gradient information. If the gradient shows non-linear dominance ( increasing), then reduce the adaptive activation coefficient of the feature and enhance the weight of the Sigmoid activation function. In addition, this adaptive activation coefficient is dynamically adjusted according to the gradient of the loss function, enabling the model to adaptively capture the co-variation pattern of key features in lubrication faults while suppressing irrelevant noise.
[0046] Furthermore, the adaptive activation coefficient and the non-linear adjustment factor of the feature are adaptively adjusted according to the error feedback during the training process, and the calculation method is expressed as:
[0047]
[0048]
[0049] In the formula, ← represents the parameter update operation; η p is the learning rate of the neural network for feature space embedding; Lp is the loss function of the neural network for feature space embedding; is the partial derivative symbol. Among them, η p as the learning rate, is dimensionless, with a dimension of [1], and the loss function L p is dimensionless; because L p is dimensionless, α p is dimensionless, so the dimension of is still [1]; α p , β p are dimensionless hyperparameters, with a dimension of [1]. The dimensions of each added term are consistent, and the dimensions of all terms are [1], so they can be directly added.
[0050] Step S208, based on the feature domain perception result, use the pre-constructed feature extraction model to perform non-linear activation on the operation monitoring data, map the operation monitoring data to a low-dimensional space, and determine the target features of the operation monitoring data.
[0051] In specific implementation, the embodiment of the present invention uses a preset feature extraction model to capture target features. Among them, by using a non-linear activation function and combining the above-mentioned feature domain perception function, non-linear features in device data can be effectively captured. For example, complex frequency components in the lubrication state signal. In device data, there may be features highly related to lubrication faults, and the weights of these key features can be more accurately enhanced, thereby improving the accuracy of diagnosis. It is expressed as:
[0052]
[0053] In the formula, is the feature output of the l-th layer of the feature space embedding neural network, representing the representation extracted by this layer from the input device data; is the weight of the l-th layer of the feature space embedding neural network; is the output data of the (l - 1)-th layer of the feature space embedding neural network, representing the features after feature extraction or transformation in the previous layer; is the bias of the l-th layer of the feature space embedding neural network; f p () is a non-linear activation function; is the feature domain perception function. Among them, as the input feature, represents the device operation sample, such as the real-time data related to the lubrication state during the operation of the above device, with the dimension of [F]; as the weight matrix, with the dimension of [F], is used to map the input feature from [F] to the next layer dimension. as the feature domain perception function, the output dimension is the same as the input, with the dimension of [F]; as the bias term, with the dimension of [F], is consistent with the output feature dimension; f p as the activation function, is a dimensionless operation (such as ReLU, Sigmoid), and does not change the dimension; therefore, the linear part has the dimension of [F], and the bias term with the dimension of [F] can be directly added.
[0054] In the embodiments of the present invention, the feature extraction model adopted is to establish a corresponding model based on a preset training sample set through a machine learning modeling unit, so as to realize the intelligent evaluation and prediction of the equipment lubrication state. Aiming at the problems faced by feature extraction and dimensionality reduction in high-dimensional equipment data, such as local optimal traps, gradient instability, and difficulty in removing redundant features, in one implementation, the feature extraction model of the embodiments of the present invention adopts a feature space embedding neural network to achieve efficient feature extraction. Further, the non-linear feature capture ability is improved through the feature domain perception of adaptive non-linear mapping, and the robustness and generalization ability of hierarchical features are enhanced through an adaptive divergence regularization mechanism, so that the model can complete feature representation more stably and accurately, avoid falling into local optimum, and reduce the interference of noise and redundant features. Specifically, the construction method of the feature extraction model of the embodiments of the present invention is as follows:
[0055] 1) Obtain a pre-constructed training sample set, input the training sample set into the initial neural network, and capture data of the training sample set through the neural network to obtain a feature output.
[0056] First, the feature extraction model of the embodiments of the present invention is constructed based on a preset initial neural network. By initializing the network structure, a multi-layer perceptron model including an input layer, multiple hidden layers, and an output layer is constructed, where the number of nodes in the input layer matches the equipment feature dimension, the number of nodes in the hidden layer is compressed proportionally, and the output layer corresponds to the number of fault categories. Then, the training stability is ensured by randomly initializing the network parameters and performing normalization processing, and the non-linear features in the equipment data are extracted layer by layer by using a non-linear activation function and a feature domain perception function. In one implementation, the way to initialize the parameters of the neural network is random initialization, and the initialized parameters follow a normal distribution with a mean of 0 and a variance of the identity matrix. During the training process, the input equipment data will be mapped layer by layer and processed through the non-linear activation function of each layer of the neural network, and finally the features after feature extraction, that is, the feature output, will be output. Among them, the embodiments of the present invention adopt a feature space embedding neural network, which is composed of a multi-layer perceptron model, including an input layer, multiple hidden layers, and an output layer. The number of nodes in each layer is set according to the feature dimension of the equipment data. For example, the number of nodes in the input layer is the same as the number of input features, the number of nodes in multiple hidden layers is a hyperparameter set artificially, usually set between 0.1 times and 0.5 times the number of nodes in the input layer, the number of nodes in the output layer is the same as the number of categories of the equipment data, and the output layer is only used for the neural network to predict the sample label and is not used as the output feature of feature extraction. The last hidden layer of the neural network is used as the output feature of feature extraction.
[0057] 2) Determine the local error corresponding to the feature output and perform divergence measurement on the feature output.
[0058] After performing feature domain perception processing on its training sample set, the corresponding local error and divergence metric can be determined. Among them, the local error is determined based on global information fusion of the feature outputs of each layer of the initial neural network. Local error correction and global information integration are performed to timely correct the biases generated in feature extraction, so that the model can still maintain global accuracy and consistency under complex mappings, which helps to maintain global accuracy and consistency under complex mappings. For example, there may be local anomalies in the lubrication state signal, and the weights of these features can be adjusted in a timely manner through local error correction. At the same time, global information is combined to avoid overfitting to specific anomalies, and the model's ability to perceive long-term trends is improved.
[0059] In specific implementation, by calculating the relative change rate of the feature output of the current layer to capture the mutation of local features (such as instantaneous shocks in vibration signals). Depending on the feature difference between the current layer and the previous layer, the calculation of local information is characterized. Further, by calculating the consistency constraint between hierarchical features, which characterizes the calculation of global information. For example, in lubrication diagnosis, it is ensured that the time-domain features captured by the shallow layer (such as current ripple) are consistent with the frequency-domain features extracted by the deep layer (such as vibration harmonics). Further, the way to correct each layer of features by calculating the local error is expressed as:
[0060]
[0061] In the formula, is the local error of the l-th layer of the feature space embedding neural network; is the feature gradient of the l-th layer of the feature space embedding neural network; is the feature output of the l-th layer of the feature space embedding neural network; is the feature output of the (l - 1)-th layer of the feature space embedding neural network; γ pcr is the global information fusion adjustment factor; is the feature output of the (l + 1)-th layer of the feature space embedding neural network. Preferably, γ pcr is set to 0.1. is the ratio calculation of norms, dimensionless, with a dimension of [1], is the multiplication of the hyperparameter and the ratio of norms, dimensionless, with a dimension of [1]. Therefore, and have the same dimension, both [F].
[0062] In order to measure the divergence of features, embodiments of the present invention calculate the divergence of feature distribution based on the difference between features and their means to measure discreteness. For example, when highly correlated features such as temperature and vibration coexist, features highly related to the lubrication state can be highlighted while suppressing irrelevant information with high randomness, improving the reliability of diagnosis. The calculation method is expressed as:
[0063]
[0064] In the formula, is the divergence metric of the feature distribution in the l-th layer; is the i-th feature of the output data of the l-th layer of the feature space embedding neural network; is the mean of the output data of the l-th layer of the feature space embedding neural network. After both the numerator and denominator are accumulated, a specific value is obtained, so it is dimensionless, with a dimension of [1], the same as that of . Embodiments of the present invention capture the dispersion of feature values in the distribution space by squaring and summing the differences between each feature value and its mean. For example, if the feature value is far from the mean (such as some vibration features increasing abnormally during lubrication failure), the value of the numerator will increase significantly; and the absolute scale of the feature value is normalized to eliminate the influence of differences in the dimensions of different sensors, calculate the proportion of the deviation degree of each feature value from the feature mean of this layer relative to the overall feature scale, quantify the discreteness of feature values in each layer of the neural network, and dynamically identify key feature patterns related to faults. The larger the finally obtained divergence value, the more obvious abnormal discrete patterns exist in the features of this layer (such as several vibration parameters deviating significantly from the normal range), and these patterns often correspond to typical features of equipment faults.
[0065] 3) Based on the divergence metric, calculate the loss function of the initial neural network.
[0066] In specific implementation, calculate the cross-entropy loss of the initial neural network according to the feature output, and calculate the adaptive divergence regularization loss term of the initial neural network based on the divergence metric. Further, according to the cross-entropy loss, the adaptive divergence regularization loss term, and a preset regularization weight, calculate the loss function of the neural network. Embodiments of the present invention combine the adaptive divergence regularization loss term with the cross-entropy loss to form the loss function of the initial neural network, and the calculation method is expressed as:
[0067] Lp = Lp cla + γ pyr Lp ADR
[0068] In the formula, Lp is the loss function of the feature space embedding neural network; Lp cla is the cross-entropy loss; γ pyrγ is the weight of the regularization term, which is used to balance the relative influence of the adaptive divergence regularization loss term and the cross-entropy loss. Preferably, γ pyr is set to 0.3. In this formula, each term is a dimensionless parameter, so this formula is a dimensionless numerical calculation. Based on this, through the synergistic effect of the domain-aware mechanism and adaptive regularization, corresponding problems can be effectively addressed to achieve accurate diagnosis. For the adaptive divergence regularization loss term, the embodiment of the present invention adopts a hierarchical feature enhancement method based on adaptive divergence regularization to solve problems such as the sparsity, noise, and uneven distribution of high-dimensional features, enabling the feature embedding neural network to focus more on useful features and suppress redundant features. Moreover, an adaptive divergence regularization term is used to dynamically adjust the sparsity and importance of each layer of features for calculating the adaptive divergence regularization loss term, which is expressed as:
[0069]
[0070] In the formula, Lp ADR is the adaptive divergence regularization loss term; λ pbv is the regularization coefficient of the feature space embedding neural network; is the divergence metric of the feature distribution of the l-th layer; Dp() is the feature distribution divergence calculation function; is the output data of the l-th layer of the feature space embedding neural network; || is the absolute value operation. Preferably, λ pbv is set to 0.3. Among them, as the divergence metric, is dimensionless, with a dimension of [1]; represents the output of the activation function, after accumulation, is dimensionless, with a dimension of [1]; λ pbv is a hyperparameter, dimensionless, with dimensions all of [1]. Therefore, the loss is calculated from dimensionless terms and is also dimensionless, with dimensions all of [1]. In this formula, the embodiment of the present invention first calculates the feature dispersion index (i.e., distribution divergence), by comparing the deviation amplitude of each feature value from the average value of the layer where it is located, and then performing a normalization process in combination with the overall feature scale. When the equipment has abnormal lubrication, relevant features such as vibration and temperature will produce collaborative fluctuations, and at this time, the dispersion index will increase significantly; then, the constraint intensity is dynamically adjusted according to the dispersion index, automatically reducing the restriction strength for the feature layer with high dispersion to ensure that the fault-sensitive features are retained, while strengthening the constraint for the stable feature layer with low dispersion to filter out random interference. For example, in the monitoring task of a generator set, it can not only keenly capture the tiny changes in vibration features caused by early wear, but also effectively ignore the irrelevant fluctuations caused by environmental factors, greatly improving the accuracy of fault warning while maintaining a low false positive probability. Further, in order to adaptively adjust the regularization strength of each layer, the regularization coefficient is updated based on the divergence information of the hierarchical features, and the calculation method is expressed as:
[0071]
[0072] In the formula, and η pce are both dimensionless hyperparameters, is the initial regularization coefficient of the feature space embedding neural network; η pce is the learning rate for regularization adjustment. Preferably, η pce is set to 0.01. is also dimensionless and represents the divergence metric of the feature distribution of the l-th layer. Therefore, all the formulas in this formula are dimensionless numerical calculations.
[0073] 4) Update the parameters of the initial neural network according to the loss function and the local error.
[0074] In specific implementation, perform similarity measurement on the feature output. Based on the similarity measurement, update the learning rate of the neural network; based on the updated learning rate, calculate the weight update value of the neural network; according to the local error, the weight update value, and the loss function, update the parameters of the weight of the neural network. In the embodiment of the present invention, the learning rate of the initial neural network is dynamically adjusted according to the current training state, the learning rate is adaptively corrected through gradient information, and the learning rate is more refinedly dynamically adjusted by combining the similarity information between features, so as to avoid excessive oscillation or local divergence in the training process in a complex high-dimensional space. The calculation method is expressed as:
[0075]
[0076] In the formula, ← is the parameter update operation; α pec is the learning rate adjustment factor; β pec is the similarity regularization factor; is the similarity measurement of the features of the l-th layer of the initial neural network. Specifically, the distance between feature vectors is calculated through cosine similarity to measure their similarity. For example, finding patterns with higher consistency in multiple sensor signals helps to enhance the generalization ability of the diagnostic model under specific working conditions. Preferably, α pec is set to 0.3, and β pec is set to 0.2. Among them, is the ratio of the norm of the gradient to the weight, dimensionless, with a dimension of [1]; It is the ratio of the similarity measure to the feature norm, also dimensionless, with a dimension of [1]. Therefore, all terms are dimensionless and have a dimension of [1], and can be directly added. Further, a gradient optimization method based on the gradient vector is adopted to overcome the phenomenon that the conventional gradient descent is prone to falling into local optima. By selecting the optimal activation function and dynamically adjusting the gradient, the problem that the traditional method is prone to falling into local optima in the high-dimensional feature space can be avoided, and the gradual change mode of lubrication faults can be better modeled to ensure that useful diagnostic information can still be extracted in edge cases. In order to adaptively select the most suitable non-linear activation function according to the error signal and gradient of the current layer, it is realized by the way of activation function selection, expressed as:
[0077]
[0078] In the formula, is the selected optimal activation function; is the set of alternative activation functions, including candidate activation functions such as ReLU, LeakyReLU, Sigmoid, etc.; is the gradient of the i-th activation function; the i-th activation function. Among them, is the selection function, and there is no operation represented by dimension. Further, combined with the above loss function, the weights of the initial neural network are updated and overall convergence is achieved to ensure the final accuracy and robustness of feature extraction, expressed as:
[0079]
[0080] In the formula, β pcb is the weight update learning rate of the feature space embedding neural network; is the weight gradient of the l-th layer of the feature space embedding neural network. Preferably, β pcb is set to 0.02. have the same dimension, both [F], β pcb is a hyperparameter, dimensionless, with a dimension of [1], is the ratio of the norms, dimensionless, with a dimension of [1]. Therefore, in this update formula, the dimensions of all terms are [F] and can be directly added, and the dimensions of all terms are consistent.
[0081] The calculation method of the weight update value is expressed as:
[0082]
[0083] In the formula, is the update value of the weight of the l-th layer of the initial neural network, representing the direction and amplitude of the weight adjustment of this layer in this iteration; is the weight gradient of the l-th layer of the feature space embedding neural network; is the gradient vector of the l-th layer of the feature space embedding neural network, representing the gradient information for all weights of the l-th layer of the feature space embedding neural network; γ p is the regularization coefficient of the feature space embedding neural network; || || F is the Frobenius norm; |||| is the L2 norm. Preferably, η p is set to 0.3. Among them, is the updated value of the weight, with the same dimension as the weight matrix and the dimension is [F]; As the gradient of the weight, since is the weight matrix with the dimension of [F], and the loss L p is dimensionless, so has the dimension of [F]; As the square of the gradient norm, it is dimensionless and the dimension is [1]; As the square of the weight norm, it is dimensionless and the dimension is [1]; γ p is a hyperparameter, dimensionless and the dimension is [1]; η p is a hyperparameter, dimensionless and the dimension is [1]. Therefore, has the dimension of [F], which is consistent with the dimension of .
[0084] Among them, in order to reduce the probability of instability of the neural network, the weights are normalized according to the input distribution information, which is expressed as:
[0085]
[0086] In the formula, is the weight of the l-th layer of the neural network; is the number of neurons in the l-th layer of the feature embedding neural network; is a uniform random number in the interval [0,1]; Var() is the variance calculation function. Among them, has the dimension of [F 2 , and after taking the square root, it is [F], is a dimensionless factor, and the number of neurons is dimensionless. Therefore, has the dimension of [F]·[1] = [F].
[0087] 5) Until the initial neural network meets the preset training requirements, a feature extraction model is constructed based on the initial neural network.
[0088] Step S210, classify the target features for the equipment failure probability to determine the lubrication state fault diagnosis result of the target equipment.
[0089] Step S212: Determine the lubrication condition of the target device based on the lubrication state fault diagnosis result.
[0090] Another device lubrication fault diagnosis method based on artificial intelligence provided by the embodiments of the present invention performs feature domain perception on the operation monitoring data from multiple sensors, dynamically calculates the correlation weights between features through a feature domain perception mechanism, and adaptively adjusts the feature mapping method by combining the advantages of ReLU and Sigmoid activation functions. Through explicit modeling of dynamic weights, adaptive activation, and gradient optimization, fully mine the data patterns, capture their regular structures or statistical characteristics, and mine the essential characteristics of the device lubrication state to correctly classify the features and the relationships between features, so as to achieve efficient extraction and representation of the target features. Among them, the feature extraction model adopted is constructed based on a feature space embedding neural network, and, combined with an adaptive divergence regularization mechanism, dynamically constrain the sparsity of each layer of features, highlight key features and suppress noise, and enhance the resistance of the model to noise and redundant features in high-dimensional device data. Also, by dynamically adjusting the activation function and learning rate, optimize the stability of feature extraction and model training, and solve the problem of local optimal traps in high-dimensional data.
[0091] Furthermore, the embodiments of the present invention also design a training sample set. Figure 3 The flowchart of a method for constructing a training sample set provided by the embodiments of the present invention is shown. Refer to Figure 3 This method includes the following steps:
[0092] Step S10: Obtain the pre-collected device operation samples, and based on the device lubrication state corresponding to the device operation samples, perform data annotation on the device operation samples to construct an initial sample set.
[0093] The device operation samples refer to the above embodiments, and the embodiments of the present invention will not be elaborated further. Further, annotate the collected data to construct an initial sample set. In one implementation, the annotation method of the present invention is manual annotation, where the annotation categories include: normal, mild lubrication deficiency, moderate lubrication deficiency, and severe lubrication deficiency. Further, since the acquisition, annotation, and preprocessing of device training data are time-consuming and laborious, if the training samples are insufficient, it is easy to cause poor generalization ability of the model and affect the accuracy of the model. Therefore, it is necessary to expand the data of the initial sample set. However, traditional data expansion methods include random noise enhancement, fixed interpolation coefficients, etc., which are usually fixed or vary within a relatively wide range, and it is difficult to generate samples with strong diversity and uniform distribution, resulting in insufficient fitting ability of the model for data in sparse regions and easy classification errors of the model on the boundary samples of device lubrication faults.
[0094] In response to this, the embodiment of the present invention adopts a bilinear interpolation data augmentation method based on local perturbation and random noise, and uses a preset data augmentation unit to perform data augmentation, avoiding the samples from being overly concentrated in a certain local area, so as to increase the diversity and richness of the data and improve the generalization ability of the model. At the same time, the embodiment of the present invention also dynamically optimizes the sample weights through an adaptive learning rate adjustment mechanism, and uses gradient enhancement and perturbation functions to locally optimize the augmented data, ensuring the stability and efficiency of the model training process, mining the potential features of the device data, and effectively improving the adaptability of the model to the complex device data distribution. In specific implementation, refer to the following steps S11-S14.
[0095] Step S11: Perform preliminary sample augmentation on the initial sample set according to the data distribution of the initial sample set to construct a first augmented sample set.
[0096] First of all, the embodiment of the present invention analyzes the distribution of device training data and generates new device data samples in the high-dimensional feature space, so as to effectively augment the original device data. Among them, after finding the concentrated area and scarce area of the device data, the embodiment of the present invention performs data augmentation based on this distribution situation to ensure that the new device data samples are more evenly distributed in the space. The kernel density estimation method can be used to calculate the local density of sample points in the feature space. If the sample density in a certain area exceeds the threshold (such as the top 30% quantile), it is determined as the concentrated area, otherwise it is the scarce area. In one implementation manner, preliminary sample augmentation is preferentially performed on the scarce area to construct a first augmented sample set.
[0097] In specific implementation, the expansion is realized based on local perturbation and noise terms, expressed as:
[0098] X′ c1i =f c (X ci ,α c ,β c )
[0099] In the formula, X′ c1i is the sample after the first augmentation, representing the device data of the i-th sample after distribution expansion; X ci is the original device data sample, representing the position of the i-th sample in the feature space; f c () is the expansion operator, which generates new samples based on the original device data sample; α c is the parameter controlling the expansion amplitude, representing the expansion range in the feature space; β c is the parameter controlling the noise amplitude, representing the diversity of the augmented data. Preferably, α c is set to 0.3, β cSet to 0.2. In specific implementation, the embodiment of the present invention generates local perturbations of the initial sample set based on the data gradient of the initial sample set, and then expands the data of the initial sample set based on the local perturbations and the preset random noise to construct the first expanded sample set. The specific process is expressed as:
[0100]
[0101] In the formula, δ c (X ci ) is the local perturbation term of the original device data sample, and the perturbation is generated based on the local gradient; ∈ c (X ci ) is the random noise term of the original device data sample, and the amplitude is controlled by the adaptive mechanism; is the sample weight of the t-th iteration and is a training parameter. Among them, the calculation method of the perturbation based on the implementation characteristics of the local perturbation and the noise term is expressed as:
[0102]
[0103] In the formula, σ c is the regularization parameter of the local perturbation, representing the perturbation amplitude; is the gradient of the i-th sample in the feature space; is the second-order gradient of the i-th sample in the feature space; μ c is the noise mean; is the noise variance; is the normal distribution. Preferably, σ c is set to 0.5, μ c is set to 0.1, is set to 0.3. Among them, X ci is represented as the feature vector of the original sample, and the dimension is [F], that is, the dimension of the feature vector; is the sample weight, dimensionless, that is, the dimension is [1], because the weight is a scalar proportionality factor; α c is the expansion amplitude, dimensionless, that is, the dimension is [1], and is used to adjust the amplitude of the perturbation term; δ c (X ci ) is the local perturbation term, and the dimension is the same as that of X ci and the dimension is [F].
[0104] Furthermore, derive Among them, is the gradient, and the dimension is [F]; is the second-order gradient, and the dimension is [F]; σ c is the regularization parameter, dimensionless, that is, the dimension is [1], and is used to balance the amplitude of the gradient term; β c is the noise amplitude, dimensionless, that is, the dimension is [1]; ∈c (X ci ) is used as a random noise term, with the same dimension as X ci , and the dimension is [F]; further, the noise distribution is the same as that of X ci with the same dimension. Therefore, all terms of each added item have the same dimension, that is, α c ·δ c (X ci )([F]), β c ·∈ c (X ci )([F]) can be directly added.
[0105] Step S12: Determine the target sample points to be augmented from the first augmented sample set based on the sample dimension of the first augmented sample set; perform linear interpolation processing on the target sample points to be augmented based on a preset interpolation ratio coefficient to construct a second augmented sample set.
[0106] In order to make the distribution of the augmented device data smoother in the feature space, the embodiment of the present invention further performs secondary augmentation on the first augmented sample set, and constructs a second augmented sample set by performing linear interpolation processing on it. Among them, the embodiment of the present invention uses bilinear interpolation calculation to estimate the new sample through the eigenvalue of adjacent samples, avoiding the generation of device data being too concentrated in a certain local area and enhancing the generalization ability of the device data. In one implementation, the corresponding target neighborhood sample points and the number of target neighborhood sample points can be determined according to the dimension number of the feature space and the meaning represented by each dimension. By dividing the high-dimensional feature space into grids according to the quantiles of each dimension (for example, the two-dimensional space is divided into rectangular regions), and positioning the local grid unit where it is located to determine the target interpolation point X ci . Among them, 4 nearest neighbor samples at the vertex of the unit can be selected as X c11 , X c12 , X c21 , X c22 . For example, if the feature space only contains two dimensions of temperature and vibration, and X ci falls within the interval [50°C, 60°C] × [0.3g, 0.4g], then the 4 neighborhood points are the samples at the four corners of the rectangle. In the embodiment of the present invention, the determined multiple target neighborhood sample points (such as X c11 , X c12 , X c21 , X c22 ) are characterized as the 4 vertex sample points of the feature space of the target interpolation point X ci , which can indicate the 4 sample features with the most differentiation. For example: X c11 is used to indicate the lower left neighborhood (low temperature, low vibration) of the target point X ci ; Xc12 For indicating the lower right neighborhood (high temperature, low vibration) of the target point X ci ; X c21 For indicating the upper left neighborhood (low temperature, high vibration) of the target point X ci ; X c22 For indicating the upper right neighborhood (high temperature, high vibration) of the target point X ci . The target neighborhood sample points in the embodiments of the present invention represent the characteristic extreme cases around the target point (such as combinations of high / low temperature and high / low vibration), which come from the samples actually existing in the corresponding training set, ensuring that the interpolation result conforms to the actual data distribution. If the target point X ci moves, the neighborhood points will be reselected. When the feature space is a 3D space and the sample has 3 features, then, the corresponding neighborhood points are 8, corresponding to the vertices of the cube respectively, and trilinear interpolation can be used for data augmentation. Similarly, N-dimensional hypercube interpolation requires 2 N neighborhood points, but to avoid a sharp increase in the calculation amount due to too many dimensions, the embodiments of the present invention can use the method of uniformly selecting 4 neighborhood points for data augmentation. If the sample quantity of all neighborhood points is greater than 4, then randomly select 4 of them.
[0107] Further, corresponding to the above neighborhood points, the data interpolation process is expressed as:
[0108] X′ c21 =(1 - λ c )(1 - γ c )X c11 +λ c (1 - γ c )X c12 +(1 - λ c )γ c X c21 +λ c γ c X c22
[0109] In the formula, X′ c2i is the second augmented sample, representing the device data after bilinear interpolation of the i-th sample; X c11 is the first neighborhood sample point of the first augmented sample, X c12 is the second neighborhood sample point of the first augmented sample, X c21 is the third neighborhood sample point of the first augmented sample, X c22 is the fourth neighborhood sample point of the first augmented sample, representing 4 device data sample points in the adjacent area of the first augmented sample in the feature space; λ c is the first interpolation ratio coefficient, γ c is the second interpolation ratio coefficient. Among them, X c11,X c12 ,X c21 ,X c22 are all neighborhood samples, with the same dimension, and the dimension is [F]; λ c ,γ c As the interpolation coefficient, it is dimensionless, that is, the dimension is [1], and it is the proportionality coefficient after distance normalization. Among them, the above first interpolation proportionality coefficient and the second interpolation proportionality coefficient are based on adaptive weight adjustment to achieve smooth distribution of interpolation, and are dynamically updated according to the distance from the neighborhood sample points. The calculation method is expressed as:
[0110]
[0111]
[0112] In the formula, || || is the L2 norm, representing Euclidean distance calculation; α c is the first smoothness adjustment parameter, β c is the second smoothness adjustment parameter. Preferably, α c is set to 0.1, and β c is set to 0.2. Among them, λ c and γ c values are determined by the distance between the target point X ci and the neighborhood points. The closer the neighborhood point is to the target point, the greater its corresponding weight (because the smaller the distance term in the denominator, the closer the coefficient is to 1). For example, if X ci is close to X c11 , then λ c →0 and γ c →0. At this time, the weight of X c11 (1 - λ c ) (1 - γ c ) → 1. Such a coefficient design ensures that the sum of weights is 1, and the interpolation result is always within the convex hull formed by the 4 neighborhood points, avoiding extrapolation. λ c and γ c are both dimensionless calculations; therefore, all terms of each added term have the same dimension, and all 4 terms are [F], and can be directly weighted and added.
[0113] Furthermore, the above first augmented sample set and the second augmented sample set are both generated by the data augmentation unit using corresponding augmentation methods. Corresponding to the first augmented sample set, its local perturbation term is determined based on the sample weight. In addition, the second augmented sample set achieves smooth distribution of interpolation through adaptive weight adjustment. The embodiment of the present invention also adaptively regulates the learning rate of the sample weight during the training process, and adjusts the learning rate in real time according to the error change of the local characteristics of the device data, which can avoid the problem of model divergence caused by too large a rate or slow training caused by too small a rate, and is expressed as:
[0114]
[0115] In the formula, is the sample weight learning rate for the (t + 1)-th iteration, is the sample weight learning rate for the t-th iteration; ΔL c is the change in loss between the current and the previous iteration; γ c is the learning rate adjustment factor. As the learning rate, it is dimensionless and has the same dimension as the loss function L c , both being dimensionless, i.e., with dimension [1]; ΔL c As the change in loss, it is dimensionless, i.e., with dimension [1]; γ c As the adjustment factor, it is dimensionless, i.e., with dimension [1].
[0116] Where is the gradient norm, also dimensionless, i.e., with dimension [1]; therefore, all terms of the above summation terms have the same dimension, and (1 - γ c ·ΔL c ) is also dimensionless, i.e., with dimension [1], and its dimension remains unchanged after multiplying with . Among them, the data expanded in the above steps can be input into a preset classifier model (such as random forest, support vector machine, etc.) to obtain its predicted label, and then the cross-entropy loss function is calculated between the predicted label and the true label of the expanded data to determine the loss corresponding to the expanded sample in this iteration. Further, the calculation methods of the change in loss between the current and the previous iteration and the learning rate adjustment factor are expressed as:
[0117]
[0118] In the formula, is the loss of the t-th iteration of the second expanded sample; is the loss of the (t - 1)-th iteration of the second expanded sample; n c is the total number of training samples; is the gradient of the change in loss between the current and the previous iteration; δ c is the loss gradient adjustment factor. Preferably, δ c is set to 0.1.
[0119] Step S13: Perform data augmentation on the second expanded sample set to construct a third expanded sample set.
[0120] In order to improve the model's ability to recognize nonlinear relationships in feature space, the embodiment of the present invention also performs local device data enhancement based on the expanded device data, corrects the gradient, and strengthens the learning of edge samples. Since the single feature caused by relying solely on linear interpolation is avoided, the potential features of the device data can be fully explored. In specific implementation, the embodiment of the present invention performs data perturbation on the second expanded sample set to determine the gradient change of the second expanded sample set; based on the gradient change, the second expanded sample set is data enhanced to construct a third expanded sample set. The calculation method is expressed as:
[0121] X′ c3i =X′ c2i +α csw ·h c (X′ c2i )
[0122] In the formula, X′ c3i is the third expanded sample, representing the i-th sample enhanced according to the perturbation function; α csw h is the parameter to control the enhancement strength; c () is the perturbation function. Among them, α csw The larger the value, the more significant the change of the perturbation function on the sample characteristics, the higher the diversity of the generated data, but it may deviate from the true distribution. csw The smaller the value, the weaker the perturbation, and the generated data is closer to the original interpolation result, but the diversity may be insufficient. Preferably, α csw Set to 0.3. c2i As the interpolated sample, the dimension is the same as the above X c11 ,X c12 ,X c21 ,X c22 are the same, with the dimension [F]; h c (X′ c2i ) as the perturbation function, dimension is [F], and the gradient correction term The dimension of is also [F]; α csw As an enhancement strength, it is dimensionless, that is, the dimension is [1]; therefore, all terms in the added terms have the same dimension, X′ c2i and α csw ·hc(X′ c2i ) have the same dimension as [F] and can be added directly.
[0123] Furthermore, the perturbation function realizes device data enhancement through sample gradient, and the calculation method is expressed as:
[0124]
[0125]
[0126] In the formula, is the gradient of the second expanded sample after perturbation; is the gradient of the second expanded sample before disturbance; α ccr is the gradient scaling factor; β ccr is the adjustment parameter related to noise and gradient; θ c is the gradient direction angle; ∈ ccr is the random noise term; is the gradient of the perturbation function between the current and last iterations. Preferably, β ccr Set to 0.2, θ c Set to α ccr Set to 0.01.
[0127] Step S14, performing data correction on the third expanded sample set to construct a training sample set.
[0128] To ensure the validity of the expanded data, the embodiment of the present invention also compares the prediction error of the newly generated device data with the original device data for verification and correction. When the difference between the generated device data and the original device data is too large, it is corrected or eliminated to ensure the quality and reliability of the final expanded data, which is expressed as:
[0129] X′ c4i =X′ c3i -α ccd ·E c (X′ c3i )
[0130]
[0131] In the formula, X′ c4i is the fourth expanded sample, representing the i-th corrected sample; α ccd is the correction intensity; E c () is the prediction error function; f ccx () is a preset classification function, which is used to classify the samples after the third expansion to predict their labels; is the true label of the third expanded sample. Preferably, α ccd Set to 0.001. Among them, X′ c3i As the enhanced sample dimension and the aforementioned X c11 ,X c12 ,X c21 ,X c22 are the same, with the dimension [F]; E c (X′ c3i ) is used as the prediction error vector, and its dimension is consistent with that of the sample, which is [F]. For example, if the prediction error is 0.01, then the difference between each dimension in the enhanced sample and the prediction error is calculated; α ccdAs the correction intensity, it is dimensionless, i.e., the dimension is [1]; therefore, all the terms of each summand have the same dimension, and the dimension of each summand is [F], which is consistent with X'. c3i In summary, the embodiment of the present invention effectively solves the problem of scarce equipment lubrication state data through the method of combining local perturbation and random noise with linear interpolation. In addition, based on the adaptive learning rate, the weights of the augmented samples are dynamically optimized to improve the data utilization efficiency during the model training process. Further, the gradient enhancement and perturbation function methods are used to process the data perturbation while augmenting the data, avoiding the data being concentrated in a specific distribution area.
[0132] Further, in combination with the fourth augmented sample set, the distribution of the augmented data and the sample weight parameters can be continuously iteratively optimized. After generating and verifying the augmented data in each round, the model parameter weights of the above data augmentation unit are updated, so that the data augmentation and model training promote each other, gradually improving the representativeness of the augmented data and the adaptability of the model to different types of device data, which is expressed as:
[0133]
[0134] In the formula, is the sample weight of the t-th iteration; is the sample weight of the (t + 1)-th iteration; η cyg is the sample weight learning rate of the t-th iteration; is the gradient of the loss change between the current and the previous iteration. Further, the above steps are repeatedly iterated until the preset iteration stop condition is satisfied, which means the sample augmentation is completed. In one embodiment, the preset iteration stop condition is to reach the preset maximum number of iterations. Preferably, the preset maximum number of iterations is set to 100 times. Among them, As the weight, it is dimensionless, i.e., the dimension is [1]; As the loss gradient, it is the same as the loss L c and is dimensionless, i.e., the dimension is [1]; η cyg As the learning rate, it is dimensionless, i.e., the dimension is [1]; therefore, all the terms of each summand have the same dimension, and both have the dimension of [1] and can be directly added.
[0135] Further, through experimental verification, the experimental data and analysis results are referred to Figure 4 , by comparing the test accuracies of different data augmentation methods, the experiment selected conventional methods such as random noise augmentation, traditional bilinear interpolation, and SMOTE oversampling as benchmarks. The results show that the method of the present invention leads other methods with an accuracy of 92.7%, indicating the synergistic effect of local perturbation and bilinear interpolation, enhancing the diversity while maintaining the data authenticity. Refer to Figure 5, the training loss convergence curve shows that the method of the present invention can quickly approach the optimal solution in the initial stage of training, and automatically reduce the step size in the later stage to prevent oscillation. Refer to Figure 6 , through the comparison of the accuracy rates under different training sample sizes, it is verified that the method of the present invention can specifically fill the gaps in the feature distribution. However, the traditional method is limited by the fixed interpolation coefficient and is difficult to generate discriminative new samples in the sparse sample area, resulting in insufficient fitting ability of the model for the long-tail distribution. Refer to Figure 7 , through the noise robustness experiment, it is proved that the method of the present invention is reliable in a complex environment. When the noise intensity coefficient increases from 0 to 0.5, the accuracy rate of the method of the present invention only drops by 9.8%, while the traditional method drops by as much as 25.6%, indicating that the algorithm can effectively distinguish effective features from noise interference.
[0136] Furthermore, on the basis of the above embodiments, the embodiment of the present invention further provides a device lubrication poor fault diagnosis device based on artificial intelligence, Figure 8 shows a structural schematic diagram of a device lubrication poor fault diagnosis device based on artificial intelligence provided by the embodiment of the present invention. As Figure 8 shown, the device includes: a data acquisition module 100, configured to acquire operation monitoring data of a target device during operation, where the operation monitoring data includes various real-time data related to the lubrication state during the operation of the device, and the various real-time data are derived from real-time monitoring data and periodic record data of sensors in the industrial field; a data processing module 200, configured to perform feature domain perception on the correlation between the real-time data in the operation monitoring data based on the data characteristics of the operation monitoring data, and determine a feature domain perception result corresponding to the operation monitoring data; a feature extraction module 300, configured to perform data mapping on the operation monitoring data based on the feature domain perception result, and determine target features in the operation monitoring data; a classification module 400, configured to classify the probability of device faults for the target features, and determine a lubrication state fault diagnosis result of the target device; an output module 500, configured to determine the lubrication condition of the target device based on the lubrication state fault diagnosis result. For the device lubrication poor fault diagnosis device based on artificial intelligence provided by the embodiment of the present invention, its implementation principle and the technical effects generated are the same as those of the foregoing method embodiments. For the sake of brief description, for the parts not mentioned in the device embodiment, reference may be made to the corresponding content in the foregoing method embodiments.
[0137] Further, the above classification module 400 is further configured to determine a fault category score corresponding to the target feature; based on the fault category score, determine a category probability distribution corresponding to the target feature; and determine the lubrication state fault diagnosis result of the target device by taking the highest probability category in the category probability distribution. The above data processing module 200 is further configured to: calculate a perception weight corresponding to each feature according to the feature distance between each feature of the operation monitoring data; based on the perception weight, perform non-linear activation on the operation monitoring data to determine a feature domain perception result corresponding to the operation monitoring data. The above feature extraction module 300 is further configured to: use a pre-constructed feature extraction model to perform non-linear activation on the operation monitoring data based on the feature domain perception result, map the operation monitoring data to a low-dimensional space, and determine the target feature of the operation monitoring data. Among them, the above feature extraction module 300 is further configured to: obtain a pre-constructed training sample set, input the training sample set into an initial neural network, capture data of the training sample set through the neural network, and obtain a feature output; where the training sample set includes device operation samples related to the lubrication state during the operation of the device, and sample labels corresponding to the device operation samples; determine a local error corresponding to the feature output, and determine a divergence metric of the feature output; based on the divergence metric, calculate a loss function of the initial neural network; update the parameters of the initial neural network according to the loss function and the local error; until the initial neural network meets the preset training requirements, construct a feature extraction model based on the initial neural network. Further, the local error is determined by performing global information fusion on the feature output of each layer of the initial neural network; the above feature extraction module 300 is further configured to: determine a similarity metric result corresponding to the feature output, update the learning rate of the neural network based on the similarity metric result; based on the updated learning rate, calculate a weight update value of the neural network; update the parameters of the weight of the neural network according to the local error, the weight update value, and the loss function. The above feature extraction module 300 is further configured to: calculate a cross-entropy loss of the initial neural network according to the feature output; calculate an adaptive divergence regularization loss term of the initial neural network based on the divergence metric; calculate a loss function of the neural network according to the cross-entropy loss, the adaptive divergence regularization loss term, and a preset regularization weight.
[0138] The device further includes a construction module, which is configured to: obtain pre-collected device operation samples, perform data annotation on the device operation samples based on the device lubrication state corresponding to the device operation samples, and construct an initial sample set; perform preliminary sample expansion on the initial sample set according to the data distribution of the initial sample set to construct a first expanded sample set; determine target samples to be expanded from the first expanded sample set based on the sample dimensions of the first expanded sample set; perform linear interpolation processing on the target samples to be expanded based on a preset interpolation ratio coefficient to construct a second expanded sample set; perform data augmentation on the second expanded sample set to construct a third expanded sample set; perform data correction on the third expanded sample set to construct a training sample set. The above construction module is further configured to: perform data perturbation on the second expanded sample set to determine the gradient change of the second expanded sample set; perform data augmentation on the second expanded sample set based on the gradient change to construct a third expanded sample set. The above construction module is further configured to: determine the data gradient of the initial sample set, generate local perturbations of the initial sample set based on the data gradient; perform data expansion on the initial sample set based on the local perturbations and a preset random noise to construct a first expanded sample set.
[0139] An embodiment of the present invention further provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of any of the methods shown above are implemented. Figures 1 to 3 An embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is run by a processor, the steps of the method shown above are executed. Figures 1 to 3 An embodiment of the present invention further provides a schematic structural diagram of an electronic device. As shown in Figure 9 the figure, it is a schematic structural diagram of the electronic device. The electronic device includes a processor 91 and a memory 90. The memory 90 stores computer-executable instructions that can be executed by the processor 91. The processor 91 executes the computer-executable instructions to implement any of the methods shown above. Figures 1 to 3 In Figure 9In the illustrated embodiment, the electronic device further includes a bus 92 and a communication interface 93, wherein the processor 91, the communication interface 93, and the memory 90 are connected through the bus 92. The memory 90 may include a high-speed random access memory (RAM), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 93 (which may be wired or wireless), a communication connection is established between the system network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 92 may be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc., and may also be an AMBA (Advanced Microcontroller Bus Architecture) bus. Among them, AMBA defines three types of buses, including an APB (Advanced Peripheral Bus) bus, an AHB (Advanced High-performance Bus) bus, and an AXI (Advanced eXtensible Interface) bus. The bus 92 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 9It is represented only by a bidirectional arrow in the figure, but it does not mean that there is only one bus or one type of bus. The processor 91 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 91 or the instructions in the form of software. The above-mentioned processor 91 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor 91 reads the information in the memory and combines its hardware to complete the foregoing Figures 1 to 3 any of the illustrated methods.
[0140] A computer program product for a method and device for diagnosing equipment lubrication failure based on artificial intelligence provided by an embodiment of the present invention includes a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the method described in the foregoing method embodiments. For specific implementation, reference can be made to the method embodiments and will not be elaborated here. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working process of the above-described system can refer to the corresponding process in the foregoing method embodiments and will not be elaborated here. If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage media include: USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs, etc., which can store program code. In the description of the present invention, it should be noted that the terms "first", "second", and "third" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance.
[0141] Finally, it should be noted that the above embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, and are not intended to limit them. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify the technical solutions described in the foregoing embodiments, or easily conceive of changes, or equivalently replace some of the technical features within the technical scope disclosed by the present invention; and these modifications, changes, or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered by the protection scope of the present invention. Therefore, the protection scope of the present invention should be subject to the protection scope of the claims.
Claims
1. A method for diagnosing the fault of poor equipment lubrication based on artificial intelligence, characterized in that, The method includes: Obtaining operation monitoring data during the operation of a target device, where the operation monitoring data includes various real-time data related to the lubrication state during the device operation, and the various real-time data are from the real-time monitoring data and periodic record data of sensors in the industrial field; Based on the data characteristics of the operation monitoring data, performing feature domain perception on the correlation between the real-time data in the operation monitoring data to determine the feature domain perception result corresponding to the operation monitoring data; Based on the feature domain perception result, performing data mapping on the operation monitoring data to determine the target feature in the operation monitoring data; Performing equipment failure probability classification on the target feature to determine the lubrication state fault diagnosis result of the target device; Based on the lubrication state fault diagnosis result, determining the lubrication condition of the target device.
2. The method according to claim 1, characterized in that, The step of performing equipment failure probability classification on the target feature to determine the lubrication state fault diagnosis result of the target device includes: Determining the fault category score corresponding to the target feature; Based on the fault category score, determining the category probability distribution corresponding to the target feature; Determining the highest probability category in the category probability distribution as the lubrication state fault diagnosis result of the target device.
3. The method according to claim 1, characterized in that, The step of performing feature domain perception on the correlation between the real-time data in the operation monitoring data based on the data characteristics of the operation monitoring data to determine the feature domain perception result corresponding to the operation monitoring data includes: Calculating the perception weight corresponding to each feature according to the feature distance between each feature of the operation monitoring data; Based on the perception weight, performing non-linear activation on the operation monitoring data to determine the feature domain perception result corresponding to the operation monitoring data.
4. The method according to claim 1, wherein The step of performing data mapping on the operation monitoring data based on the feature domain perception result to determine the target feature in the operation monitoring data includes: Using a pre-constructed feature extraction model to perform non-linear activation on the operation monitoring data based on the feature domain perception result, mapping the operation monitoring data to a low-dimensional space, and determining the target feature of the operation monitoring data; Wherein, the construction method of the feature extraction model includes: Obtaining a pre-constructed training sample set, inputting the training sample set into an initial neural network, and performing data capture on the training sample set through the neural network to obtain a feature output; where the training sample set includes device operation samples related to the lubrication state during the device operation and the sample labels corresponding to the device operation samples; Determining the local error corresponding to the feature output and determining the divergence metric of the feature output; Based on the divergence metric, calculating the loss function of the initial neural network; According to the loss function and the local error, updating the parameters of the initial neural network; Until the initial neural network meets the preset training requirements, constructing a feature extraction model based on the initial neural network.
5. The method according to claim 4, characterized in that The local error is determined based on global information fusion of the feature output of each layer of the initial neural network; The step of updating the parameters of the neural network according to the loss function and the local error includes: Determine the similarity measurement result corresponding to the feature output, and update the learning rate of the neural network based on the similarity measurement result; Calculate the weight update value of the neural network based on the updated learning rate; Update the parameters of the weights of the neural network according to the local error, the weight update value, and the loss function.
6. The method according to claim 4, wherein The step of calculating the loss function of the initial neural network based on the divergence metric includes: Calculate the cross-entropy loss of the initial neural network according to the feature output; Calculate the adaptive divergence regularization loss term of the initial neural network based on the divergence metric; Calculate the loss function of the neural network according to the cross-entropy loss, the adaptive divergence regularization loss term, and a preset regularization weight.
7. The method according to claim 4, characterized in that The method for constructing the training sample set includes: Obtain the pre-collected device operation samples, perform data annotation on the device operation samples based on the device lubrication state corresponding to the device operation samples, and construct an initial sample set; Perform preliminary sample expansion on the initial sample set according to the data distribution of the initial sample set to construct a first expanded sample set; Based on the sample dimension of the first expanded sample set, determine the target sample points to be expanded from the first expanded sample set; perform linear interpolation processing on the target sample points to be expanded based on a preset interpolation ratio coefficient to construct a second expanded sample set; Perform data augmentation on the second expanded sample set to construct a third expanded sample set; Perform data correction on the third expanded sample set to construct a training sample set.
8. The method according to claim 7, characterized in that, The step of performing data augmentation on the second expanded sample set to construct a third expanded sample set includes: Perform data perturbation on the second expanded sample set to determine the gradient change of the second expanded sample set; Perform data augmentation on the second expanded sample set based on the gradient change to construct a third expanded sample set.
9. The method according to claim 7, characterized in that, The step of performing preliminary sample expansion on the initial sample set according to the data distribution of the initial sample set to construct a first expanded sample set includes: Determine the data gradient of the initial sample set, and generate local perturbations of the initial sample set based on the data gradient; Perform data expansion on the initial sample set based on the local perturbations and a preset random noise to construct a first expanded sample set.
10. A device for diagnosing lubrication malfunction of equipment based on artificial intelligence, characterized in that, The device includes: A data acquisition module for acquiring operation monitoring data during the operation of a target device. Among them, the operation monitoring data includes various real-time data related to the lubrication state during the operation of the device, and the various real-time data are from the real-time monitoring data and periodic record data of sensors in the industrial field; A data processing module for performing feature domain perception on the correlation between the real-time data in the operation monitoring data based on the data characteristics of the operation monitoring data, and determining the feature domain perception result corresponding to the operation monitoring data; A feature extraction module, configured to perform data mapping on the operation monitoring data based on the feature domain perception result, and determine target features in the operation monitoring data; A classification module, configured to classify the target features according to the equipment failure probability, and determine the lubrication state fault diagnosis result of the target equipment; An output module, configured to determine the lubrication condition of the target equipment based on the lubrication state fault diagnosis result.