Equipment fault prediction method and system based on deep learning
Through the cross-attention mechanism of deep learning, the multimodal data of the equipment is integrated, and the problem of insufficient integration of multimodal data in the prior art is solved, high-precision equipment failure prediction is achieved, and the stability and adaptability of the model are enhanced.
Patent Information
- Application Number
- CN202510830843.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing equipment fault prediction methods lack systematic solutions, making it difficult to effectively integrate multimodal data, resulting in incomplete feature expression, difficulty in processing high-dimensional features, and low efficiency in fusion between numerical and text, affecting the accuracy of fault prediction.
A deep learning-based method is adopted, and the cross-attention mechanism is used to deeply fusion of numerical data and text-attributed data. Numerical features are extracted through multi-layer perceptrons, text features are processed by self-attention, and cross-modal information fusion is carried out through cross-attention, and feature representation and model training are optimized.
It improves the accuracy and interpretability of equipment failure prediction, enhances the stability and adaptability of the model in the face of data fluctuations and uncertainties, and realizes high-precision fault prediction throughout the process.
Smart Images

Figure CN120337163A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly relates to a device fault prediction method and system based on deep learning. Background Technique
[0002] The statements in this part only provide background techniques related to the present invention and do not necessarily constitute prior art.
[0003] Device status data usually includes numerical sensor time-series data (such as temperature, vibration frequency, pressure) and text-based maintenance logs (such as fault descriptions, repair records). Facing complex data sources inside (device sensors, work order systems) and outside (environmental parameters, industry standards), how to efficiently fuse and extract valuable information to achieve more accurate fault prediction has become a major problem.
[0004] Existing fault prediction methods lack a systematic solution to integrate heterogeneous data, are prone to missing key details (such as the association between device historical repair records and sensor data mutations), and have problems such as lack of full-process integration, incomplete feature expression, difficulty in processing high-dimensional feature data, and low efficiency in fusing numerical and text data, making the accuracy of device fault prediction unable to meet the actual application requirements. Summary of the Invention
[0005] To solve the deficiencies of the existing technology, the present invention provides a device fault prediction method and system based on deep learning, optimizes multi-modal feature representation, uses a cross-attention mechanism to achieve deep fusion of numerical data and text data, and greatly improves the accuracy of device fault prediction.
[0006] To achieve the above object, the present invention adopts the following technical solutions: In the first aspect, the present invention provides a device fault prediction method based on deep learning.
[0007] A device fault prediction method based on deep learning includes the following processes: Obtain the numerical features and text feature vectors of the device to be predicted; The numerical features are used to extract high-order non-linear representations through a multi-layer perceptron to obtain encoded numerical features, and key vectors and value vectors of cross-attention are obtained according to the encoded numerical features; The text feature vectors are used to obtain encoded text vectors through an encoder, and query vectors of cross-attention are obtained by self-attention calculation of the encoded text vectors; According to the key vectors, the value vectors, and the query vectors, cross-attention is used for cross-modal information fusion, and the encoded numerical features are concatenated with the output of cross-attention to obtain fusion features; Based on the fusion features and the pre-trained deep learning network model, obtain the fault prediction result of the device to be predicted.
[0008] In an implementation manner of the first aspect of the present invention, the numerical features of the device to be predicted are derived from a numerical dataset, and the text feature vectors are derived from a text vector dataset. The construction of the numerical dataset and the text vector dataset includes: Obtain the monitoring data of the device to be predicted; Perform data screening on the monitoring data, and the data screening includes: removing invalid columns and removing invalid rows; The dataset after data screening includes numerical data and text data. Preprocess all the numerical data to obtain a numerical dataset, and preprocess all the text data to obtain a text vector dataset.
[0009] As a further limitation, preprocessing all the numerical data to obtain a numerical dataset includes: Use mean filling, median filling, linear interpolation, or model-based methods to fill in the missing values for all the numerical data; Use statistical methods or machine learning algorithms to identify and process outliers; Normalize or standardize the numerical data after outlier processing; Combine all the normalized or standardized numerical data to obtain a numerical dataset.
[0010] As a further limitation, preprocessing all the text data to obtain a text vector dataset includes: Remove invalid words, perform text processing through stemming or lemmatization, convert the result of text processing into a numerical vector representation, and combine all the numerical vector representations to obtain a text vector dataset.
[0011] As a further limitation, the numerical data includes at least: temperature, vibration frequency, and pressure sensor data; the text data includes at least: fault description, repair record, maintenance work order, technical parameters, component specifications, and historical fault cases.
[0012] In an implementation manner of the first aspect of the present invention, cross-attention is used to dynamically allocate weights to different numerical features according to the query vector. When a fault-sensitive word is identified according to the query vector, increase the weight of the numerical feature corresponding to the fault-sensitive word by a set threshold or a set percentage, and the sum of the weights of each numerical feature is 1.
[0013] In one implementation of the first aspect of the present invention, the reconstruction loss of the feature extraction network, the prediction loss of the deep learning network model, and the regularization intensity loss are used as the total loss function, and the feature extraction network and the deep learning network model are trained with the goal of minimizing the total loss function; Among them, the feature extraction network includes an encoder, a self-attention module, a multi-layer perceptron, and a cross-attention module. The feature extraction network is used to generate the fusion feature according to the numerical feature and the text feature vector.
[0014] In a second aspect, the present invention provides a device fault prediction system based on deep learning.
[0015] A device fault prediction system based on deep learning includes: A data acquisition unit configured to: acquire the numerical feature and the text feature vector of the device to be predicted; A numerical feature processing unit configured to: the numerical feature extracts a high-order non-linear representation through a multi-layer perceptron to obtain an encoded numerical feature, and obtains the key vector and the value vector of the cross-attention according to the encoded numerical feature; A text feature processing unit configured to: the text feature vector obtains an encoded text vector through an encoder, and the encoded text vector obtains a query vector of the cross-attention through self-attention calculation; A multi-dimensional feature fusion unit configured to: according to the key vector, the value vector, and the query vector, perform cross-modal information fusion by using cross-attention, and splice the encoded numerical feature with the output of the cross-attention to obtain a fusion feature; A prediction result generation unit configured to: obtain the fault prediction result of the device to be predicted according to the fusion feature and the pre-trained deep learning network model.
[0016] In a third aspect, the present invention provides a computer device, including: a processor and a computer-readable storage medium; The processor is adapted to execute a computer program; The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the device fault prediction method based on deep learning as described in the first aspect of the present invention.
[0017] In a fourth aspect, the present invention provides a computer-readable storage medium, which stores a computer program, and the computer program is adapted to be loaded and executed by a processor to implement the device fault prediction method based on deep learning as described in the first aspect of the present invention.
[0018] Compared with the prior art, the beneficial effects of the present invention are: 1. The present invention innovatively proposes a device fault prediction method based on deep learning, which optimizes multi-modal feature representation. While efficiently dealing with high-dimensional complex data, it deeply explores the correlation between features and retains the key information crucial for business decisions. By using the cross-attention mechanism, it realizes the deep interaction and effective fusion of numerical features and text features, enhances the stability and adaptability of the prediction scheme in the face of data fluctuations or uncertainties, and improves the accuracy of device fault prediction.
[0019] 2. The present invention innovatively proposes a device fault prediction method based on deep learning. Through the cross-attention mechanism, it realizes the deep interaction between numerical data and text data. After processing, the numerical data serves as the key vector and value vector, and the text data serves as the query vector. The cross-attention mechanism dynamically adjusts the weights of different numerical features according to the fault-sensitive words in the maintenance text data, so that the final fused features not only contain the physical state information of the numerical data, but also incorporate the context semantics in the text data, enabling more accurate identification of compound faults and significantly improving the accuracy and interpretability of fault prediction.
[0020] 3. The present invention innovatively proposes a device fault prediction method based on deep learning, which realizes the end-to-end processing from data acquisition, data preprocessing, multi-dimensional feature extraction to the training of the deep learning network model. First, it integrates the multi-source heterogeneous data of industrial equipment, conducts strict screening and cleaning, optimizes the feature representation and then performs feature fusion. Based on the fused features, it trains the deep learning network model, and finally realizes high-precision device fault prediction, ensuring the full-link reliability from the original data to the prediction result, and providing a complete solution for predictive maintenance in industrial scenarios.
[0021] 4. The present invention innovatively proposes a device fault prediction method based on deep learning. The present invention uses an autoencoder to reduce the dimension of high-dimensional text features, effectively compresses redundant information, and at the same time retains the key state features. It adopts the self-attention mechanism to dynamically allocate weights, captures the key entities in the text data, and automatically enhances the importance of the text features corresponding to the key entities. This not only improves the model's ability to understand complex device data, but also reduces the computational complexity through compact feature representation, providing technical support for real-time fault warning.
[0022] The advantages of the additional aspects of the present invention will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0024] Figure 1 A schematic flowchart of a device fault prediction method based on deep learning provided for an exemplary embodiment of the present invention; Figure 2 A schematic diagram of a feature extraction network provided for an exemplary embodiment of the present invention; Figure 3 A schematic diagram of a device fault prediction system based on deep learning provided for an exemplary embodiment of the present invention; Figure 4 A schematic diagram of a computer device provided for an exemplary embodiment of the present invention. Detailed implementation manners
[0025] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0026] It should be noted that the following detailed description is exemplary and is intended to provide further description of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0027] As described in the background art, currently, it is necessary to design a full-process method covering data acquisition, preprocessing, feature engineering to model training to improve the accuracy of fault prediction and retain details crucial for equipment health management, thereby improving the operation and maintenance efficiency and competitiveness of enterprises. The pain points and defects faced by existing device fault prediction technologies are mainly reflected in the following aspects: (1) Lack of a systematic full-process method. Existing technologies usually handle a single link in isolation (such as sensor data preprocessing or text classification), lacking the full-process integration from data acquisition, multi-modal preprocessing, feature engineering to model training and prediction; for example, ignoring the semantic value of maintenance logs in the data acquisition stage, or only relying on numerical data in the feature engineering stage, resulting in one-sided model input information; this fragmented processing makes it difficult to synergistically optimize data quality and model performance, ultimately affecting the reliability of fault prediction.
[0028] (2) Incomplete feature representation. Device data usually contains various types of features, such as time-series data collected by sensors, text information in maintenance logs, and parameter information related to device models. Traditional methods have obvious deficiencies in dealing with such multi-modal data: for numerical sensor data, they often only rely on simple statistical features (such as mean and variance), while ignoring the temporal dependence relationships contained therein, such as the periodic changes in vibration frequency; for text-based maintenance logs, shallow processing methods such as keyword extraction are often used, and the implicit fault patterns and their correlations are not deeply explored, such as the potential connection between "bearing wear" and "insufficient lubricating oil"; in addition, in terms of multi-modal feature fusion, existing methods are difficult to dynamically capture the correlation between sudden increase signals of sensors and descriptions such as "abnormal noise" in the logs, resulting in limited detection ability for complex faults and prone to missed detections.
[0029] (3) Difficulties in processing high-dimensional feature data. After the sensor data is segmented by time-series windows, the dimension surges (such as temperature, vibration, and pressure data collected per minute). Traditional dimensionality reduction methods (such as PCA) cannot effectively distinguish noise from key features (such as abnormal vibration peaks); in addition, the text vectors generated from maintenance logs by word embedding have redundant information (such as irrelevant operation records), further exacerbating the "curse of dimensionality", increasing the model calculation cost and reducing real-time performance.
[0030] (4) Low efficiency in numerical and text fusion. Existing methods have obvious inefficiencies in fusing sensor data and maintenance logs; firstly, the common practice is to simply concatenate numerical sensor features and text-based log vectors, lacking in-depth exploration of the semantic correlation between the two. For example, the strong correlation between high-temperature alarm logs and sudden increases in temperature sensors is often ignored; secondly, most methods adopt a static weight allocation strategy, fixing the fusion ratio of sensor data and text features, and it is difficult to dynamically adjust according to the actual input content. This inefficient fusion method results in low-quality feature representations generated finally, limiting the model's ability to capture complex fault patterns (such as the combined effect of mechanical wear and electrical overload), thus affecting the accuracy and reliability of fault diagnosis.
[0031] In view of the above problems, the present invention proposes a device fault prediction method based on deep learning, which realizes the full-process processing from data acquisition, preprocessing, feature extraction to model training and prediction. First, multi-source and multi-modal data such as real-time sensor data (temperature, vibration frequency, current, etc.) from the industrial Internet of Things platform, maintenance logs (fault descriptions, repair records) of the device management system, and environmental parameters (humidity, dust concentration) are integrated, and the data quality is ensured through strict screening and cleaning; then, a custom feature extraction network is used to optimize the data, aiming to efficiently handle high-dimensional complex data, deeply explore the correlation between features and retain key information crucial for business decisions. In particular, the present invention utilizes advanced text processing techniques and cross-attention mechanisms to achieve deep interaction and effective fusion of numerical features and text features, enhancing the stability and adaptability of the model in the face of data fluctuations or uncertainties.
[0032] More specifically, as Figure 1 shown, it includes the following processes: S101: Obtain the numerical feature and text feature vector of the device to be predicted; S102: The numerical feature extracts a high-order non-linear representation through a multi-layer perceptron to obtain an encoded numerical feature, and the key vector and value vector of cross-attention are obtained according to the encoded numerical feature; S103: The text feature vector (i.e., the original text vector) obtains an encoded text vector through an encoder, and the query vector of cross-attention is calculated through self-attention according to the encoded text vector; S104: According to the key vector, the value vector and the query vector, cross-attention is used for cross-modal information fusion, and the encoded numerical feature is concatenated with the output of cross-attention to obtain a fused feature; S105: According to the fused feature and the pre-trained deep learning network model, the fault prediction result of the device to be predicted is obtained.
[0033] Before step S101 of this implementation manner, it also includes the processes of data acquisition, data set construction, data screening and data classification preprocessing. Specifically, it includes: To achieve device health management and fault prediction, data needs to be collected from multi-source heterogeneous channels. This includes real-time sensor data (such as vibration, temperature, pressure, etc.) provided by enterprise internal systems like industrial Internet of Things platforms, maintenance work orders and repair records in the device management system; technical parameters, component specifications, and historical fault cases provided by device manufacturers; health standards, fault mode libraries, and environmental monitoring data in industry databases; as well as text information such as operation records, manual reports, and third-party inspection reports. At the same time, external environmental parameters and production scheduling data are collected. The key data types of this invention cover time series signals, device static information, text descriptions, environmental working condition data, and industry reference information. The data needs to meet accuracy, integrity, timeliness, and reliable sources to cover the learning needs of the entire device life cycle and complex fault modes.
[0034] In terms of constructing the dataset, the above various types of data are integrated into a dataset in matrix form , where each row represents an observation value, each column represents different features, and the last column represents the actual device health status. Suppose there are observation values and features, then the dataset can be expressed as: (1); Among them, represents the value of the th feature of the th observation value, corresponds to the device health status of each observation value, represents the real number field. It should be noted that for text-type data, through feature engineering, it needs to be converted into numerical-type features for matrix expression; for example, methods such as the bag-of-words model, TF-IDF, or Word2Vec are used to convert the text into numerical-type vector representations. For example, the product description field is converted into a 300-dimensional vector, and the customer review is converted into a 300-dimensional vector. After this step, the text-type data is also converted into numerical-type data.
[0035] In this invention, the obtained dataset is screened for data. Specifically, it includes: ① Remove invalid columns: Invalid columns refer to those feature columns that contribute nothing to model training or have serious missing values. First, for each column j, calculate its missing value ratio: (2); Among them, represents the counting operation, represents the missing value. Then, define a missing value ratio threshold , usually with a value range of . If the missing value ratio of a certain column , then this column is marked as an invalid column. Remove all columns that satisfy from the dataset to obtain a new dataset , where is the number of valid columns. The mathematical expression is: (3); ② Remove invalid rows: Invalid rows refer to those sample rows that contain too many missing values or outliers. For each row i, calculate its missing value ratio: (4); Define a missing value ratio threshold , usually with a value range of . If the missing value ratio of a certain column , then this column is marked as an invalid column. Remove all columns that satisfy from the dataset to obtain a new dataset , where is the number of valid columns, and the mathematical expression is: (5); Among them, represents the number of valid samples, and represents the number of valid features.
[0036] Perform necessary data preprocessing on the dataset after data screening to provide reliable high-quality data for subsequent feature engineering and model training effects. On the dataset , it can be divided into numerical data and text data , where , represents the number of numerical data, represents the number of text data. Different preprocessing strategies can be adopted for different types of data.
[0037] ① Process numerical data: For numerical data, mainly focus on missing value handling, outlier detection and handling, and standardization or normalization. First, use mean filling, median filling, linear interpolation, or model-based methods to fill in missing values; secondly, use statistical methods (such as Z-score or IQR) or machine learning algorithms (such as Isolation Forest) to identify and handle outliers; in order to ensure that data of different scales can be treated fairly by the model, it is usually necessary to standardize or normalize numerical data. After these steps, a preliminary processed numerical dataset will be obtained.
[0038] ② Process text data: The preprocessing of text data mainly includes steps such as removing stop words, stemming or lemmatization, and vectorization. First, remove those high-frequency words that are not helpful for analysis. Then, reduce the variant forms of words through stemming or lemmatization, and finally form a text vector dataset , where represents the dimension of the text vector.
[0039] In step S101 of the present invention, the input of the feature extraction network consists of numerical features and the original text vector (i.e., the text feature vector), and the numerical features are derived from the numerical dataset , and the original text vector is derived from the text vector dataset .
[0040] As Figure 2 shown, the processes of steps S102, S103, and S104 in this implementation manner are the feature extraction processes of the feature extraction network of the present invention, including the following processes: The goal of the feature extraction network is to construct a more expressive feature set through in-depth mining and fusion of existing features, so as to provide high-quality input for the subsequent equipment fault prediction model. In the above-obtained text vector dataset contains a large amount of high-dimensional sparse data. These high-dimensional vectors often have a lot of redundant information or noise, which will interfere with the learning process of the model; at the same time, the obtained numerical data needs to be fused with the text vector to achieve the complementarity of multi-source information, capture the potential laws that cannot be reflected by single-modal data, and finally obtain high-quality features and input them into the subsequent equipment fault prediction model (i.e., the deep learning network model).
[0041] More specifically, the feature extraction network of the present invention can not only extract deeper text vector feature representations through non-linear transformation and dimensionality reduction techniques, but also efficiently integrate numerical features and text vectors. The overall structure of the feature extraction network is as Figure 2 shown, mainly including modules such as an encoder 203, a decoder 204, a multi-layer perceptron 201, and a cross-attention 202.
[0042] The encoder 203 is a neural network, which is set as a fully connected layer here (the specific number of layers depends on the data and business situation): (6); where is a set of original text vectors, represents the weight matrix of the encoder 203, is a number much smaller than d, represents the bias vector of the encoder 203, Denote activation functions (such as ReLU, Sigmoid), Denote the encoded text vector.
[0043] ② The decoder 204 is similar to the encoder 203 and is also a neural network (which is only used in the training stage for reconstructing the text vector). Here, it is set as a fully connected layer (the specific number of layers depends on the data and business situation): (7); Among them, Denote the encoded text vector, Denote the weight matrix of the decoder 204, Denote the bias vector of the decoder 204, Denote a set of reconstructed text vectors.
[0044] For the encoder 203 and the decoder 204, during the training process, the mean square error or cross - entropy is used as the loss function to optimize the network parameters. Here, the mean square error (not limited to this function) is selected as the reconstruction loss function: (8); ③ The multi - layer perceptron 201 consists of fully connected layers, including the input layer, hidden layers and the output layer. Among them, the layer has neurons. Let Denote the output vector of the layer. Then the calculation result from the layer to the layer can be expressed as: (9); Among them, Denote the weight matrix from the layer to the layer, is the bias vector, Denote the activation function. The size of and the number of neurons in each layer
[0045] ④ In the cross - attention 202, according to the encoded numerical features Generate and matrices: (10); (11); Among them, , is the dimension of the attention space, is a learnable weight matrix, is the bias term.
[0046] Then, the output of the self-attention 205 is used as , calculate the attention scores and output the results: (12); where, is the output of the cross-attention 202, function normalizes the similarity to a probability distribution.
[0047] The numerical features extract high-order non-linear representations through the multi-layer perceptron 201 to obtain the encoded numerical features; meanwhile, the text feature vector realizes feature dimensionality reduction through the encoder 203 to remove redundant information and noise, obtaining the encoded text vector; the encoded text vector uses self-attention to capture the dependencies between text fields and extract global context information, and its output is used as the of the cross-attention 202; in addition, the and of the cross-attention 202 are constructed based on the encoded numerical features, and subsequently, the cross-attention 202 is used to achieve dynamic weight allocation and cross-modal information fusion; finally, the encoded numerical features are concatenated with the output of the cross-attention 202 (implemented through the concatenation module 206) to comprehensively utilize the original numerical features and the features enhanced by the cross-attention mechanism to generate more accurate and rich feature representations.
[0048] In step S105 of this implementation manner, during the construction stage of the deep learning network model, an ensemble learning algorithm is used to build and train the deep learning network model. Here, the Boosting model is selected as an example, but it is not limited to this model. The following is the complete training process of the deep learning network model and its related mathematical expressions.
[0049] The goal of the Boosting algorithm is to minimize the loss function while introducing a regularization term to prevent overfitting: (13); where, is the loss function, such as mean squared error (MSE) or cross-entropy loss, is the regularization term used to control the model complexity, and the formula is: (14); where, is the number of trees, is the leaf node weight, and are the regularization parameters.
[0050] To further improve the model performance, the custom multi-modal feature extractor in the feature extraction stage is combined with the deep learning network model to form an overall device fault prediction model. The loss function for the entire process (feature extraction process and fault prediction process) is designed as follows: (15); where is the weight for controlling the reconstruction loss of the feature extraction network, is the weight for the prediction loss of the deep learning network model, is the weight for controlling the regularization strength, represents the weight parameters of the overall network model, represents the set of weight parameters.
[0051] In the end-to-end training process, 80% of the preprocessed dataset is used for model training, and the hyperparameters of the model are tuned through cross-validation. To comprehensively measure the model performance, mean squared error (MSE), mean absolute error (MAE), and coefficient of determination ( ) evaluation metrics are used to comprehensively measure the model performance.
[0052] After completing the training and hyperparameter tuning of the custom multi-modal feature extraction network and the deep learning network model, the remaining 20% of the preprocessed dataset is used as the test set to evaluate the prediction ability of the entire system. The specific steps are as follows: ① Prepare the test set: Divide 20% of the data from the preprocessed complete dataset as an independent test set. This part of the data has not been used during the entire training process, ensuring the authenticity and objectivity of the evaluation results; ② Model prediction: Use the trained custom multi-modal feature extractor and deep learning network model to predict the data in the test set. This step simulates the model performance in the actual application scenario and provides an accurate reference for device fault prediction; ③ Performance evaluation: To comprehensively measure the prediction ability of the model, mean squared error (MSE), mean absolute error (MAE), and coefficient of determination ( ) evaluation metrics are used to comprehensively measure the model performance.
[0053] In summary, the device fault prediction method based on deep learning provided by the present invention has the following advantages: (1) Systematic process coverage: The present invention realizes the full process of equipment sensor data collection, multi-modal data preprocessing, feature extraction to model training and prediction. First, the multi-source heterogeneous data of industrial equipment (such as sensor time series data such as temperature and vibration frequency, as well as text information such as fault descriptions and maintenance records in maintenance logs) are integrated, and after strict screening and cleaning (such as removing sensor noise and repairing missing values), the feature representation is optimized using a custom feature extraction network. Finally, a deep learning network model is trained based on the fusion features to achieve high-precision equipment fault prediction. This method ensures the reliability of the entire link from raw data to prediction results, and provides a complete solution for predictive maintenance in industrial scenarios.
[0054] (2) Refined multimodal data processing: The present invention uses an autoencoder to perform feature dimensionality reduction on the text feature vectors extracted from the unstructured text in the maintenance log, effectively compressing redundant information while retaining key semantic features. Subsequently, the self-attention mechanism is used to perform global association modeling on the encoded text vectors corresponding to different texts, and the weight distribution of each part of the text features is dynamically adjusted to accurately capture key fault entities (such as "bearing wear" and "insufficient lubricant") and their contextual dependencies. This mechanism enables the model to automatically focus on the most discriminative semantic units according to different text contents, and enhance the importance of sensor features with the highest semantic match with the current semantics during the multimodal fusion process. For example, when "abnormal noise" appears in the maintenance log, the model will automatically increase the attention weight of the vibration sensor-related features to achieve semantic alignment and linkage analysis between text and equipment status. Through the effective modeling of the self-attention mechanism, not only the model's ability to understand and express diverse and heterogeneous text information is enhanced, but also the accuracy and interpretability of multimodal feature fusion are improved.
[0055] (3) Efficient numerical and text fusion method: This invention realizes deep interaction between sensor time series data (i.e. numerical data) and maintenance log text (i.e. text data) through a cross-attention mechanism. Specifically, by processing sensor time series data as key vectors and value vectors, and by processing maintenance log text as query vectors, the semantic information in the text data (such as "high temperature alarm") dynamically adjusts the weight information of different sensor features. , The obtained weight matrix is Each element in is assigned a different weight. The larger the weight, the better the corresponding The greater the contribution of the elements in [it] to the final attention output. For example, when "cooling system failure" appears in the text data, the present invention will focus on the abnormal fluctuations of the temperature sensor, rather than evenly distributing the weights of all sensors. The final fused features not only contain the physical state information of the original sensor data, but also incorporate the context semantics in the maintenance records, enabling the deep learning network model to more accurately identify complex faults (such as the combined effect of mechanical wear and electrical overload), significantly improving the accuracy and interpretability of fault prediction.
[0056] The above content elaborates in detail the method of the embodiment of the present invention. To facilitate better implementation of the above method of the embodiment of the present invention, correspondingly, the system of the embodiment of the present invention is provided below.
[0057] Figure 3 There is shown a device fault prediction system based on deep learning, including: A data acquisition unit 301, configured to: acquire the numerical features and text feature vectors of the device to be predicted; A numerical feature processing unit 302, configured to: extract high-order non-linear representations of the numerical features through a multi-layer perceptron to obtain encoded numerical features, and obtain the key vector and value vector of cross-attention according to the encoded numerical features; A text feature processing unit 303, configured to: obtain an encoded text vector from the text feature vector through an encoder, and obtain the query vector of cross-attention by calculating the encoded text vector through self-attention; A multi-dimensional feature fusion unit 304, configured to: perform cross-modal information fusion using cross-attention according to the key vector, the value vector, and the query vector, and splice the encoded numerical features with the output of cross-attention to obtain fused features; A prediction result generation unit 305, configured to: obtain the fault prediction result of the device to be predicted according to the fused features and a pre-trained deep learning network model.
[0058] It can be understood that the above-mentioned respective units can be separately or all combined into one or several other units to form, or some of them can be further split into multiple smaller units with functional division to form, which can achieve the same operation without affecting the realization of the technical effects of the embodiments of the present application. The above units are divided based on logical functions. In practical applications, the function of one unit can also be realized by multiple units, or the functions of multiple units can be realized by one unit. In other embodiments of the present application, the system may also include other units. In practical applications, these functions can also be assisted by other units and can be realized by the cooperation of multiple units.
[0059] According to another embodiment of the present application, the system described in this embodiment can be constructed by running a computer program (including program code) capable of executing the steps involved in the present invention on a general computing device such as a computer including processing elements and storage elements such as a Central Processing Unit (CPU), a Random Access Memory (RAM), and a Read-Only Memory (ROM). The computer program can be recorded on, for example, a computer-readable recording medium, loaded into the above computing device through the computer-readable recording medium, and run therein.
[0060] Figure 4 A computer device is shown. The electronic device includes a processor 401, a communication interface 402, and a computer-readable storage medium 403. Among them, the processor 401, the communication interface 402, and the computer-readable storage medium 403 can be connected through a bus or other means.
[0061] Among them, the communication interface 402 is used to receive and send data. The computer-readable storage medium 403 can be stored in the memory of the electronic device. The computer-readable storage medium 403 is used to store a computer program. The computer program includes program instructions. The processor 401 is used to execute the program instructions stored in the computer-readable storage medium 403.
[0062] The processor 401 is the computing core and control core of the electronic device, and is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions to implement the corresponding method flow or corresponding function.
[0063] The processor 401 is configured to execute the following process: Obtain the numerical features and text feature vectors of the device to be predicted; The numerical features are used to extract high-order non-linear representations through a multi-layer perceptron to obtain encoded numerical features, and key vectors and value vectors of cross-attention are obtained according to the encoded numerical features; The text feature vectors are used to obtain encoded text vectors through an encoder, and query vectors of cross-attention are obtained by calculating the encoded text vectors through self-attention; According to the key vectors, the value vectors, and the query vectors, cross-attention is used for cross-modal information fusion, and the encoded numerical features are concatenated with the output of the cross-attention to obtain fusion features; According to the fusion features and a pre-trained deep learning network model, a fault prediction result of the device to be predicted is obtained.
[0064] The present invention also provides a computer-readable storage medium. The computer-readable storage medium is a memory device in an electronic device and is used to store programs and data. It can be understood that the computer-readable storage medium here can include both the built-in storage medium in the electronic device and, of course, the extended storage medium supported by the electronic device. The computer-readable storage medium provides a storage space, and this storage space stores the processing system of the electronic device.
[0065] Moreover, in this storage space, one or more instructions suitable for being loaded and executed by a processor are also stored. These instructions can be one or more computer programs (including program codes). It should be noted that the computer-readable storage medium here can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory; optionally, it can also be at least one computer-readable storage medium located far from the aforementioned processor.
[0066] In one embodiment, one or more instructions are stored in the computer-readable storage medium; the one or more instructions stored in the computer-readable storage medium are loaded and executed by the processor to implement the following process: Obtain the numerical features and text feature vectors of the device to be predicted; The numerical features extract high-order non-linear representations through a multi-layer perceptron to obtain encoded numerical features, and based on the encoded numerical features, the key vector and value vector of cross-attention are obtained; The text feature vectors are encoded into encoded text vectors through an encoder, and the encoded text vectors are used to calculate the query vector of cross-attention through self-attention; According to the key vector, the value vector, and the query vector, cross-attention is used for cross-modal information fusion, and the encoded numerical features are concatenated with the output of cross-attention to obtain fused features; According to the fused features and a pre-trained deep learning network model, a fault prediction result of the device to be predicted is obtained.
[0067] The present invention also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and these computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the electronic device to perform the following process: Obtain the numerical features and text feature vectors of the device to be predicted; The numerical features extract high-order non-linear representations through a multi-layer perceptron to obtain encoded numerical features, and based on the encoded numerical features, the key vector and value vector of cross-attention are obtained; The text feature vector is encoded into an encoded text vector through an encoder, and the encoded text vector is used to calculate the query vector of cross-attention through self-attention; According to the key vector, the value vector and the query vector, cross-attention is used for cross-modal information fusion, and the encoded numerical features are concatenated with the output of the cross-attention to obtain fused features; According to the fused features and a pre-trained deep learning network model, a fault prediction result of the device to be predicted is obtained.
[0068] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this application can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Skilled artisans can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0069] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions according to the embodiments of this application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted through the computer-readable storage medium. The computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital line) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data processing device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive), etc.
[0070] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A device fault prediction method based on deep learning, characterized in that It includes the following processes: Obtain the numerical features and text feature vectors of the device to be predicted; The numerical features extract high-order non-linear representations through a multi-layer perceptron to obtain encoded numerical features, and obtain the key vector and value vector of cross-attention based on the encoded numerical features; The text feature vectors obtain encoded text vectors through an encoder, and the encoded text vectors obtain the query vector of cross-attention through self-attention calculation; According to the key vector, the value vector and the query vector, adopt cross-attention for cross-modal information fusion, and splice the encoded numerical features with the output of cross-attention to obtain fusion features; According to the fusion features and the pre-trained deep learning network model, obtain the fault prediction result of the device to be predicted.
2. The deep learning-based device fault prediction method according to claim 1, wherein The numerical features of the device to be predicted are derived from a numerical dataset, and the text feature vectors are derived from a text vector dataset. The construction of the numerical dataset and the text vector dataset includes: Obtain the monitoring data of the device to be predicted; Perform data screening on the monitoring data, and the data screening includes: removing invalid columns and removing invalid rows; The dataset after data screening includes numerical data and text data; Preprocess all numerical data to obtain a numerical dataset; Preprocess all text data to obtain a text vector dataset.
3. The deep learning-based device fault prediction method according to claim 2, wherein Preprocessing all numerical data to obtain a numerical dataset includes: Use mean filling, median filling, linear interpolation or model-based methods to fill in missing values for all numerical data; Use statistical methods or machine learning algorithms to identify and process outliers; Normalize or standardize the numerical data after outlier processing; Combine all the normalized or standardized numerical data to obtain a numerical dataset.
4. The deep learning-based device fault prediction method according to claim 2, wherein Preprocessing all text data to obtain a text vector dataset includes: Remove invalid words, perform text processing through stemming or lemmatization, convert the results of text processing into numerical vector representations, and combine all numerical vector representations to obtain a text vector dataset.
5. The deep learning-based device fault prediction method according to any one of claims 2-4, wherein The numerical data at least includes: temperature, vibration frequency, and pressure sensor data; The text data at least includes: fault descriptions, repair records, maintenance work orders, technical parameters, component specifications, and historical fault cases.
6. The deep learning-based device fault prediction method according to any one of claims 1-4, wherein Cross-attention is used to dynamically assign weights to different numerical features according to the query vector. When a fault-sensitive word is identified according to the query vector, increase the weight of the numerical feature corresponding to the fault-sensitive word by a set threshold or a set percentage, and the sum of the weights of each numerical feature is 1.
7. The method for predicting device faults based on deep learning according to any one of claims 1-4, characterized in that the reconstruction loss of the feature extraction network, the prediction loss of the deep learning network model, and the regularization intensity loss are used as the total loss function, and the feature extraction network and the deep learning network model are trained with the goal of minimizing the total loss function; wherein, the feature extraction network includes an encoder, a self-attention module, a multi-layer perceptron, and a cross-attention module, and the feature extraction network is used to generate the fusion feature according to the numerical feature and the text feature vector.
8. A device fault prediction system based on deep learning, characterized in that, It includes: a data acquisition unit, configured to: acquire the numerical feature and the text feature vector of the device to be predicted; a numerical feature processing unit, configured to: the numerical feature extracts a high-order non-linear representation through a multi-layer perceptron to obtain an encoded numerical feature, and obtains the key vector and value vector of cross-attention according to the encoded numerical feature; a text feature processing unit, configured to: the text feature vector obtains an encoded text vector through an encoder, and the encoded text vector obtains a query vector of cross-attention through self-attention calculation; a multi-dimensional feature fusion unit, configured to: according to the key vector, the value vector, and the query vector, perform cross-modal information fusion using cross-attention, and splice the encoded numerical feature with the output of cross-attention to obtain a fusion feature; a prediction result generation unit, configured to: obtain the fault prediction result of the device to be predicted according to the fusion feature and the pre-trained deep learning network model.
9. A computer device, characterized in that, It includes: a processor and a computer-readable storage medium; the processor is adapted to execute a computer program; the computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the method for predicting device faults based on deep learning according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is adapted to be loaded and executed by the processor to implement the method for predicting device faults based on deep learning according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method for predicting drug target associativity based on combined cross-domain attention model
CN116646001A
Attention mechanism-based electric energy prediction method and system for multiple fields
CN116738212A
Visual language cross-modal learning method for structural health diagnosis large model
CN117253112A
Deep fusion network production line fault prediction method based on deep learning
CN119357769A
VF value multi-modal prediction method and system based on deep learning
CN119851332A
Cited By
Equipment risk prediction method and device based on feature selection
CN120579070A
Intelligent fault analysis method and device for indoor and outdoor equipment of full-electronic interlocking system, electronic equipment and storage medium
CN120951081A
Humanoid robot decision interaction method, system and equipment and medium
CN120974368A
Behavior labeling method based on end-to-end humanoid robot large model and related equipment
CN121705817A