An elevator risk warning method based on local patch and enhanced position coding
By using local patch and enhanced position coding methods in elevator fault prediction, and using the Transformer model for dual attention modeling, the problem of difficult to explore the intrinsic connections of time series and sparse fault records in the existing technology is solved, and more accurate elevator fault prediction and maintenance requirements extraction are achieved.
Patent Information
- Application Number
- CN202510065942.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-01-16
AI Technical Summary
The existing elevator fault prediction algorithms are difficult to effectively explore the intrinsic connections of the time series, and due to sparse fault records and trapped records, it is difficult to accurately judge the elevator fault status, which makes it difficult for the elevator maintenance system to adapt to the diversity of usage conditions and the uncertainty of faults.
The elevator risk warning method based on local patch and enhanced position coding is adopted. By collecting IoT elevator data, multivariate timing characteristics are generated, and preprocessing and patch operations are performed. The Transformer model is used for dual attention modeling, combining position coding and self-attention mechanism to extract the fault type and fault probability distribution.
Through local patch and enhanced position coding technology, the model can more effectively capture local and global patterns in the time series, improve the accuracy and practicality of elevator fault prediction, and more accurately extract the elevator maintenance needs.
Smart Images

Figure CN119503573B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of elevator fault warning, and relates to an elevator risk warning method based on local patch and enhanced position coding. Background Art
[0002] Most of the current elevator fault prediction algorithms are based on the elevator's temperature, humidity, running speed, maintenance interval and other information to model the elevator's operating status and predict faults. The prediction methods used are the time series prediction algorithm LSTM and the machine learning method random forest algorithm. These methods do not fully explore the intrinsic connections of time series. The original data studied in this project include various types of information such as the elevator's manufacturing unit, equipment model, performance parameters, geographic location, fault records, trapped records, weather and temperature. It is difficult to model all input data with a single structure; fault records and trapped records are event-type information, which are relatively sparse in time series, and the information used to extract features is relatively insufficient; the internal connections between various types of data are weak, and it is difficult to use a single structure to model and predict the probability of elevator failure.
[0003] With the development of artificial intelligence and big data technologies, artificial intelligence and big data technologies have gradually been reflected in the elevator industry. After the integration of elevators with artificial intelligence and big data technologies, the safety and reliability of elevators have been further improved. Elevators need to be maintained and inspected regularly during use to avoid elevator accidents. However, traditional elevator maintenance mostly relies on manual regular inspection and fixed-cycle maintenance. Although it guarantees the safe operation of elevators to a certain extent, it still has the following disadvantages: 1. The maintenance standards of elevators are different, and the professional quality of maintenance personnel is uneven. It cannot be ensured that all maintenance personnel perform maintenance on elevators correctly in accordance with regulations, and the maintenance needs of elevators are uneven. Maintenance personnel can only perform maintenance on elevators in accordance with the prescribed maintenance standards, which greatly increases the workload of maintenance personnel and reduces maintenance efficiency; 2. In some remote elevator maintenance methods, static data of the elevator is obtained when the elevator fails, and these static data cannot reflect all the fault problems of the elevator because they cannot accurately judge the fault status of the elevator. At the same time, when the elevator fails, the sample data generated is small, and the data obtained is uneven, which leads to the problem of unstable data. Therefore, the current elevator maintenance system and early warning prediction system are difficult to adapt to the diversity of elevator usage conditions and the uncertainty of failures. Summary of the invention
[0004] In order to overcome the deficiencies of the prior art, the present invention provides an elevator risk warning method based on local patch and enhanced position coding.
[0005] In order to achieve the above object, the present invention adopts the following technical solutions:
[0006] An elevator risk warning method based on local patch and enhanced position coding includes the following steps:
[0007] S1: Collect IoT elevator data as raw data for risk warning and store it in the database;
[0008] S2: Clean the collected raw data to generate multivariate time series features, which include dense features, sparse features and static features;
[0009] S3: Preprocess and patch the generated multivariate time series features. Use the patch module to segment the input time series data. Segment the input time series according to a certain window size and step size, and use the segmented patch blocks as token input.
[0010] S4: The input token is position-encoded through the relative position encoding layer to obtain the comprehensive vector feature;
[0011] S5: The comprehensive vector features are transmitted to the dual attention module according to the input dimension for dual attention modeling. The dual attention module includes intra-patch attention and inter-patch attention. The softmax function is used to normalize the input time series features based on the calculated weights. According to the processing results of the self-attention mechanism, the comprehensive vector features are extracted and crossed through convolution pooling operations.
[0012] S6: Reduce the dimension of the output vector calculated by the dual attention mechanism, output the corresponding fault type, and convert the dimension reduction result through the softmax function to obtain the fault probability distribution;
[0013] S7: Train the multivariate Transformer model, adjust the parameters of the multivariate Transformer model, use the loss function and Adam optimizer, evaluate the model performance in combination with the validation set, and adjust the model structure or training process based on the evaluation results;
[0014] S8: A customized multi-layer convolutional block is used to post-process the data of predicted fault types and fault probability distribution to extract specific maintenance requirements.
[0015] Furthermore, in step S3, an embedding layer is used to preprocess the multivariate time series features, and the static features and dense features are converted into dense features through the embedding layer. The converted dense features and the dense features generated after data cleaning are extracted and spliced. The conversion formula is:
[0016] x1=embedding(Xsparse)
[0017] x2=embedding(Xstatic)
[0018] x3=WXdense+b
[0019] F=cat(x1,x2,x3)
[0020] Among them, x1 is the dense feature vector generated after the static feature is converted, x2 is the vector of sparse features converted to dense features, x3 is the vector of dense feature extraction, Xdense is the dense feature, Xsparse and Xstatic are the one-hot encoded features that need to be converted, W is the fully connected layer, F is the obtained comprehensive vector, b is the bias parameter, Xdense extracts features through the fully connected layer W and performs cat splicing operation.
[0021] Furthermore, in step S3, the patch operation uses the following formula:
[0022]
[0023] Among them, L is the length of the original time series, S is the length of each patch block, and P is the number of patch blocks. This is a rounding-down operation, which means that only the number of patches that can be fully adapted is considered.
[0024] Furthermore, in step S4, the position coding includes geometric position coding, and the specific steps of performing geometric position coding include:
[0025] S4.1: define a position encoding function, wherein the position encoding function generates a unique position code according to the position of the element in the time series;
[0026] S4.2: Generate the position code corresponding to each time series position through the position code function;
[0027] S4.3: Combine the generated position encoding with the feature vector of the corresponding position in the time series.
[0028] Furthermore, in step S4, the position encoding also includes semantic position encoding, which is injected into the attention map, and the L2 regularization term is used to constrain the difference between the similarity matrix of each layer of the attention map and the initial comprehensive vector feature as the semantic PE, using the following formula:
[0029]
[0030] in, is the temporal semantic regularization term, is the number of attention layers in the model, is the L2 distance, is the softmax operation, Indicates Attention map in the layer attention layer, is the similarity matrix of the token of each patch block of the single variable at the initial time, is the query vector for each token, A key vector representing each token, is the vector in the similarity matrix of the token of each patch block, and T represents the symbol of the transposition operation. Express Transpose, Express Transpose, D represents the dimension of query vector, key vector and value vector.
[0031] Furthermore, in step S5, the specific steps of dual attention modeling include:
[0032] S5.1: The input time series is a matrix ,in is the length of the original time series, As feature dimension, the matrix Divide into multiple patch blocks. Assume that the length of each patch block is S, and get the matrix of each patch block after division. ;
[0033] S5.2: Perform feature embedding for each patch block to obtain the feature matrix representing the i-th patch block after embedding into the internal time series ,in The query vector, key vector, and value vector are generated through the time step feature vector in each patch block, and the internal attention output of each patch block is obtained by combining the softmax operation. , concatenate the internal attention output of each patch block , generate the attention output matrix within the patch , where P is the number of patch blocks;
[0034] S5.3: Convert the matrix of the original time series Combine the length S and the feature dimension The feature embedding operation is performed to obtain the embedded matrix ,in, , calculate the attention weights between patches and obtain the attention output matrix between patches ;
[0035] S5.4: By reordering the attention outputs within the patch and performing a linear transformation on the length dimension of the patch block, we obtain the matrix , merge the time steps in each patch block and output them with the attention between patches Add together to get the dual attention output .
[0036] Furthermore, in step S5.2, the internal attention output of each patch block is The calculation formula is:
[0037]
[0038] in, are the query vector, key vector, and value vector of the time step in each patch block, respectively. T represents the symbol of the transpose operation. Express Transpose, is the feature dimension after embedding.
[0039] Furthermore, in step S5.3, the formula for calculating the attention weight between patches is:
[0040]
[0041] in, , , They are query matrix, key matrix and value matrix respectively. T represents the symbol of transposition operation. Express Transpose, .
[0042] Furthermore, the original data of the risk warning in step S1 includes: basic elevator information, fault information, weather temperature information, and maintenance information.
[0043] Furthermore, in step S8, the customized multi-layer convolution block includes: two convolution layers and one fully connected layer, and the steps of extracting specific maintenance requirements by the multi-layer convolution block include:
[0044] S8.1: Preliminary extraction of elevator fault type and fault probability features through the first convolutional layer;
[0045] S8.2: Reduce the dimension of the extracted features through the maximum pooling layer;
[0046] S8.3: Extract features of elevator failure and failure probability in the second convolutional layer;
[0047] S8.4: The features extracted by the first convolutional layer and the second convolutional layer are converted into a one-dimensional array through a flattening layer, and the converted one-dimensional array is integrated at the fully connected level;
[0048] S8.5: The output layer predicts the maintenance cycle of each elevator based on the integrated features.
[0049] In summary, the present invention is beneficial in that:
[0050] The present invention provides data support and guarantee for model training and prediction of maintenance needs by collecting various raw data of IoT elevators.
[0051] After the raw data is acquired, it is cleaned to generate multivariate time series features. During the cleaning process, abnormal data is removed, missing values are filled, and unified formatting is performed to ensure data quality, providing a reliable foundation for data analysis and decision support.
[0052] The generated multivariate time series can capture more local semantic information and the connection between multivariate variables through patch operation, extracting more extensive information. And by splitting the data into smaller patches, the amount of computation can be significantly reduced.
[0053] Traditional global computation methods often require processing large amounts of data, which can bring computational and storage challenges. Using patches can make the model more efficient in processing local information, thereby improving overall performance.
[0054] The token features after the patch operation need to be relatively positionally encoded before being input into the Transformer model, so that variants of the same features at different positions can be distinguished, helping the model to more accurately identify and process local patterns and long-term dependencies in sequence data.
[0055] A dual attention mechanism is used for the token after position encoding. The intra-patch attention mechanism uses a trainable query matrix Qinit to integrate the temporal information of the time steps within each patch. The inter-patch attention mechanism performs attention calculations on the p patches after the patch operation, obtains the interactions between patches, and finally outputs p time blocks that capture the global correlation of the time series. In order to fuse the global correlation and the local details captured by the inter-patch attention, this study rearranges the intra-patch attention output into the shape of each patch, and then adds it to the time block generated by the inter-patch attention to obtain the final dual-attention output patch feature. Through such multi-scale modeling and dual attention mechanism, the model can fully explore the complex patterns and dependencies in the time series, from local to global, and capture richer temporal information.
[0056] After the patch features are convolved and pooled, the extracted features are reduced in dimension through a fully connected layer to generate the corresponding fault type. The fault type is converted into a probability distribution of the fault type through a softmax function, so that the model can output a quantitative, probabilistic prediction result, thereby providing an intuitive and explainable representation of the possible types of elevator faults.
[0057] During the model training process, the Class Balanced Loss loss function is used to calculate the loss, and the parameters of the Transformer model are updated through the Adam optimizer to reduce the loss and improve the prediction performance of the model.
[0058] In the process of model prediction of maintenance requirements, a customized convolutional block is used to generate the final maintenance requirements based on the fault type and its corresponding probability distribution, which effectively improves the accuracy and practicality of the model's fault prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 A flowchart of an elevator risk warning method based on local patch and enhanced position coding provided by the present invention.
[0060] Figure 2 Schematic diagram of the original data table category of the Internet of Things elevator in an embodiment of the present invention.
[0061] Figure 3 This is a schematic diagram of the main abnormal type information included in the original data of the present invention.
[0062] Figure 4 A schematic diagram of a data cleaning process provided by an embodiment of the present invention.
[0063] Figure 5 A sampling diagram of model fault error data provided by an embodiment of the present invention.
[0064] Figure 6 This is a structural diagram of the elevator risk warning big data model system provided in an embodiment of the present invention.
[0065] Figure 7 A schematic diagram of a self-attention mechanism provided for an embodiment of the present invention. DETAILED DESCRIPTION
[0066] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0067] The present invention provides an elevator risk warning method based on local patch and enhanced position coding, comprising the following steps:
[0068] S1: Collect IoT elevator data as raw data for risk warning and store it in the database;
[0069] S2: Clean the collected raw data to generate multivariate time series features. Multivariate time series features include dense features, sparse features, and static features, such as elevator type, number of operations, humidity, and operating speed.
[0070] S3: Preprocess and patch the generated multivariate time series features, use the patch module to segment the input time series data, split the input time series according to a certain size of window and step size, and use the patch blocks formed by segmentation as token input, where token is the smallest processing unit of input data in deep learning. It is a meaningful representation extracted and encoded from a patch block, which is used to represent fragments of time series data and serves as the basic input unit of the model.
[0071] S4: Position encoding of the input token through the relative position encoding layer;
[0072] S5: The concatenated and position-encoded comprehensive vector features are transmitted to the DualAttention Block according to the input dimension for dual attention modeling. This can avoid the Transformer focusing only on the connection between single-point time steps and extract more extensive information. The dual attention mechanism consists of intra-patch attention and inter-patch attention. The inter-patch attention method allows us to better pay attention to the global correlation of the time series. The intra-patch attention uses a trainable query matrix (Qinit) to capture the fine-grained time series information within each patch, that is, to focus on the local details within the patch. The softmax function is used to normalize the input time series features based on the calculated weights; according to the processing results of the self-attention mechanism, the comprehensive vector features are extracted and crossed through convolution pooling operations;
[0073] S6: Reduce the dimension of the output vector calculated by the dual attention mechanism, output the corresponding fault type, and convert the dimension reduction result through the softmax function to obtain the fault probability distribution;
[0074] S7: Train the multivariate Transformer model, adjust the parameters of the multivariate Transformer model, use the loss function and Adam optimizer, evaluate the model performance in combination with the validation set, and adjust the model structure or training process based on the evaluation results;
[0075] S8: A customized multi-layer convolutional block is used to post-process the predicted fault type and fault probability distribution data to extract specific maintenance requirements.
[0076] Specifically, the present invention provides an elevator risk warning method based on local patch and enhanced position coding, which is applied to the elevator risk warning big data model system. The present invention extracts IoT elevator data in the database based on time series, and uses multiple dual attention mechanisms to extract and fuse features, and uses enhanced position coding to achieve prediction and evaluation of IoT elevator failures, and extracts the corresponding maintenance requirements of the elevator based on the evaluation results. The maintenance company formulates exclusive maintenance plans to repair the elevator according to the maintenance requirements, which saves maintenance time and reduces the workload of elevator maintenance.
[0077] like Figure 2 As shown in Figure 1, in step S1, the elevator risk warning big data model collects real-time data of the elevator through the Internet of Things technology and stores the data in the database. The original data mainly includes the following types of information: basic elevator data, fault records, weather temperature data and maintenance information. The basic elevator data (as shown in Table 1) covers the important characteristics of the elevator, including manufacturing date, model, location, etc., providing key data for evaluating the basic condition of the elevator.
[0078]
[0079] Table 1
[0080] The fault record (see Table 2) records in detail the elevator’s equipment code, fault occurrence time, maintenance arrival time, repair time, and fault type. These data can be used in the multivariate Transformer model to predict fault types, analyze the time characteristics of fault occurrence, perform fault statistics, and evaluate maintenance effectiveness, further improving the accuracy of predictions and the effectiveness of maintenance decisions.
[0081]
[0082] Table 2
[0083] The fault types are divided into five categories, as shown in Table 3: trapped people, open door operation, door opening and closing, overspeed, and others.
[0084]
[0085] Table 3
[0086] Weather temperature data (see Table 4) includes historical meteorological information related to the elevator location, such as temperature and weather conditions. This data helps analyze the potential impact of different weather conditions on elevator performance.
[0087]
[0088] Table 4
[0089] Finally, the maintenance information table (see Table 5) records the maintenance details of the elevator, including registration code, maintenance time, maintenance unit, maintenance personnel, maintenance type, failed component code, and upload time, etc. This information is crucial for managing the maintenance status of the elevator, identifying potential maintenance issues, and tracking maintenance history, which helps improve maintenance efficiency and quality.
[0090]
[0091] Table 5
[0092] In step S1, the elevator risk warning big data model system extracts key features from multiple information tables through a large amount of data collection and analysis to build a comprehensive time series data set. This data set not only contains the basic information and failure history of the elevator, but also combines external environmental factors such as weather conditions to ensure the comprehensiveness of the data. These integrated data are then used to train the machine learning model to achieve accurate prediction of elevator failures. By analyzing the relationship between the operating status of the elevator, historical failure modes and the external environment, the model can predict the potential future failure types and occurrence times. This predictive capability not only improves maintenance efficiency and reduces the occurrence of sudden failures, but also provides solid data support for the long-term and stable operation of the elevator.
[0093] like Figure 3 As shown, in step S2, after the raw data of the IoT elevator is uploaded, the data will be cleaned. The data cleaning in this embodiment mainly processes four types of abnormal data, including error values, outliers, duplicate data, and missing values. Error values refer to format, type, or numerical errors that appear in a data set. Such problems may arise from missing information, input errors, or data type mismatches during the upload process. In order to effectively identify error values, it is necessary to have an in-depth understanding of the business logic and cross-validate with multiple data sources. When correcting error values, it is necessary to clarify the source of the error, which requires an in-depth understanding of the business and data.
[0094] In this embodiment, in combination with steps S1 and S2, the present invention combines the advantages of internal data association and external data source in data processing, effectively solves the problems of integrity and accuracy of IoT elevator data, and provides reliable data support for elevator fault prediction.
[0095] like Figure 4 As shown, the steps of data cleaning in the present invention include:
[0096] S2.1: Fill missing values in the original data to ensure data integrity by completing or deleting missing information;
[0097] S2.2: Outliers in the raw data are corrected, such as values that do not conform to logical or range constraints.
[0098] S2.3: Adjust the format of all data to a unified format and correct the consistency and standardization issues of data representation;
[0099] S2.4: All data are digitized and non-numerical data are converted into numerical data to facilitate subsequent analysis and processing to generate the final normalized data set.
[0100] The data is cleaned through the above steps to ensure data quality, providing a reliable basis for subsequent data analysis and decision support. After the original data is cleaned through the above steps, multivariate time series features are generated, among which the multivariate time series features include: elevator operating temperature, elevator operation times and other elevator operating status characteristics.
[0101] like Figure 5 As shown in the figure, in the process of processing raw data, the elevator risk warning big data model system counts multiple alarms of the same fault type in one day as one alarm, and converts the number of alarms on that day into new features and inputs them into the Transformer model for data training. In this way, the input data of the Transformer model is optimized, improving the accuracy and efficiency of fault prediction.
[0102] In step S3, the elevator risk warning big data model system preprocesses the generated multivariate time series features. In this embodiment, an embedding layer is used to preprocess the multivariate time series features. Specifically, static features and dense features are converted into dense features through an embedding layer, and the converted dense features and the dense features generated after data cleaning are extracted and spliced; the conversion formula is:
[0103] x1=embedding(Xsparse)
[0104] x2=embedding(Xstatic)
[0105] x3=WXdense+b
[0106] F=cat(x1,x2,x3)
[0107] Among them, x1 is the dense feature vector generated after the static feature is converted, x2 is the vector of sparse features converted to dense features, x3 is the vector of dense feature extraction, Xdense is the dense feature, Xsparse and Xstatic are the one-hot encoded features that need to be converted, W is the fully connected layer, F is the obtained comprehensive vector, b is the bias parameter, Xdense extracts features through the fully connected layer W and performs cat splicing operation.
[0108] In step S3, a patch operation is performed on the multivariate time series features to divide the original time series into multiple small blocks (patch blocks) to reduce the computational complexity. Assume that the length of the original time series is , the length of each patch block is , the number of patch blocks is , can be calculated by the following formula:
[0109]
[0110] in, Indicates a round-down operation, which means calculating the number of blocks that can be fully adapted. If the last block cannot be fully filled, it is padded by repeating the last value of the sequence to ensure that the length of the sequence can be filled. In this way, the number of time steps in the original sequence is divided by Reduced to approximately , that is, the number of input tokens is reduced to approximately the original This approach significantly reduces computational complexity and memory usage, especially when processing longer time series, which can reduce the computational burden by about times.
[0111] In step S4, before the block-based time series features are input into the Transformer model, the block-based time series features are encoded in relative positions to more accurately capture the relative distance and direction between tokens in the sequence, thereby more effectively processing time series information. This encoding method helps the model better understand the local patterns and long-term dependencies in the sequence, so that the sequence is not lost and the time series features are more vivid.
[0112] In this embodiment, the specific steps of performing position encoding include:
[0113] S4.1: Define a position encoding function that generates a unique position code based on the position of an element in a time series;
[0114] Specifically, a mathematical function such as sine and cosine function is used to generate a position-encoded vector to ensure that each element position has a unique and distinguishable representation. In this embodiment, the method of performing position encoding includes but is not limited to mathematical functions such as sine and cosine functions, and can also be other methods that can achieve position encoding. The formula for generating a position-encoded vector using mathematical functions such as sine and cosine functions is:
[0115]
[0116] in, is the position code, For location, is the dimension index, is the dimension of the model. These position vectors are independent of the data content and only represent the position of the elements in the sequence. In addition, these position vectors are added to the feature vector of each element in the sequence.
[0117] S4.2: Generate the position code corresponding to the position pos of each time series through the position coding function;
[0118] This step ensures that each location will have a unique location code, which not only reflects the absolute location information of the location, but also contains the relative location information with other locations.
[0119] S4.3: Combine the generated position encoding with the feature vector of the corresponding position in the time series.
[0120] Specifically, the combination process can be completed by addition and subtraction. The combined vector can be expressed as:
[0121]
[0122] in, is the new feature vector after combining position encoding, is the original eigenvector, is the position encoding vector.
[0123] The semantic position encoding can be seen as the similarity matrix of the initial token, while the geometric position encoding is generated based on the position index. This fundamental difference means that the semantic position encoding cannot be directly added to the input embedding like the geometric position encoding. Since the attention map also represents the similarity between tokens, and its shape is consistent with the similarity matrix between the initial tokens, this study injects the semantic position encoding into the attention map. Specifically, this study uses the L2 regularization term to constrain the difference between each layer of the attention map and the initial token similarity matrix, as the semantic PE, which is implemented using the following formula:
[0124]
[0125] in, is the temporal semantic regularization term, is the number of attention layers in the model, is the L2 distance, is the softmax operation, Indicates The attention map in the attention layer, and is the similarity matrix of the token of each patch block of the single variable at the beginning, is the query vector of each token, which is used to compare with the key vectors of other tokens ( ) to perform similarity calculation to generate attention weights. A key vector representing each token, usually with Participate in calculating the attention weight together. and The dot product (or after scaling) is used to measure the similarity between tokens and thus determine the degree of weighted summation. is the vector in the similarity matrix of the token of each patch block, indicating the similarity of the initial input data. T represents the symbol of the transposition operation. Express Transpose, Express Transpose. D is the dimension of each token embedding or vector. Here, D represents the dimension of the query, key, and value vectors, which are usually hyperparameters specified when setting up the model. The injection method of semantic position encoding is significantly different from that of traditional geometric position encoding. However, the two have similar purposes: they embed the proximity relationships between tokens into the model. Specifically, semantic position encoding injects these relationships in the semantic space, while traditional geometric position encoding includes them in the geometric space. This enhanced position encoding method enables the model to more accurately capture the position information in the time series and improve the prediction accuracy of the model.
[0126] As for step S4, the core of this step is to fuse the position information into the features of each element, so that the Transformer model can not only understand the information of each element itself, but also identify its position in the sequence. This fusion distinguishes the variants of the same feature in different positions, helping the Transformer model to more accurately identify and process local patterns and long-term dependencies (positional relationships between data) in sequence data. Through these features, the Transformer model not only learns the independent information of each element, but also their relationships and patterns in the entire sequence. This process enables the Transformer model to effectively process time series data, especially in understanding the relative positional relationships between elements, temporal dependencies in sequences, and contextual relationships.
[0127] like Figure 6 As shown, it is a structural diagram of the elevator risk warning big data model system used in the present invention. The system includes: a database module, a preprocessing module and a deep learning module. The database module is used to store the information collected by the IoT elevator and save it as an elevator basic information table, an elevator fault information table and a weather temperature table. Step S1 in the present invention is completed in the database module. The prediction module is used to clean, discretize and normalize the data to generate a time series. The deep learning module is used to predict the category and probability of elevator failure in the next cycle based on the input time series.
[0128] As for step S5, the core of this step is to use a dual attention mechanism to simultaneously capture the relationship within and between time series patches. Here, for the convenience of expression, we use a single variable time series To describe, firstly, the patch module is used to segment the input univariate time series data, and the input time series is segmented according to a certain size of window and step size, and the segmented patch blocks ( , ,…, ) as the token input and transmitted to the Dual Attention Block for dual attention modeling, which can avoid the Transformer focusing only on the connection between single-point time steps and extract more extensive information.
[0129] The dual attention mechanism consists of intra-patch attention and inter-patch attention. The inter-patch attention method allows us to better pay attention to the global correlation of the time series, and the intra-patch attention helps to capture the local details within a single patch. The intra-patch attention mechanism uses a trainable query matrix Qinit to integrate the temporal information of the time steps within each patch block. Patch blocks After this patch-by-patch attention mechanism, each patch block is transformed from the original time series length Reduced to length 1. Finally, the attention results in all patch blocks are connected to form a length of The local details of the adjacent time steps in the time series are represented. The inter-patch attention mechanism performs attention calculation on the p patch blocks after the patch operation, obtains the interaction between the patch blocks, and finally outputs p time blocks that capture the global correlation of the time series. In order to fuse the global correlation and the local details captured by the inter-patch attention, this study rearranges the intra-patch attention output into the shape of each patch block, and then adds it to the time block generated by the inter-patch attention to obtain the final dual-attention output patch. The specific steps include:
[0130] S5.1: In this step, the first input time series is the matrix ,in is the length of the original time series, As the feature dimension, the time series is divided into patches of different sizes. Assume that the length of each patch is .
[0131] S5.2: For each patch matrix , first perform feature embedding and obtain , represents the embedded feature matrix of the time series inside the i-th patch block, and represents the time series information within the patch block. is the feature dimension after embedding. For each time step in the patch block, self-attention operation is performed. Please refer to Figure 7 As shown in Figure 2. The calculation of attention involves the calculation of query, key and value. It is done by linear transformation from the embedded feature matrix The feature vector obtained at each time step will generate a query vector, key vector and value vector. The specific formula is:
[0132]
[0133] in, , , are the learned weight matrices corresponding to the transformations of queries, keys, and values, respectively.
[0134] The specific calculation formula of its self-attention mechanism is:
[0135]
[0136] T represents the symbol of the transpose operation. and The inner product of the transpose of measures the similarity between the query and the key, divided by This is to scale and avoid values that are too large. Operation, get the attention distribution of each time step , each patch block transitions from the original input length S to a length of 1. The above formula is used to calculate the relationship between the time steps within each patch block. Note: Through the above formula, we calculate the relationship (similarity) between the time steps within each patch block through the self-attention mechanism, and The attention output is obtained by weighted summation. This enables the model to capture the dependencies between time steps within the patch block and extract the local features of each patch block. Thus, the internal attention output of each patch block is obtained. Finally, the attention results in each patch block are concatenated to generate a size of The specific formula is as follows:
[0137]
[0138] Where P is the number of patches, and the final concatenated matrix Contains the attention outputs within all patch blocks, indicating the temporal relationship within each patch block.
[0139] S5.3: For inter-patch attention, the matrix of the original time series Along the feature dimension Perform feature embedding operation to obtain , and then rearrange the data to change the length S of the patch block and the embedding feature dimension The two are combined to get the embedded matrix ,in, (That is, combining the length of the patch block and the embedding feature dimension). Then the attention weights between patches are calculated to capture the global correlation between patches. Through the self-attention mechanism, the attention weights between patches capture the relationship between each patch block and all other patches (global comparison), and achieve global fusion of features through the weighted sum value matrix. In the end, the generated features not only contain the local information of each patch block, but also integrate the global influence of all other patch blocks, thereby capturing the global correlation of the time series. The formula is as follows:
[0140]
[0141] in, It captures the global temporal dependencies between different patch blocks. Among them, the query matrix: , the bond matrix: , value matrix: , T represents the symbol of the transposition operation, which means transposing the matrix K. Through the transposition operation, the query and key matrices can perform dot product calculations to obtain the similarity or attention weight between each patch block, capture the global relationship between the patch blocks, and obtain the attention output matrix between the patch blocks. .
[0142] S5.4: Finally, the dual attention mechanism captures local and global correlations in the time series by fusing the attention from within each patch block and the attention between different patches. The matrix is obtained by reordering the attention outputs within the patch block and performing a linear transformation on the length dimension of the patch block. The time steps in each patch are then merged and combined with the attention output between the patches. Add together to get the final dual attention output .
[0143] In step S6, the features extracted by the dual attention mechanism are passed to a specially configured fully connected layer for final fault prediction. The main function of this fully connected layer is to reduce the dimensionality of the time series features extracted by the Transformer block in the Transformer model. Specifically, this layer reduces the high-dimensional feature vector to 5 dimensions, each dimension representing a specific fault type. This dimensionality reduction process effectively maps the complex time series features extracted by the Transformer model into a lower-dimensional vector space. Subsequently, the reduced 5-dimensional vector is input into the softmax activation function, which converts the vector into a probability distribution, in which each element represents the probability of the corresponding fault type occurring. This conversion enables the model to output prediction results in the form of probabilities, thereby providing intuitive and explainable fault type predictions, providing strong support for elevator fault diagnosis.
[0144] In step S7, the Transformer model is trained to optimize its performance by adjusting the parameters of the model. During the training process, the optimized parameters include hyperparameters and other parameters of the model itself. In order to improve the predictive ability of the model, the Class Balanced Loss loss function and the Adam optimizer are used, and the effect of the model is evaluated through the validation set. The validation set consists of 36 months of time series data, of which the data of the last 6 months is not used for training but is evaluated as a validation set. Specifically, the Class Balanced Loss loss function is used to process class-imbalanced data sets. During the training process, it improves the classification performance of the model for minority classes by assigning greater weights to minority classes and reducing the weights of majority classes. This method can prevent the model from being biased towards the majority class when facing class imbalance, and help the model better learn the characteristics of minority classes, thereby improving the prediction accuracy of these more difficult to identify categories. In addition, the Adam optimizer combines the advantages of the Momentum and RMSprop optimizers, and can calculate an adaptive learning rate for each parameter, thereby improving the training efficiency of the model. The personalized learning rate adjustment feature of the Adam optimizer ensures a more stable and rapid convergence process while reducing the need to adjust hyperparameters. This optimization method is particularly effective in the training of complex neural network models, which can accelerate training and improve model performance. The training process is optimized by continuously increasing the number of training times and evaluating the learning progress of the model after each training. The neural network model in this embodiment adopts a Transformer structure, and its timing prediction performance is optimized by the above method during the training process.
[0145] In step S8, this step is a prediction process and predicts the maintenance needs of the elevator. Specifically, the five fault types predicted in step S6 and their probability distribution data are post-processed through a complex multi-layer convolution block. The custom multi-layer convolution block has two convolution layers and one fully connected layer, and each layer is equipped with corresponding weights ( , , , ) and the bias term ( , , , ). The steps of extracting specific maintenance requirements using a custom multi-layer convolutional block include:
[0146] S8.1: The first convolutional layer is used to preliminarily extract the features f1(x) of the elevator fault type and fault probability. The calculation formula of the first convolutional layer is:
[0147] f1(x)=ReLU( *x+ )
[0148] Among them, x is the input fault type and the corresponding fault probability distribution data, is the fault weight of the first convolutional layer and the fully connected layer, is the bias term of the first convolutional layer, and ReLU is the activation function.
[0149] S8.2: The extracted features are reduced in dimension through the maximum pooling layer; the calculation formula is:
[0150] p(x)=MaxPool(f1(x))
[0151] Among them, p(x) is the feature after dimensionality reduction, MaxPool is the maximum pooling operation, and f1(x) is the feature of the first-layer fault type and fault probability.
[0152] S8.3: Extract the features of elevator failure and failure probability of the second convolutional layer. The extraction formula is:
[0153] f2(x)=ReLU( *p(x)+ )
[0154] Among them, f2(x) is the feature of the second layer fault type and fault probability, ReLU is the activation function, is the weight of the second convolutional layer, is the bias term of the second convolutional layer.
[0155] S8.4: The features extracted by the first and second convolutional layers are converted into a one-dimensional array through a flattening layer , integrate the one-dimensional arrays generated by the transformation at the fully connected level.
[0156] The flattening layer converts the features of the two convolutional layers into an array The formula is:
[0157] =Flatten(f2(x))
[0158] Among them, Flatten is an operation that converts multi-dimensional features into a one-dimensional array, and f2(x) is the feature of the second-layer fault type and fault probability.
[0159] Fully connected layer The formula for integrating the array is:
[0160] =ReLU( * + )
[0161] in, is the weight of the fully connected layer, is the bias term of the fully connected layer, is an array.
[0162] S9.5: The output layer predicts the maintenance cycle of each elevator based on the integrated features.
[0163] Specifically, in this step, the activation function σ of the output layer predicts the maintenance cycle of each elevator based on the integrated features, thereby achieving effective interpretation of the fault type and probability distribution and accurate formulation of the maintenance strategy, and also improving the accuracy and practicality of fault prediction. The formula for the activation function σ to predict the maintenance cycle is:
[0164] y=σ( * + )
[0165] Among them, y is the maintenance cycle, is the weight of the output layer, is the fully connected layer, is the bias term of the output layer.
[0166] As shown in Tables 6-10, it is a schematic diagram of the prediction results of the system provided by the example of the present invention for different fault types. The actual measurement objects are 8,000 elevators of KONE Elevator. The predictions are made for the four time periods of early March, late March, early April and late April in KONE data. The results are shown in the table: The model shows stability in terms of door opening and closing, and can achieve a hit rate of about 30%. It can also show a hit rate of about 20% in terms of other faults. This proves the effectiveness of the model. Moreover, even in the case of a small sample, faults such as trapped people, door opening and overspeeding can be predicted.
[0167]
[0168] Table 6
[0169]
[0170] Table 7
[0171]
[0172] Table 8
[0173]
[0174] Table 9
[0175]
[0176] Table 10
[0177] Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
Claims
1. An elevator risk warning method based on local patch and enhanced position coding, characterized in that: The following steps are involved: S1: Collect IoT elevator data as raw data for risk warning and store it in the database; S2: Clean the collected raw data to generate multivariate time series features, which include dense features, sparse features and static features; S3: Preprocess and patch the generated multivariate time series features. Use the patch module to segment the input time series data. Segment the input time series according to a certain window size and step size, and use the segmented patch blocks as token input. S4: The input token is position-encoded through the relative position encoding layer to obtain the comprehensive vector feature; S5: The comprehensive vector features are transmitted to the dual attention module according to the input dimension for dual attention modeling. The dual attention module includes intra-patch attention and inter-patch attention. The softmax function is used to normalize the input time series features based on the calculated weights. According to the processing results of the self-attention mechanism, the comprehensive vector features are extracted and crossed through convolution and pooling operations; S6: Reduce the dimension of the output vector calculated by the dual attention mechanism, output the corresponding fault type, and convert the dimension reduction result through the softmax function to obtain the fault probability distribution; S7: Train the multivariate Transformer model, adjust the parameters of the multivariate Transformer model, use the loss function and Adam optimizer, evaluate the model performance in combination with the validation set, and adjust the model structure or training process based on the evaluation results; S8: A customized multi-layer convolutional block is used to post-process the data of predicted fault types and fault probability distribution to extract specific maintenance requirements.
2. The elevator risk warning method based on local patch and enhanced position coding according to claim 1 is characterized in that: In step S3, the embedding layer is used to preprocess the multivariate time series features, the static features and dense features are converted into dense features through the embedding layer, and the converted dense features and the dense features generated after data cleaning are extracted and spliced; the conversion formula is: x1=embedding(Xsparse) x2=embedding(Xstatic) x3=WXdense+b F=cat(x1,x2,x3), where x1 is the dense feature vector generated after the static feature is converted, x2 is the vector of sparse features converted to dense features, x3 is the vector of dense feature extraction, Xdense is the dense feature, Xsparse and Xstatic are the one-hot encoded features that need to be converted, W is the fully connected layer, F is the obtained comprehensive vector, b is the bias parameter, Xdense extracts features through the fully connected layer W and performs cat splicing operation.
3. The elevator risk warning method based on local patch and enhanced position coding according to claim 1 is characterized in that: In step S3, the patch operation uses the following formula: Among them, L is the length of the original time series, S is the length of each patch block, and P is the number of patch blocks. This is a rounding-down operation, which means that only the number of patches that can be fully adapted is considered.
4. The elevator risk warning method based on local patch and enhanced position coding according to claim 1 is characterized in that: In step S4, the position coding includes geometric position coding, and the specific steps of performing geometric position coding include: S4.1: define a position encoding function, wherein the position encoding function generates a unique position code according to the position of the element in the time series; S4.2: Generate the position code corresponding to each time series position through the position code function; S4.3: Combine the generated position encoding with the feature vector of the corresponding position in the time series.
5. The elevator risk warning method based on local patch and enhanced position coding according to claim 4 is characterized in that: In step S4, the position encoding also includes semantic position encoding, which is injected into the attention map, and the L2 regularization term is used to constrain the difference between the similarity matrix of each layer of attention map and the initial comprehensive vector feature as the semantic PE, using the following formula: in, is the temporal semantic regularization term, is the number of attention layers in the model, is the L2 distance, is the softmax operation, Indicates Attention map in the layer attention layer, is the similarity matrix of the token of each patch block of the single variable at the initial time, is the query vector for each token, A key vector representing each token, is the vector in the similarity matrix of the token of each patch block, and T represents the symbol of the transposition operation. Express Transpose, Express Transpose, D represents the dimension of query vector, key vector and value vector.
6. The elevator risk warning method based on local patch and enhanced position coding according to claim 5 is characterized in that: In step S5, the specific steps of dual attention modeling include: S5.1: The input time series is a matrix ,in is the length of the original time series, As feature dimension, the matrix Divide into multiple patch blocks. Assume that the length of each patch block is S, and get the matrix of each patch block after division. ; S5.2: Perform feature embedding for each patch block to obtain the feature matrix representing the i-th patch block after embedding into the internal time series ,in The query vector, key vector, and value vector are generated through the time step feature vector in each patch block, and the internal attention output of each patch block is obtained by combining the softmax operation. , concatenate the internal attention output of each patch block , generate the attention output matrix within the patch , where P is the number of patches; S5.3: The matrix of the original time series Combine the length S and the feature dimension The feature embedding operation is performed to obtain the embedded matrix ,in, , calculate the attention weights between patches and obtain the attention output matrix between patches ; S5.4: By reordering the attention outputs within the patch and performing a linear transformation on the length dimension of the patch block, we obtain the matrix , merge the time steps in each patch block and output them with the attention between patches Add together to get the dual attention output .
7. The elevator risk warning method based on local patch and enhanced position coding according to claim 6 is characterized in that: In step S5.2, the internal attention output of each patch block is The calculation formula is: in, are the query vector, key vector, and value vector of the time step in each patch block, respectively. T represents the symbol of the transpose operation. Express Transpose, is the feature dimension after embedding.
8. The elevator risk warning method based on local patch and enhanced position coding according to claim 6 is characterized in that: In step S5.3, the formula for calculating the attention weight between patches is: in, , , They are query matrix, key matrix and value matrix respectively. T represents the symbol of transpose operation. Express Transpose, .
9. The elevator risk warning method based on local patch and enhanced position coding according to claim 1 is characterized in that: The original data of risk warning in step S1 includes: basic information of elevator, fault information, weather and temperature information, and maintenance information.
10. The elevator risk warning method based on local patch and enhanced position coding according to claim 1, characterized in that: In step S8, the customized multi-layer convolution block includes: two convolution layers and one fully connected layer. The steps of extracting specific maintenance requirements by the multi-layer convolution block include: S8.1: Preliminary extraction of elevator fault type and fault probability features through the first convolutional layer; S8.2: Reduce the dimension of the extracted features through the maximum pooling layer; S8.3: Extract features of elevator failure and failure probability in the second convolutional layer; S8.4: The features extracted by the first convolutional layer and the second convolutional layer are converted into a one-dimensional array through a flattening layer, and the converted one-dimensional array is integrated at the fully connected level; S8.5: The output layer predicts the maintenance cycle of each elevator based on the integrated features.
Citation Information
Patent Citations
Elevator failure recognition method based on convolutional neural network
CN108178037A
Image classification method, and training method and device of image classification model
CN114418030A