Elevator intelligent operation and maintenance management method and system based on multi-modal neural network
By fusing multi-source sensor data from elevators using a multimodal neural network and dynamically adjusting feature weights, the problem of unstable data fusion in elevator operation and maintenance is solved. This enables three-dimensional perception of elevator status and fault diagnosis, improving the stability and safety of operation and maintenance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-08
- Publication Date
- 2026-04-14
AI Technical Summary
Existing intelligent elevator operation and maintenance methods struggle to effectively integrate multi-source sensor data in complex environments, leading to unstable status assessment conclusions. In particular, in scenarios with high personnel density and continuous operation, fixed-weight fusion strategies are susceptible to interference from occasional factors, resulting in misjudgments or fluctuations in assessments.
A multimodal neural network is used for multi-source time-series data acquisition and feature extraction. A multimodal joint feature vector is generated through a cross-modal feature fusion layer. Combined with a multidimensional health status representation polyhedron, the feature weights are dynamically adjusted to make fault diagnosis and maintenance decisions.
It improves the stability and reliability of elevator status assessment, can accurately locate specific component anomalies, reduce ineffective operation and maintenance, improve the efficiency of operation and maintenance resource allocation, avoid sudden failures, and enhance the safety and reliability of elevator operation.
Smart Images

Figure CN121493741B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of operation and maintenance management technology, and in particular to an intelligent operation and maintenance management method and system for elevators based on multimodal neural networks. Background Technology
[0002] In the field of intelligent elevator operation and maintenance, using multiple types of sensors for status monitoring is a common practice. Existing methods typically involve installing vibration sensors, sound sensors, and image sensors at key locations in the elevator to collect multi-source data reflecting the equipment's operating status, and then conducting fault analysis or health assessments. A typical approach is to extract features from each type of sensor data separately and then perform feature-level or decision-level fusion using predefined fixed weights or rules. However, the actual operating environment of elevators is complex and variable, especially in public transportation scenarios with high passenger density and long continuous operating times (such as escalators or elevators in subway stations). The data collection process is easily affected by various incidental factors unrelated to the equipment's own health status. In this situation, a fixed data fusion strategy may not be able to adequately adapt to the real-time changes in the quality of different data sources, potentially affecting the stability of the final status assessment conclusions.
[0003] Taking the alignment assessment of a vertical elevator guide rail in a subway station as an example, existing methods typically integrate vibration data (reflecting operational stability) and visual data (reflecting the relative position of the guide rail). In actual operation, when a subway train approaches or exits the station, the low-frequency vibrations it generates are transmitted through the building structure, which may cause regular disturbances in the vibration sensor signals of the elevator guide rail. At the same time, due to the concentrated flow of passengers entering and exiting the elevator car during this period, the monitoring angle of some visual sensors deployed in the shaft may be obstructed, resulting in temporary target loss or blurring in the collected image sequence. If the system still mechanically integrates these two types of features according to a preset fixed weight, the temporarily disturbed vibration features and the degraded visual features may not be effectively identified during the integration process. Their negative impact may be carried into subsequent analysis, potentially leading to misjudgment of the actual alignment status of the guide rail or fluctuations in the assessment results. Summary of the Invention
[0004] The technical problem to be solved by this invention is to provide an intelligent operation and maintenance management method and system for elevators based on multimodal neural networks, so as to realize intelligent operation and maintenance throughout the entire life cycle of elevators.
[0005] To solve the above-mentioned technical problems, the technical solution of the present invention is as follows:
[0006] Firstly, an intelligent elevator operation and maintenance management method based on multimodal neural networks, the method comprising:
[0007] Step 1: Collect multi-source time-series monitoring data and construct a multimodal neural network. Input the multi-source time-series monitoring data into the multimodal neural network to obtain vibration feature vector, acoustic feature vector and visual feature vector;
[0008] Step 2: Input the vibration feature vector, acoustic feature vector and visual feature vector into the cross-modal feature fusion layer in the multimodal neural network for fusion to generate a multimodal joint feature vector;
[0009] Step 3: Based on the multimodal joint feature vector, a steady-state reference feature vector is calculated in the feature space as the benchmark feature anchor point, and five key health status characterization vectors are defined: traction machine bearing wear, brake shoe clearance, guide rail alignment, car vertical stability, and landing door opening and closing synchronization. Using the benchmark feature anchor point as the vertex, the five key health status characterization vectors are synthesized as edge vectors to construct a multidimensional health status characterization polyhedron.
[0010] Step 4: Map the multidimensional health status representation polyhedron to a reference space defined by a pre-defined regular polyhedron, and calculate the vertex deviation, volume overlap ratio, and surface area difference coefficient.
[0011] Step 5: Calculate the multimodal data reliability assessment vector based on the vertex deviation vector, volume overlap ratio, and surface area difference coefficient.
[0012] Step 6: Dynamically adjust the feature weights in cross-modal fusion using the multimodal data reliability assessment vector, and make fault diagnosis and maintenance decisions based on the fusion features generated after weight adjustment.
[0013] Secondly, an elevator intelligent operation and maintenance management system based on multimodal neural networks includes:
[0014] The acquisition module is used to acquire multi-source time-series monitoring data and construct a multimodal neural network. The multi-source time-series monitoring data is input into the multimodal neural network to obtain vibration feature vectors, acoustic feature vectors and visual feature vectors.
[0015] The fusion module is used to input the vibration feature vector, acoustic feature vector and visual feature vector into the cross-modal feature fusion layer in the multimodal neural network for fusion to generate a multimodal joint feature vector;
[0016] The module is used to calculate a steady-state reference feature vector in the feature space based on the multimodal joint feature vector as the benchmark feature anchor point, and to define five key health status characterization vectors: traction machine bearing wear, brake shoe clearance, guide rail alignment, car vertical stability, and landing door opening and closing synchronization. Using the benchmark feature anchor point as the vertex, the five key health status characterization vectors are synthesized as edge vectors to construct a multidimensional health status characterization polyhedron.
[0017] The quantization module is used to map the multidimensional health status representation polyhedron to a reference space defined by a preset regular polyhedron, and to calculate the vertex deviation, volume overlap ratio and surface area difference coefficient.
[0018] The evaluation module is used to calculate the reliability evaluation vector of multimodal data based on the vertex deviation vector, volume overlap ratio, and surface area difference coefficient.
[0019] The decision module is used to dynamically adjust the feature weights in cross-modal fusion using multimodal data reliability assessment vectors, and to make fault diagnosis and maintenance decisions based on the fusion features generated after weight adjustment.
[0020] Thirdly, a computer-readable storage medium storing a program that, when executed by a processor, implements the method.
[0021] The above-described solution of the present invention has at least the following beneficial effects:
[0022] By acquiring multi-source time-series data and extracting multi-modal features, vibration, acoustic, and visual information are integrated to overcome the limitations of single-modal data in characterizing elevator operation and achieve a three-dimensional perception of equipment status. Based on the reliability of multi-modal data, the fusion weights are dynamically adjusted to address the problem of fixed fusion strategies struggling to cope with environmental interference, reducing the negative impact of disturbed data on assessment results and improving the stability and reliability of status assessments. By constructing a multi-dimensional health status representation polyhedron, abstract features are transformed into intuitive geometric models. Combined with key health status vectors and deviation analysis, anomalies in specific components such as traction machines and brakes can be located, improving the targeting of fault diagnosis. Fault identification and lifespan prediction are achieved based on dynamic fusion features, promoting the shift from periodic maintenance to on-demand maintenance, reducing ineffective maintenance operations, and improving the efficiency of maintenance resource allocation. Through multi-modal data cross-validation and early anomaly warning, potential fault risks are identified in advance, avoiding operational interruptions caused by sudden failures and improving the safety and reliability of elevator operation. Attached Figure Description
[0023] Figure 1 This is a flowchart illustrating the intelligent elevator operation and maintenance management method based on a multimodal neural network provided in an embodiment of the present invention.
[0024] Figure 2 This is a schematic diagram of an elevator intelligent operation and maintenance management system based on a multimodal neural network, provided by an embodiment of the present invention. Detailed Implementation
[0025] Exemplary embodiments of the present disclosure will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure may be implemented in various forms and should not be limited to the embodiments set forth herein. Rather, these embodiments are provided so that this disclosure will be thorough and complete, and will fully convey the scope of the disclosure to those skilled in the art.
[0026] like Figure 1 As shown, embodiments of the present invention propose an intelligent elevator operation and maintenance management method based on a multimodal neural network, the method comprising the following steps:
[0027] Step 1: Collect multi-source time-series monitoring data and construct a multimodal neural network. Input the multi-source time-series monitoring data into the multimodal neural network to obtain vibration feature vector, acoustic feature vector and visual feature vector;
[0028] Step 2: Input the vibration feature vector, acoustic feature vector and visual feature vector into the cross-modal feature fusion layer in the multimodal neural network for fusion to generate a multimodal joint feature vector;
[0029] Step 3: Based on the multimodal joint feature vector, a steady-state reference feature vector is calculated in the feature space as the benchmark feature anchor point, and five key health status characterization vectors are defined: traction machine bearing wear, brake shoe clearance, guide rail alignment, car vertical stability, and landing door opening and closing synchronization. Using the benchmark feature anchor point as the vertex, the five key health status characterization vectors are synthesized as edge vectors to construct a multidimensional health status characterization polyhedron.
[0030] Step 4: Map the multidimensional health status representation polyhedron to a reference space defined by a pre-defined regular polyhedron, and calculate the vertex deviation, volume overlap ratio, and surface area difference coefficient.
[0031] Step 5: Calculate the multimodal data reliability assessment vector based on the vertex deviation vector, volume overlap ratio, and surface area difference coefficient.
[0032] Step 6: Dynamically adjust the feature weights in cross-modal fusion using the multimodal data reliability assessment vector, and make fault diagnosis and maintenance decisions based on the fusion features generated after weight adjustment.
[0033] In this embodiment of the invention, multi-source time-series data acquisition and multi-modal feature extraction are used to integrate vibration, acoustic, and visual information, overcoming the limitations of single-modal data in characterizing elevator operation and achieving a three-dimensional perception of equipment status. Based on the reliability of multi-modal data, the fusion weights are dynamically adjusted to address the problem of fixed fusion strategies being unable to cope with environmental interference, reducing the negative impact of interfered data on the evaluation results and improving the stability and reliability of the status assessment. By constructing a multi-dimensional health status representation polyhedron, abstract features are transformed into intuitive geometric models. Combined with key health status vectors and deviation analysis, anomalies in specific components such as traction machines and brakes can be located, improving the targeting of fault diagnosis. Fault identification and lifespan prediction are achieved based on dynamic fusion features, promoting the transformation of operation and maintenance from periodic inspections to on-demand maintenance, reducing ineffective operation and maintenance, and improving the efficiency of operation and maintenance resource allocation. Through multi-modal data cross-validation and early anomaly warning, potential fault risks are identified in advance, avoiding operational interruptions caused by sudden faults and improving the safety and reliability of elevator operation.
[0034] In a preferred embodiment of the present invention, step 1 includes:
[0035] Step 100: Synchronously collect multi-source time-series monitoring data using a heterogeneous sensor array deployed in key elevator components. The heterogeneous sensor array includes a vibration acceleration sensor mounted on the traction machine bearing housing, a broadband acoustic sensor mounted on the top of the car, and image sensing units deployed in the hoistway and landing door areas. Specifically, this includes: first, deploying the heterogeneous sensor array and synchronously collecting multi-source time-series monitoring data; second, accurately installing the heterogeneous sensor array according to the operating characteristics of key elevator components; and third, fixing the vibration acceleration sensor to the outer end face of the traction machine bearing housing, ensuring the sensor probe is completely flush with the bearing housing surface. The system is designed to capture real-time vibration acceleration changes generated during bearing rotation. A wideband acoustic sensor is installed at the center of the car's top, with its pickup direction facing the inside of the hoistway, to collect wide-frequency acoustic signals generated by components such as the traction machine, guide rails, and brakes. Image sensing units are deployed at the top, bottom, and middle of the hoistway, and an image sensing unit is also installed above the inner side of each landing door. This ensures that the shooting angle of all image sensing units fully covers the entire length of the guide rails, the landing door opening and closing mechanism, and key components on the car top, enabling continuous image acquisition of the relative position of the guide rails and the opening and closing actions of the landing doors. This ensures the spatiotemporal consistency of multi-source data. To ensure consistency, a unified clock synchronization mechanism is used to calibrate the time of all sensors, ensuring that the start time of data acquisition for each sensor is consistent and that the same acquisition frequency is set. The acquisition frequency for both the vibration acceleration sensor and the broadband acoustic sensor is set to 1000 Hz. This value is determined based on the typical rotational frequency range of elevator traction machine bearings. The typical speed of an elevator traction machine is 100 to 1500 rpm, corresponding to a rotational frequency of 1.67 to 25 Hz. According to the Nyquist sampling theorem, the sampling frequency must be higher than twice the highest frequency of the signal. A sampling frequency of 1000 Hz can adequately cover the high-frequency harmonic signals generated by bearing failure (not...). (Exceeding 500 Hz) to avoid signal aliasing; the acquisition frequency of the image sensing unit is set to 30 frames / second. This value takes into account the dynamic capture requirements of elevator operation status and data storage costs. The frame rate of 30 frames / second can clearly record dynamic processes such as the opening and closing of landing doors (about 1 to 3 seconds in total) and the relative position changes of the guide rails during the car's operation, meeting the requirements of temporal continuity for visual feature extraction; during the acquisition process, the vibration acceleration sensor continuously acquires the original temporal data of vibration acceleration that changes over time, the broadband acoustic sensor continuously acquires the original temporal data of sound pressure level that changes over time, and the image sensing unit continuously acquires the original image sequence.
[0036] For the raw time-series vibration acceleration data, a moving average filtering method is first used for noise reduction. A sliding window with a length of 5 data points is selected, and each data point within the window is summed with its adjacent data points according to weights (the weight of the data point at the center of the window is set to 0.4, the weights of the adjacent data points on both sides of the center are each set to 0.2, and the weights of the two data points at the edge of the window are each set to 0.1, with the sum of all weight coefficients being 1). This weight allocation follows the principle that the central data point contributes the most and the edge data points contribute the least, which can smooth noise while preserving the abrupt change characteristics of the original signal to the greatest extent and avoiding the loss of fault characteristics due to over-filtering; then normalization is performed. The normalized vibration acceleration time series data is obtained by subtracting the minimum value of the time series data from the value of each data point, and then dividing by the difference between the maximum and minimum values of the time series data. For broadband acoustic raw time series data, the frequency range signals related to elevator operation are first screened out using a bandpass filter (the filter frequency range is set to 20 Hz to 5000 Hz) to remove low-frequency environmental interference (below 20 Hz) and high-frequency noise (above 5000 Hz). The setting of this frequency range is based on the distribution characteristics of acoustic signals of elevator key component failures. Fault signals such as traction machine bearing wear and brake noise are mainly distributed in the range of 20 to 20 Hz. The guide rail friction noise is distributed between 2000 and 5000 Hz. Covering this range can completely preserve the effective acoustic characteristics. Next, frame segmentation is performed, dividing the continuous time-series data into frames of 0.02 seconds each. The root mean square value of the sound pressure level data within each frame is calculated to obtain the sound pressure level feature value corresponding to each frame, forming the processed acoustic time-series data. For the original image sequence, each frame is first converted to grayscale, and the red, green, and blue channel values of each pixel are weighted and summed according to a fixed ratio (red channel weight set to 0.299, green channel weight set to 0.587, blue channel weight set to 0.587, and blue channel weight set to 0.587). The grayscale value of the pixel is obtained by setting the weight ratio to 0.114. This weight ratio makes the grayscale image more consistent with the human eye's perception of brightness, improving the accuracy of subsequent feature extraction. Then, Gaussian filtering is performed (Gaussian kernel size is set to 3×3, standard deviation is set to 1.0) to reduce noise. The 3×3 Gaussian kernel size can effectively filter out salt-and-pepper noise in the image while avoiding blurring of image edges. The standard deviation of 1.0 corresponds to a moderate filtering intensity, which meets the image preprocessing requirements under the complex lighting environment in the elevator shaft. Finally, the pixel difference between adjacent frames is calculated to extract motion features in the image sequence, forming the processed image sequence data.
[0037] Step 101: Construct a multimodal neural network. This multimodal neural network includes parallel vibration feature extraction networks, acoustic feature extraction networks, and visual feature extraction networks. Time-series data collected by a vibration acceleration sensor is input into the vibration feature extraction network, which outputs a vibration feature vector. Time-series data collected by a broadband acoustic sensor is input into the acoustic feature extraction network, which outputs an acoustic feature vector. Image sequences collected by an image sensing unit are input into the visual feature extraction network, which outputs a visual feature vector. Specifically, this involves constructing a multimodal neural network using a parallel architecture. The core design idea of this network is to design a feature extraction sub-network adapted to the data type for each type of data. The three sub-networks operate independently and output a unified feature dimension. The specific construction process is as follows: First, a vibration feature extraction network is built, which is designed for vibration acceleration... The design of one-dimensional sequence characteristics for time-series data begins with setting up an input layer to receive preprocessed vibration acceleration time-series data. The input data format is consistent with the preprocessed time-series data to ensure data transmission compatibility. Next, a one-dimensional convolutional layer is constructed as the network's first feature extraction layer. This layer has 16 convolutional kernels with a kernel size of 3. The choice of 16 kernels is because this number allows for the capture of local features of different frequencies and amplitudes in the vibration signal while avoiding an excessive number of kernels that could lead to a surge in network parameters, thus effectively controlling the risk of overfitting. The kernel size of 3 is chosen because the local features of vibration time-series data are often reflected in the changing trends of adjacent data points; a window size of 3 data points can accurately capture this local correlation, adapting to the characteristic distribution patterns of vibration signals.
[0038] After the convolutional layers are built, convolution operations are performed. Each convolutional kernel slides across the input vibration time series data according to a preset size (i.e., 3). During the sliding process, each element of the convolutional kernel is multiplied one by one with the corresponding element of the time series data. Then, all the multiplication results are summed to obtain a convolutional feature value. By continuously sliding the convolutional kernel, the entire vibration time series data is traversed, and finally, 16 one-dimensional convolutional feature maps with the same dimensions are generated. Each feature map corresponds to the type feature captured by a convolutional kernel. After the convolution operation is completed, a normalization layer is directly connected after the convolutional layer. Each generated one-dimensional convolutional feature map is normalized by subtracting the mean of the feature map from each feature value and then dividing by the standard deviation of the feature map. This ensures that the data distribution of each feature map is at a uniform scale, avoiding excessive gradient fluctuations during network training due to differences in data scale, and improving training stability. Batch normalization. After the dimensionality reduction is completed, a max pooling layer is applied, with a pooling window size of 2. A pooling window size of 2 can compress the dimension of each convolutional feature map to half of its original size. This reduces the subsequent computational load and improves operational efficiency. At the same time, by taking the maximum value within the pooling window, the key feature information of each local region is preserved, avoiding the loss of useful features. After the feature dimensionality reduction is completed, the 16 dimensionality-reduced one-dimensional convolutional feature maps output by the max pooling layer are sequentially concatenated into a continuous one-dimensional feature vector. The dimension of this vector is 256. This dimension is calculated by combining the length of the preprocessed vibration time series data with the stride of the convolutional layer, the convolutional kernel size, and the pooling window size of the pooling layer. The specific calculation logic is: (length of input time series data - convolutional kernel size + 1) ÷ pooling window size × number of convolutional kernels. This ensures that the concatenated feature vector can fully carry the key features of the vibration signal after multiple processing steps.
[0039] After the feature vectors are concatenated, a fully connected layer is built. The concatenated 256-dimensional one-dimensional feature vector is input into the fully connected layer. The core components of this layer are the weight matrix and the bias term. The weight matrix is set to a dimension of 64×256. The purpose of this dimension setting is to linearly map the 256-dimensional high-dimensional feature vector to a 64-dimensional low-dimensional space. The value range of the elements in the weight matrix is set to -0.05 to 0.05. This range of values can avoid the problem of gradient explosion caused by excessively large initial weights leading to neuron output saturation. The bias term is set to a vector containing 64 elements, each with a value of 0.01. This value can ensure that the output value of each neuron is not zero during network initialization, so that the gradient can be effectively propagated in the network and avoid the gradient vanishing situation. Finally, through the linear transformation of the fully connected layer (multiplying and summing the feature vector and the weight matrix and then adding the bias term), a vibration feature vector with a dimension of 64 is output. This vector is the core feature expression of the vibration data after targeted extraction.
[0040] While building the vibration feature extraction network, an acoustic feature extraction network was built in parallel. This network, based on the stationarity of acoustic time-series data, was optimized and adjusted from the vibration feature extraction network architecture. First, an input layer was set up to specifically receive the processed acoustic time-series data (sound pressure level feature values corresponding to each frame), adapting to the frame sequence format of the acoustic time-series data. Next, a one-dimensional convolutional layer was constructed, using the same convolutional kernel parameters as the vibration feature extraction network, setting 16 convolutional kernels with a kernel size of 3. Maintaining this consistent parameter is crucial to ensure that the intermediate feature dimension of the acoustic feature extraction network matches that of the vibration feature extraction network. After the convolutional layer was built, convolution operations were performed. Each convolutional kernel slid across the input acoustic time-series frame sequence, capturing the local correlation features of the sound pressure level feature values of each frame through element-wise multiplication and summation convolution operations, ultimately generating 16 acoustic features. The 16 acoustic convolutional feature maps are then connected to a batch normalization layer, which normalizes each map in the same way as the vibration feature extraction network, ensuring uniform data scale for acoustic features and improving network training stability. After batch normalization, an average pooling layer is applied. Considering that the core features of acoustic time-series data are reflected in the average level and steady change of sound pressure level within frames, the max pooling layer in the vibration feature extraction network is replaced with an average pooling layer. The pooling window size is set to 2, and the average value within the pooling window is used as the feature value after pooling, avoiding the loss of acoustic detail features that max pooling may cause. After the pooling operation is completed, the 16 acoustic convolutional feature maps output by the average pooling layer are concatenated in sequence to form a one-dimensional feature vector with a dimension of 256, which is consistent with the dimension of the intermediate feature vector of the vibration feature extraction network.
[0041] Next, a fully connected layer is constructed, using the same parameters as the vibration feature extraction network. The weight matrix is set to 64×256, with element values ranging from -0.05 to 0.05. The bias term is a 64-element vector, with each element having a value of 0.01, ensuring that the output dimension of the acoustic feature vector after transformation by the fully connected layer is consistent with that of the vibration feature vector. Finally, through linear transformation of the fully connected layer, the input 256-dimensional acoustic feature vector is multiplied by the 64×256-dimensional weight matrix, and then the resulting 64-dimensional vector is multiplied by the 64×256-dimensional weight matrix. The bias terms of the four elements are summed element-wise to output a 64-dimensional acoustic feature vector, thus completing the feature extraction of the acoustic data. Simultaneously, a third feature extraction sub-network, the visual feature extraction network, is built in parallel. This network, designed for the two-dimensional spatial characteristics of image sequence data, employs a deep feature extraction process. First, an input layer is set up to specifically receive the processed image sequence data, adapting to the two-dimensional pixel matrix format of each frame. Then, the first two-dimensional convolutional layer is constructed. Since image data contains rich spatial edges, textures, and other two-dimensional features, more convolutional kernels are needed to capture these features. To capture spatial features at different scales but in the same direction, this layer uses 32 two-dimensional convolutional kernels, each 3×3 in size. This size ensures sufficient spatial coverage for capturing local pixel relationships while controlling the number of kernel parameters. After the convolutional layer is built, two-dimensional convolution operations are performed. Each kernel slides across the pixel matrix of the input image frame, multiplying each element of the kernel by the corresponding pixel grayscale value in the image frame. The results are then summed to obtain a two-dimensional convolutional feature value. By traversing the entire image frame, 32 two-dimensional features are generated. Each convolutional feature map corresponds to a spatial feature type, including basic spatial features such as horizontal edges, vertical edges, texture details, and local contours in the image. A batch normalization layer is then connected to normalize each of the 32 two-dimensional convolutional feature maps, unifying the data scale and improving training stability. After batch normalization, a max pooling layer is applied, with a pooling window size of 2×2. By taking the maximum value within the pooling window, the dimensionality of the two-dimensional convolutional feature maps is reduced, compressing spatial dimensions and improving computational efficiency while preserving key edge and texture features in the image.
[0042] After completing the first convolutional pooling operation, to fully extract deep semantic features of the image (such as the geometry of the guide rails, the opening and closing states of the gates, etc.), three more convolutional and pooling layers with the same architecture are stacked after the first convolutional pooling structure. The number of convolutional kernels in the subsequent convolutional layers is set to 64, 128, and 256 respectively, with the kernel size remaining at 3×3 and the pooling window size at 2×2. Doubling the number of convolutional kernels layer by layer is a classic and effective strategy for extracting deep features in deep learning. As the number of network layers increases, more complex and abstract high-level semantic features are captured through more convolutional kernels, gradually realizing the transformation from low-level pixel features to high-level semantic features. After completing the stacking of multiple convolutional pooling layers, the 256 two-dimensional feature maps output by the last pooling layer are unfolded sequentially in row or column order, transforming them into a continuous one-dimensional feature vector with a dimension of 1024. This dimension is based on the resolution of the input image. The kernel size and stride of the four convolutional layers and the pooling window size of the four pooling layers are calculated to ensure that all key semantic features extracted from the image by the deep network can be fully contained. After the feature vector is flattened, a fully connected layer is built. The 1024-dimensional one-dimensional feature vector is input into the fully connected layer. The weight matrix of this layer is set to 64×1024. Through linear transformation, the high-dimensional visual feature vector is mapped to a 64-dimensional low-dimensional space, which is consistent with the dimension of the vibration and acoustic feature vectors to meet the cross-modal fusion requirements. The value range of the elements in the weight matrix is set to -0.05 to 0.05. The bias term is set to a vector containing 64 elements, each with a value of 0.01. The parameter settings are consistent with those of the vibration and acoustic feature extraction network to ensure the stability of network training. Finally, through the linear transformation of the fully connected layer, a 64-dimensional visual feature vector is output, completing the deep feature extraction of visual data.
[0043] By constructing the three feature extraction sub-networks in parallel as described above, the overall construction of the multimodal neural network is completed. This network outputs vibration feature vectors, acoustic feature vectors, and visual feature vectors with uniform dimensions through the three parallel feature extraction sub-networks.
[0044] In this embodiment, by deploying and synchronously collecting heterogeneous sensor arrays at key parts of the elevator, comprehensive acquisition of multi-dimensional operational data including vibration, acoustics, and vision is achieved, covering key operational status information of the elevator's core components and overcoming the limitations of single sensor data in characterizing equipment status. Targeted preprocessing effectively filters out environmental interference and data noise, improving the purity of the original data. The construction of parallel feature extraction networks in the multimodal neural network can adapt exclusive feature extraction logic according to the characteristics of different types of data, fully mining the equipment health status information contained in each modality of data and generating single-modality feature vectors with strong representational capabilities. The three feature extraction networks operate independently and have a unified output dimension, thus fully preserving the unique information of each modality of data.
[0045] In a preferred embodiment of the present invention, step 2 includes:
[0046] Step 200: The vibration feature vector, acoustic feature vector, and visual feature vector are simultaneously input into the cross-modal feature fusion layer. In this layer, a multi-head attention mechanism is used to calculate the pairwise correlation weights between the vibration, acoustic, and visual feature vectors. Specifically, the cross-modal feature fusion layer receives three 64-dimensional feature vectors from the multimodal neural network: the vibration feature vector, the acoustic feature vector, and the visual feature vector. Correlation weights are calculated synchronously for all three in this layer. To fully explore the potential correlations between different modal features, a multi-head attention mechanism is used to split the feature processing dimensions. The specific process is as follows: For the vibration feature... The acoustic feature vector, visual feature vector, and acoustic feature vector are each subjected to independent linear projection processing. Each feature vector is multiplied by a preset projection weight matrix (elements ranging from -0.05 to 0.05) and a bias term is added (each element is initially set to 0.01) to generate three intermediate feature vectors with the same dimensions (the dimensions remain 64 after projection to ensure computational consistency). Subsequently, each intermediate feature vector is split into 8 parallel attention head feature segments, each segment having 8 dimensions (64 dimensions ÷ 8 attention heads). Parallel computation of multiple attention heads improves the comprehensiveness of association capture. The process is then performed according to vibration-acoustic, vibration-visual, and acoustic-visual sequences. The method involves constructing query and key vectors for each pair of paired features. For example, when calculating the correlation between vibration and acoustic features, the split vibration feature attention head segment is used as the query vector, and the acoustic feature attention head segment as the key vector; when calculating the correlation between vibration and visual features, the vibration feature attention head segment is used as the query vector, and the visual feature attention head segment as the key vector; when calculating the correlation between acoustic and visual features, the acoustic feature attention head segment is used as the query vector, and the visual feature attention head segment as the key vector. Element-wise multiplication is performed on each query-key vector pair, and then all multiplication results are summed to obtain the initial correlation value for each vector pair. To avoid... Due to numerical bias caused by feature dimension, the initial correlation value is divided by the square root of the key vector dimension (8 dimensions) to complete the standardization process, resulting in the standardized correlation value of each vector pair. The softmax function is applied to each standardized correlation value to normalize it, so that each correlation value is transformed into a probability distribution with values between 0 and 1 and a sum of 1. This distribution is the correlation weight of the corresponding feature pair under a single attention head. Then, the correlation weights calculated by the 8 attention heads for the same feature pair are arithmetically averaged to finally obtain three stable pairwise correlation weights, namely vibration-acoustic correlation weight, vibration-visual correlation weight, and acoustic-visual correlation weight.
[0047] Step 201: Based on pairwise correlation weights, the vibration feature vector, acoustic feature vector, and visual feature vector are weighted, concatenated, and linearly transformed to generate a unified multimodal joint feature vector. Specifically, this includes: first, calculating the comprehensive weight of each feature vector, which is obtained by weighted summation of the correlation weights of that feature vector with the other two feature vectors. For example, the comprehensive weight of the vibration feature vector = vibration-acoustic correlation weight × 0.5 + vibration-visual correlation weight × 0.5; the comprehensive weight of the acoustic feature vector = vibration-acoustic correlation weight × 0.5 + acoustic-visual correlation weight × 0.5; the comprehensive weight of the visual feature vector = vibration-visual correlation weight × 0.5 + acoustic-visual correlation weight × 0.5. All weight coefficients are set to 0.5 to balance the influence of the two sets of correlations. Then, each dimension value of each feature vector is multiplied by its corresponding comprehensive weight to obtain three weighted feature vectors, ensuring stronger correlation. Features play a more significant role in the fusion process. The weighted vibration feature vector, acoustic feature vector, and visual feature vector are concatenated sequentially to form a 192-dimensional concatenated feature vector (64-dimensional + 64-dimensional + 64-dimensional). This vector fully preserves the weighted feature information of the three modes. The 192-dimensional concatenated feature vector is input into a preset linear transformation layer. Matrix multiplication is performed with a 64×192-dimensional weight matrix (elements ranging from -0.05 to 0.05), and then a bias term with 64 elements (each initially set to 0.01) is superimposed. This maps the high-dimensional concatenated features to a 64-dimensional low-dimensional space. During the linear transformation, the elements of the weight matrix are controlled within the range of -0.05 to 0.05 to avoid feature distortion caused by excessively large initial weights. Finally, a 64-dimensional unified multimodal joint feature vector is output, which integrates the key information of the three modes and has the same dimension as the original feature vector.
[0048] This embodiment calculates pairwise correlation weights through a multi-head attention mechanism, which can accurately capture the dynamic correlation between vibration, acoustic, and visual modal features. Compared with fixed-weight fusion, it is more adaptable to the changes in the correlation of different modal data during elevator operation, thus improving the targeting of feature fusion. The weighted splicing strategy based on correlation weights allows modal features with strong correlation and high information quality to play a greater role in the fusion process, while weakening the influence of irrelevant or interfering features, thereby improving the information purity of the multimodal joint feature vector. The linear transformation stage realizes the mapping of high-dimensional spliced features to a unified low-dimensional space, which not only preserves the core information of each modality but also ensures the simplicity of the joint feature vector.
[0049] In a preferred embodiment of the present invention, step 3 includes:
[0050] Step 300: Based on a predefined sliding time window, the mean of multiple continuously generated multimodal joint feature vectors within the window is calculated to obtain a steady-state reference feature vector representing the typical operating state within the window. This steady-state reference feature vector is then used as the benchmark feature anchor. Specifically, this includes: firstly, setting the length of the sliding time window to 30 sampling periods (this length is set based on the typical smoothness period of elevator operation to ensure coverage of continuous and stable operating states and avoid occasional fluctuations affecting the accuracy of the benchmark), and setting the sliding step size to 1 sampling period. That is, each time a new multimodal joint feature vector is generated, the window slides forward one position, always retaining the latest 30 multimodal features. The joint feature vector is calculated as follows: For the 30 historical continuous multimodal joint feature vectors (each with 64 dimensions) stored within the window, the mean is calculated for each dimension. Specifically, for each dimension from the 1st to the 64th dimension of the feature vector, the values of the 30 joint feature vectors within the window in that dimension are summed sequentially to obtain the sum of that dimension. Then, the sum of each dimension is divided by the number of joint feature vectors in the window, 30, to obtain the mean of that dimension. The means of the 64 dimensions are arranged sequentially to form a 64-dimensional feature vector. This vector is the steady-state reference feature vector representing the typical operating state of the elevator within the window, and it is used as the benchmark feature anchor point for subsequent health state modeling.
[0051] Step 301: From the multimodal joint feature vector at the current moment, extract the following five key health state representation vectors: a first state vector representing the wear of the traction machine bearings, a second state vector representing the brake shoe clearance, a third state vector representing the guide rail alignment, a fourth state vector representing the vertical stability of the car operation, and a fifth state vector representing the synchronicity of the landing door opening and closing. Specifically, this involves: first, obtaining the 64-dimensional multimodal joint feature vector generated at the current moment; simultaneously, retrieving the 64-dimensional steady-state reference feature vector calculated in step 300; and performing a difference operation based on these two vectors to obtain a difference feature vector reflecting the deviation of the current state from the reference. The specific calculation... The method involves subtracting the value of each dimension of the current multimodal joint feature vector from the first to the 64th dimension from the value of the corresponding dimension of the steady-state reference feature vector. That is, subtracting the value of the first dimension of the current vector from the value of the first dimension of the reference vector, subtracting the value of the second dimension of the current vector from the value of the second dimension of the reference vector, and so on. After completing the dimension-by-dimensional subtraction of all 64 dimensions, a 64-dimensional difference feature vector is formed. The value of each dimension of this difference vector directly corresponds to the degree to which the current elevator operating state deviates from the normal steady state in that feature dimension. A positive value indicates that the state in that dimension is higher than the reference level, and a negative value indicates that it is lower than the reference level. The magnitude of the absolute value corresponds to the strength of the deviation.
[0052] Subsequently, based on the pre-established feature dimension-health status association rules, the difference feature vectors were subjected to dimension filtering and numerical adjustment to generate five key health status representation vectors. The process of establishing the feature dimension-health status association rules is as follows: First, a full data collection was conducted, covering 100 elevator samples of different brands and models (including residential elevators, commercial elevators, heavy-duty elevators, etc.) with an operating age of 1 to 15 years. At a collection frequency of 10 times per second, multimodal sensor data, PLC operating parameters, and fault record data throughout the elevator's entire life cycle were continuously collected. The multimodal sensor data included vibration acceleration of the traction machine bearing and vibration signals of the brake action collected by vibration sensors, and noise in the machine room and inside the car collected by acoustic sensors. The system includes: stable noise; visual sensor-collected image features of guide rail verticality and landing door opening / closing gap; PLC operating parameters including traction machine speed, brake force, elevator speed, load weight, and landing door opening / closing time; fault record data covering the occurrence time, wear degree detection data, and maintenance / replacement records of traction machine bearing wear; alarm information and gap measurement data for abnormal brake shoe clearance; calibration records of guide rail deviation; operational vibration feedback data; vibration exceeding standards records for car vertical stability; passenger experience feedback; and opening / closing time difference data and jamming fault records for landing door synchronization, ensuring complete coverage of the entire development process of five key health conditions from initial minor abnormalities to severe failure; the second step involves processing the collected data... Correlation analysis was performed on historical data. First, each multimodal joint feature vector in the historical data was labeled with a corresponding health status tag. For traction machine bearing wear, the tags were divided into four levels: normal (wear ≤ 0.1 mm), slightly abnormal (0.1 mm < wear ≤ 0.3 mm), moderately abnormal (0.3 mm < wear ≤ 0.5 mm), and severely abnormal (wear > 0.5 mm). For brake shoe clearance, the tags were divided into four levels: normal (clearance ≤ 0.2 mm), slightly abnormal (0.2 mm < clearance ≤ 0.5 mm), moderately abnormal (0.5 mm < clearance ≤ 0.8 mm), and severely abnormal (clearance > 0.8 mm). For guide rail alignment, the tags were divided into four levels: normal (perpendicularity deviation ≤ 0.1 mm). For verticality deviation, the labels are categorized into four levels: normal (vibration acceleration ≤ 0.1g), slight abnormality (0.1g < vibration acceleration ≤ 0.3g), moderate abnormality (0.3g < verticality deviation ≤ 0.5g), and severe abnormality (verticality deviation > 0.5g). For car vertical stability, the labels are categorized into four levels: normal (vibration acceleration ≤ 0.1g), slight abnormality (0.1g < vibration acceleration ≤ 0.3g), moderate abnormality (0.3g < vibration acceleration ≤ 0.5g), and severe abnormality (vibration acceleration > 0.5g). For landing door opening and closing synchronization, the labels are categorized into four levels: normal (opening and closing time difference ≤ 0.1s), slight abnormality (0.1s < opening and closing time difference ≤ 0.3s), and moderate abnormality (0.3s < opening and closing time difference ≤ 0.1s).The system is divided into four levels: 5s), severe abnormality (opening / closing time difference > 0.5s), and critical abnormality. The labeled multimodal joint feature vectors are then matched with their corresponding health status labels. Pearson correlation analysis is used to calculate the correlation coefficient between each feature dimension and five health status categories: traction machine bearing wear, brake shoe clearance, guide rail alignment, car vertical stability, and landing door opening / closing synchronicity. Thirdly, the absolute value of the correlation coefficient is used as a quantitative indicator of the strength of the association between the feature dimension and the health status. Statistical analysis of 5000 fault and normal samples verifies this. When the absolute value of the correlation coefficient is ≥ 0.7, the feature dimension... When the linear correlation between the feature dimensions and the corresponding health status reaches a strong level, the fault identification accuracy can reach over 92% after constructing the state vector based on the feature dimensions selected by this threshold. If the threshold is lower than 0.7, a large number of weakly correlated dimensions will be introduced, leading to feature redundancy, and the fault identification accuracy will drop to below 75%. Furthermore, considering the industry standard that the strong correlation feature selection threshold in elevator fault diagnosis engineering practice is usually between 0.65 and 0.75, 0.7 was ultimately determined as the correlation strength threshold. All feature dimensions with a correlation strength higher than this threshold are bound and labeled with their corresponding health status, forming a complete feature dimension-health status correlation rule.
[0053] For the first state vector representing the wear degree of the traction machine bearing, all dimensions in the difference feature vector with a correlation strength higher than 0.7 with the traction machine bearing wear degree are selected according to the association rules. The original difference values of these dimensions are retained. For dimensions with a correlation strength lower than 0.7, their values are uniformly set to 0. After adjusting the values of all 64 dimensions, a 64-dimensional first state vector is formed. For the second state vector representing the brake shoe clearance, the same logic is used: the original values of the dimensions in the difference feature vector with a correlation strength higher than 0.7 with the brake shoe clearance are retained, and the values of the remaining dimensions are set to 0, generating a 64-dimensional first state vector. The second state vector is obtained; similarly, the dimensions that are strongly correlated (correlation strength greater than 0.7) with guide rail alignment, car vertical stability, and landing door opening and closing synchronization are selected according to the association rules. The original difference values are retained, and the values of non-correlated dimensions are set to 0. The third, fourth, and fifth state vectors with 64 dimensions are generated respectively. Through the above-mentioned dimension-by-dimensional calculation, rule matching and numerical adjustment operations, the key health status representation vectors with five dimensions of 64 are finally obtained. The non-zero dimension of each vector accurately corresponds to the feature changes of a key health status, realizing the accurate mapping from abstract feature vectors to specific health statuses.
[0054] Step 302: Using the steady-state reference feature vector as the spatial starting point, and taking the first, second, third, fourth, and fifth state vectors as spatial edge vectors, the five vertices of the multidimensional health state representation polyhedron are determined sequentially through vector addition. Specifically, this includes: using the 64-dimensional steady-state reference feature vector obtained in step 300 as the spatial starting point (this starting point corresponds to the reference position for normal elevator operation in the feature space), taking the first to fifth state vectors extracted in step 301 as the five edge vectors in the space, and calculating the five vertices of the polyhedron one by one through vector addition. The specific calculation process is as follows: the value of each dimension of the steady-state reference feature vector is respectively compared with the corresponding dimension of the first state vector. The values of the degrees are added together to obtain the 64-dimensional coordinates of the first vertex, which maps to the feature space position corresponding to the wear degree of the traction machine bearing. Each dimension of the steady-state reference feature vector is added to the corresponding dimension of the second state vector to obtain the 64-dimensional coordinates of the second vertex, which maps to the feature space position corresponding to the brake shoe clearance. The 64-dimensional coordinates of the third vertex are obtained by adding the steady-state reference feature vector to the third state vector dimension by dimension, which maps to the feature space position corresponding to the guide rail alignment degree. The 64-dimensional coordinates of the fourth vertex are obtained by adding the steady-state reference feature vector to the fourth state vector dimension by dimension, which maps to the feature space position corresponding to the vertical stability of the car operation.
[0055] By adding the steady-state reference feature vector to the fifth state vector dimension by dimension, the 64-dimensional coordinate value of the fifth vertex is obtained, which maps the feature space position corresponding to the synchronicity of the door opening and closing. Through the above five vector addition operations, the five vertices of the multidimensional health state representation polyhedron are finally determined, and each vertex is directly associated with a key health state of the elevator.
[0056] Step 303 involves sequentially connecting the steady-state reference feature vector to the five vertices of the multidimensional health state representation polyhedron, and then connecting edges between the five vertices to construct a closed multidimensional health state representation polyhedron. Specifically, this includes: connecting the spatial starting point corresponding to the steady-state reference feature vector to each of the five vertices determined in step 302, forming five edges pointing from the reference point to each health state vertex. The length and direction of each edge indirectly reflect the degree of deviation of the corresponding health state from the normal reference. Then, following the order of the first vertex, the second vertex, the third vertex, the fourth vertex, the fifth vertex, and the first vertex, the five vertices are connected sequentially to form five closed edges. Through the connection of these two types of edges, a closed multidimensional representation polyhedron is finally constructed, with the steady-state reference feature vector as the reference anchor point and the five key health state vertices as the core. The health status characterization polyhedron, whose overall shape, vertex positions, edge lengths, and other geometric characteristics comprehensively map the current health status and overall operation of the elevator's key components in the following ways: The length of the edge pointing from the reference point to each vertex is positively correlated with the deviation of the corresponding component's health status from the normal reference; the longer the length, the more severe the deviation. The direction of the edge corresponds to the characteristic dimension trend of the deviation. The spatial coordinates of each vertex directly map the health status of a type of key component; the farther the vertex is from the reference point, the higher the degree of component abnormality. Multiple vertices shifting in the same direction indicate a risk of coordinated abnormality in multiple components. The size of the polyhedron and the regularity of its faces map the overall operational stability. A moderate volume and regular shape represent a balanced state of each component and stable overall operation, while an excessively large volume and a stretched or deformed shape indicate significant abnormalities and a decrease in overall operational stability.
[0057] In this embodiment, the mean calculation method of the sliding time window can effectively filter out occasional fluctuations in elevator operation, extract typical steady-state operating characteristics as benchmark anchors, provide a stable and reliable reference benchmark for health status assessment, and avoid assessment distortion caused by data deviation at a single moment; five key health status representation vectors are extracted in a targeted manner, realizing accurate mapping between elevator core components and key operating characteristics, and improving the targeting of fault diagnosis; the abstract feature vectors are transformed into an intuitive multi-dimensional health status representation polyhedron, upgrading the elevator health status from numerical description to geometric modeling, reducing the difficulty of interpreting abstract features; the construction of the closed polyhedron fully integrates the correlation information of each key health status, preserving the independent state characteristics of individual components, and reflecting the overall correlation of the operating status of each component.
[0058] In a preferred embodiment of the present invention, step 4 includes:
[0059] Step 400: Perform a spatial affine coordinate transformation on each vertex of the multidimensional health state representation polyhedron, mapping the multidimensional health state representation polyhedron to a reference space defined by a predefined regular polyhedron. Specifically, the regular polyhedron in the reference space is a predefined standard geometric model with the same number of vertices as the multidimensional health state representation polyhedron, containing one reference anchor point and five health state vertices. The geometric parameters of this regular polyhedron are determined by collecting multiple sets of feature vector data from the normal operation phase of the entire life cycle of elevators of the same model and under the same working conditions. Based on these data, a large number of multidimensional health state representation polyhedra under normal conditions are generated according to the process from steps 300 to 303. Statistical analysis is performed on the vertex coordinates, side lengths, face areas, and other geometric parameters of these polyhedra, and the mean of each parameter is calculated. These mean values are used as the standard. The geometric parameters of the regular polyhedron represent the standard feature space form of the elevator's key components when they are in a healthy state. In this standard feature space form, the reference anchor point is located at the reference origin of the feature space. The five healthy state vertices are uniformly and symmetrically distributed around the reference anchor point. The spatial orientation of each vertex corresponds one-to-one with a type of key healthy state (traction machine bearing wear, brake shoe clearance, guide rail alignment, car vertical stability, and landing door opening and closing synchronization). The spatial angles between adjacent vertices are equal, forming a symmetrical standard space configuration where the vertex orientations are fixed and the dimensions of each healthy state do not interfere with each other. The affine transformation parameters include translation, scaling, and rotation matrices. The parameters are determined with the goal of ensuring that the reference anchor point of the multidimensional healthy state characterizing polyhedron completely coincides with the reference anchor point of the regular polyhedron. Let the coordinates of the reference anchor point of the multidimensional healthy state characterizing polyhedron be... The coordinates of the reference anchor point of the regular polyhedron are: Calculate the coordinate difference between the two reference anchor points in corresponding dimensions. This difference is the translation parameter; the translation matrix is a 65×65 augmented matrix, in the form of... Translation operations are used to make the spatial positions of the reference anchor points of the two polyhedra consistent.
[0060] To determine the scaling ratio and scaling matrix, first calculate the side lengths from the reference anchor point of the multidimensional health state representation polyhedron to the five health state vertices. Let the lengths of these five sides be... Calculate the mean of its side lengths. Next, retrieve the standard side lengths from the reference anchor point of the regular polyhedron to the five corresponding healthy state vertices, and set them as follows: Calculate the mean of its standard side length. Scaling ratio The scaling matrix is a 65×65 augmented matrix, in the form of... The overall size of the multidimensional health state representation polyhedron is adjusted by scaling operations to match the size of the regular polyhedron. The rotation matrix is determined as follows: First, a 64-dimensional feature space coordinate system matching the feature vector dimension is constructed, with the position where the reference anchor points of the two polyhedra coincide after translation as the origin. Each coordinate axis corresponds to one dimension of the feature vector. Then, the direction vectors of the five reference edges of the multidimensional health state representation polyhedron (i.e., the edges where the reference anchor points point to the five health state vertices) are calculated. Specifically, the 64-dimensional coordinates of each health state vertex are subtracted from the 64-dimensional coordinates of the reference anchor point to obtain the direction vectors of the five reference edges with a dimension of 64, forming the actual direction vector matrix. At the same time, the five corresponding standard edges of the regular polyhedron (where the reference anchor points point to the corresponding five health state standard vertices) are calculated using the same method. The direction vectors of the actual direction vector matrix (the edges of the points) are used to form a standard direction vector matrix. Finally, with the goal of minimizing the error between the actual direction vector matrix and the standard direction vector matrix, an orthogonal iteration method is used for iterative solution. The initial rotation matrix is set to a 64×64 identity matrix (diagonal elements are all 1s, off-diagonal elements are all 0s). This matrix is a typical form of an orthogonal matrix, ensuring that the direction vectors in the initial state are not subject to additional rotational interference, and the iteration process converges stably. The initial rotation matrix is multiplied by the actual direction vector matrix to obtain the projected direction vector matrix. The deviation between the projected direction vector matrix and the standard direction vector matrix is calculated (the deviation is calculated as the sum of the squares of the differences between corresponding elements of the two matrices). Based on the deviation, the elements of the rotation matrix are adjusted, and the above steps of projection, deviation calculation, and matrix adjustment are repeated until the deviation value is less than a preset threshold of 1× (This threshold is determined based on the data precision level of the 64-dimensional feature vectors and the allowable error range of orientation alignment in engineering practice, ensuring that the alignment accuracy of the orientation vectors meets the requirements of health status assessment.) The resulting 64×64 orthogonal matrix is the required rotation matrix R. This matrix ensures that the five reference edge direction vectors of the rotated multidimensional health status representation polyhedron are completely aligned with the five corresponding standard edge direction vectors of the regular polyhedron, satisfying that the spatial orientation of each vertex of the rotated multidimensional health status representation polyhedron is consistent with the orientation of the corresponding vertices of the regular polyhedron. The rotation matrix is then extended to a 65×65 augmented matrix, in the form of... The 64-dimensional coordinates of the reference anchor point and the five health state vertices of the multidimensional health state representation polyhedron are transformed into homogeneous coordinates, that is, the coordinates of any vertex are transformed into homogeneous coordinates. Its homogeneous coordinates are The affine transformation equation is: ,in The homogeneous coordinates are transformed. The homogeneous coordinates of each vertex are substituted into the affine transformation equation in turn, and the matrix multiplication operation is performed dimension by dimension to obtain the transformed homogeneous coordinates. Then, the last dimension of the homogeneous coordinates is removed to obtain the 64-dimensional new coordinates of each vertex in the reference space. Finally, the complete mapping of the multidimensional health state representation polyhedron to the reference space is realized.
[0061] Step 401: In the reference space, calculate the Euclidean distance from each vertex of the multidimensional health state representation polyhedron to the corresponding vertex of the regular polyhedron. Based on all Euclidean distances, calculate the distribution variance and mean of the Euclidean distances, and generate a vertex deviation vector based on the distribution variance and mean. Specifically, this includes: for the multidimensional health state representation polyhedron and the regular polyhedron in the reference space, calculating the Euclidean distance between each pair of corresponding vertices according to the vertex correspondence; for any pair of corresponding vertices, extracting the 64-dimensional coordinates of the two vertices, calculating the coordinate difference dimension by dimension, and performing coordinate difference calculations for each dimension. The Euclidean distance between the vertices is calculated by squaring the values, summing the squared differences across all dimensions, and then taking the square root of the sum. This process is repeated for the baseline anchor point and the five healthy vertices, resulting in six sets of Euclidean distance data. These six sets of Euclidean distance data are then summed to obtain the total Euclidean distance. This sum is divided by the number of Euclidean distance data points (six) to obtain the mean of the Euclidean distance distribution. The mean of the Euclidean distance distribution is then subtracted from each set of Euclidean distance data to obtain the difference between the means of that set. Each difference is then squared, and all squared values are summed. The mean differences are summed to obtain the sum of squared differences. This sum is then divided by the number of Euclidean distance data points, six, to obtain the variance of the Euclidean distance distribution. The six sets of calculated Euclidean distance data, the mean of the Euclidean distance distribution, and the variance of the Euclidean distance distribution are arranged in a preset order to form an eight-dimensional vertex deviation vector. This preset order follows the logic of individual deviations first, followed by overall characteristics. Specifically, the arrangement rules are as follows: first, the baseline anchor point and five health status vertices are arranged sequentially (based on traction machine bearing wear, brake shoe clearance, guide rail alignment, car vertical stability, and landing door opening and closing synchronization). The six sets of Euclidean distance data (corresponding to the category order of sex) are then arranged in order of Euclidean distance distribution mean and Euclidean distance distribution variance. The core purpose of this arrangement is to achieve hierarchical quantification of deviation features. The first six elements correspond to the direct deviation degree of each vertex relative to the standard vertex, which can accurately locate the abnormality of a single health status dimension; the last two elements reflect the comprehensive deviation features of all vertices from the perspectives of overall distribution central tendency and dispersion, respectively, which can intuitively evaluate the overall offset degree and stability of the polyhedron in the reference space. Each element of this vector reflects the deviation feature of the corresponding vertex in the reference space.
[0062] Step 402: Calculate the ratio of the intersection and union of the spatial volume occupied by the multidimensional health status representation polyhedron after mapping and the volume of the regular polyhedron to obtain the volume overlap ratio. Specifically, this includes: based on the coordinates of all vertices of the multidimensional health status representation polyhedron in the reference space obtained in step 400, the local volume is calculated using the spatial polyhedron volume piecewise calculation method (i.e., the convex polyhedron decomposition method). The specific operation and calculation process are as follows: First, polyhedron decomposition is performed. Using the reference anchor point as the common vertex, the multidimensional health status representation polyhedron in 64-dimensional space is decomposed into 5 independent 64-dimensional spatial simplexes. The vertices of each simplex are constructed according to the following rules: a reference anchor point, one healthy vertex, and two adjacent healthy vertices. The healthy vertices are numbered sequentially based on the traction machine bearing wear, brake shoe clearance, guide rail alignment, car vertical stability, and landing door opening / closing synchronization. The adjacent vertices of the first vertex are the last and second vertices, and the adjacent vertices of the last vertex are the fourth and first vertices. After polyhedral decomposition, the local volume of each 64-dimensional simplex is calculated using the 64-dimensional simplex volume formula. ,in It is the first The local volume of a 64-dimensional simplex; It is the factorial of 64, a fixed constant, with a value of 11.2688693218588417 × ; It is the first The determinant value of the 64th order square matrix corresponding to each simplex; It is the coordinate difference matrix of the i-th simplex. It is constructed by taking the coordinates of the 64 vertices (excluding the reference anchor point) inside the simplex, subtracting the corresponding coordinates of the reference anchor point from each of them, and obtaining 64 64-dimensional vectors. These vectors are then arranged in columns to form a 64-order square matrix. The matrix determinant is calculated using the 64-order determinant expansion rule.
[0063] After calculating the local volumes of the five simplexes, these local volumes are summed to obtain the spatial volume of the mapped multidimensional health status representation polyhedron. Based on historical data of elevator health status monitoring and the normalization results of feature vectors, the value range of this volume is [0.001, 1000], and the unit is feature space volume unit. At the same time, the predefined standard volume of the regular polyhedron is directly retrieved, with a value of 1.0 feature space volume unit. The calculation basis of this value is that the vertex coordinates of the regular polyhedron are all 64-dimensional normalized unit vectors, satisfying that the sum of the squares of the coordinates of each dimension is equal to 1, and the side length between adjacent vertices is 2. Using the same volume calculation method as the mapped polyhedron, a fixed value of 1.0 is finally obtained, which is the basic parameter of the reference space. Next, the hyperplane cutting method is used to calculate the spatial volume of the overlapping part of the mapped multidimensional health status representation polyhedron and the regular polyhedron. The specific steps are as follows: first, extract all the hyperplane equations of the two polyhedra respectively, where the hyperplane equation of the mapped polyhedron is... ( =1, 2, ..., M, where M is the number of hyperplanes of the polyhedron after mapping; For the first The first hyperplane equation The dimension coefficient, with a value range of [-1.0, 1.0], is set according to the distribution law of characteristic parameters of elevator health status monitoring and satisfies... The standardization constraint is 1; For the first 3D coordinate variables; For the first The constant term of the hyperplane equation, with a value range of [-0.5, 0.5] (determined based on the eigenvector normalization result), is: The hyperplane equation of the regular polyhedron is: ( =1, 2, ..., N, where N is the number of hyperplanes of a regular polyhedron; For the first The first hyperplane equation Dimension coefficient, taking values of ± =± This satisfies the standardization constraint of the coefficient vector, ensuring that the distance between the hyperplane and the vertices is uniform; For the first 3D coordinate variables; For the first The constant terms of the equations for the hyperplanes take values of = Then, using these hyperplane equations as constraints, solve the simultaneous equations to obtain the coordinates of all vertices in the overlapping region; finally, decompose the overlapping region into several 64-dimensional simplexes, calculate the volume of each simplex using the simplex volume formula mentioned above, and sum them up to obtain the volume of the overlapping part, which ranges from [0, 1.0], in characteristic space volume units; based on the principle of inclusion-exclusion of sets, use the obtained volume of the mapped polyhedron, the standard volume of the regular polyhedron, and the volume of the overlapping part to calculate the total volume of the space covered by the two polyhedra, with the formula: total volume = volume of the mapped polyhedron + standard volume of the regular polyhedron - volume of the overlapping part. Substituting the relevant values into the range, we can obtain the total volume range as [1.0, 1001.0], in characteristic space volume units. Finally, divide the volume of the overlapping part by the total volume to obtain the volume overlap ratio. Since the volume of the overlapping part is greater than or equal to 0 and less than or equal to the total volume, the value range of this ratio is between 0 and 1.
[0064] Step 403: Calculate the ratio of the absolute difference between the total surface area of the multidimensional health state representation polyhedron after mapping and the surface area of the regular polyhedron to the surface area of the regular polyhedron, obtaining the surface area difference coefficient. Specifically, this includes: based on the coordinates of all vertices of the multidimensional health state representation polyhedron in the reference space obtained in step 400, calculating the total surface area and related parameters of the polyhedron. The specific operation and calculation process are as follows: First, calculate the geometric area of each face of the polyhedron. Each face of the multidimensional health state representation polyhedron is a 63-dimensional hyperplane convex polygon (i.e., a 63-dimensional hyperface). Its geometric area calculation adopts the hyperface piecewise decomposition method. The specific steps are: for each face of the polyhedron, extract the 64-dimensional coordinates of all its vertices, and let... The vertex set of a certain face is Q1(x11, x12, ..., x164), Q2(x21, x22, ..., x264), ..., Qt(xt1, xt2, ..., xt64) (t is the number of vertices of the face, ranging from [5, 8], determined based on the geometric characteristics of a 64-dimensional polyhedron). Using any vertex in the face (such as Q1) as a common reference point, the 63-dimensional hyperface is decomposed into several 63-dimensional simplexes. The decomposition rule is that each simplex is formed by the reference point and 63 other non-coplanar vertices in the face, until the entire hyperface is covered. The volume of each decomposed simplex (i.e., the local area of that part of the hyperface) is calculated using the 63-dimensional simplex volume formula. The formula is... ,in, It is a 63rd order square matrix, constructed by taking the first... The coordinates of the 63 vertices (excluding the reference point) within a 63-dimensional simplex are subtracted from the corresponding coordinates of the reference point to obtain 63 63-dimensional vectors. These vectors are then arranged in columns to form a square matrix. The determinant of this square matrix is calculated using the 63rd order determinant expansion rule; It is the factorial of 63, and its value is 1.98260831540444 × The geometric area of the surface is obtained by summing up the local surfaces of all 63-dimensional simplexes after the surface is decomposed.
[0065] After calculating the area of a single face, the total surface area of the polyhedron is calculated. This process of calculating the area of a single face is repeated to sequentially calculate the geometric area of all faces of the multidimensional health state representation polyhedron. The areas of all faces are then summed to obtain the total surface area of the mapped multidimensional health state representation polyhedron. After obtaining the total surface area of the mapped polyhedron, the standard surface area of a regular polyhedron is retrieved. A predefined standard surface area value of 64.0 (in feature space units) is directly retrieved. This value is based on the number of vertices and side lengths of a 64-dimensional regular polyhedron (preset to be...). The surface area and its structural characteristics are calculated in advance using the same method as the surface area of the polyhedrons mentioned above, serving as the basic parameters for the reference space. Then, the absolute difference in surface area is calculated by subtracting the standard surface area of the regular polyhedron from the total surface area of the multidimensional health state characterization polyhedron after mapping, and taking the absolute value of the difference. Finally, the surface area difference coefficient is calculated by dividing the absolute difference in surface area by the standard surface area of the regular polyhedron.
[0066] This embodiment maps non-standard multidimensional health status representation polyhedra to a unified reference space, eliminating the differences in feature space caused by different elevator equipment models and operating environments, and realizing standardized comparative analysis of different elevator health statuses. The calculated vertex deviation vector can quantify the deviation characteristics of the actual polyhedron vertices relative to the standard vertices in the reference space from the perspective of the mean and variance of the distance distribution. The volume overlap ratio can intuitively reflect the degree of spatial overlap between the current health status polyhedron and the standard health status polyhedron. The higher the ratio, the better the overall health status of the elevator, and vice versa, indicating a significant abnormal trend. The surface area difference coefficient quantifies the difference between the actual state and the standard state from the perspective of polyhedron surface morphology, and can capture local morphological anomalies that cannot be reflected by volume parameters, further improving the evaluation dimensions of elevator health status.
[0067] In a preferred embodiment of the present invention, step 5 includes:
[0068] Step 500: Normalize the vertex deviation vector to obtain a normalized vertex deviation vector; weight the volume overlap ratio and surface area difference coefficient to obtain the geometric consistency index. Specifically, this includes: normalizing the vertex deviation vector using the L2 normalization method. The specific calculation process is as follows: first, calculate the magnitude of the vertex deviation vector, which is calculated as the square root of the sum of the squares of all components in the vector; then, divide each component of the vertex deviation vector by this magnitude to obtain the normalized result of each component; the normalized results of all components together constitute the normalized vertex deviation vector. This process ensures that all components of the vector are on a uniform numerical scale, eliminating the influence of differences in dimensions. For the geometric consistency index calculation, first, set the weights of the volume overlap ratio and the surface area difference coefficient. The weights are determined based on the importance analysis of the influence of geometric features on the consistency of multimodal data, with the weight of the volume overlap ratio set to 0. 6. The value of the volume overlap ratio reflects the degree of spatial matching between the multidimensional health status representation polyhedron and the regular polyhedron, and is a core indicator for measuring the consistency of the spatial structure of multimodal features, with a higher weight in the assessment of data reliability. The weight of the surface area difference coefficient is 0.4. This value is based on the fact that the surface area difference coefficient reflects the degree of matching of the boundary shape of the polyhedron, and its influence on the overall geometric consistency is lower than that of the spatial matching. Both weights are in the range of 0 to 1, and their sum is 1. Then, a weighted summation is performed. The calculation process is as follows: the volume overlap ratio is multiplied by a weight of 0.6 to obtain the weighted value of the volume overlap ratio; the surface area difference coefficient is multiplied by a weight of 0.4 to obtain the weighted value of the surface area difference coefficient. The two weighted values are added together, and the sum is the geometric consistency index. This index comprehensively reflects the overall consistency level between the multidimensional health status representation polyhedron and the regular polyhedron in terms of volume matching and surface area matching.
[0069] Step 501 involves multiplying each component of the normalized vertex deviation vector with the geometric consistency index to obtain a multimodal data reliability assessment vector reflecting the current data reliability of the vibration feature vector, acoustic feature vector, and visual feature vector. Specifically, this includes: first, clarifying the composition of the normalized vertex deviation vector, where each component corresponds one-to-one with the vibration feature vector, acoustic feature vector, and visual feature vector. The magnitude of each component directly reflects the degree of vertex deviation of the corresponding feature vector; a smaller component value indicates less deviation of the vertex from the standard position, and a higher basic reliability of the feature vector. Then, performing component multiplication is performed. This operation is based on the fact that the geometric consistency index comprehensively reflects the overall geometric matching level between the multidimensional health state representation polyhedron and the regular polyhedron, and is a core parameter for measuring the rationality of the multimodal feature space structure. Multiplying the components of the normalized vertex deviation vector with the geometric consistency index enables a coupled assessment of the basic deviation degree of the feature vector and the overall geometric structure rationality. The specific calculation process is as follows: The first component of the normalized vertex deviation vector is multiplied with the geometric consistency index. A higher geometric consistency index indicates a better match in the overall geometric structure and a stronger correction effect on the basic reliability of the vibration feature vector. The result is the evaluation value of the current data reliability of the vibration feature vector. The second component of the normalized vertex deviation vector is multiplied with the geometric consistency index to correct the basic deviation of the acoustic feature vector through the overall geometric consistency level. The result is the evaluation value of the current data reliability of the acoustic feature vector. The third component of the normalized vertex deviation vector is multiplied with the geometric consistency index to quantify the deviation impact of the visual feature vector in combination with the overall geometric matching. The result is the evaluation value of the current data reliability of the visual feature vector. The above three evaluation values are arranged in order and together constitute the multimodal data reliability evaluation vector. This vector, through the coupled calculation of the basic deviation degree and the overall geometric consistency, fully reflects the current data reliability level of each of the three feature vectors.
[0070] This embodiment effectively eliminates the dimensional differences among the components of the vertex deviation vector through normalization processing. Simultaneously, the geometric consistency index obtained through weighted summation integrates the spatial proportion matching characteristics reflected by the volume overlap ratio and the boundary morphology matching characteristics reflected by the surface area difference coefficient. This comprehensively and objectively reflects the consistency between the geometric model constructed from multimodal features and the standard model, providing a solid geometric feature basis for data reliability assessment. By multiplying the normalized vertex deviation vector with the geometric consistency index, the degree of vertex deviation and the level of geometric consistency are organically integrated. This accurately quantifies the current data credibility of the vibration feature vector, acoustic feature vector, and visual feature vector. The resulting multimodal data reliability assessment vector provides a clear reliability reference for subsequent multimodal data fusion processing, helping to screen high-credibility data and eliminate low-credibility interference information. This improves the accuracy and reliability of elevator health status assessment based on multimodal data, ensuring the scientific rigor and soundness of the assessment results.
[0071] In a preferred embodiment of the present invention, step 6 includes:
[0072] Step 600: Input the multimodal data reliability assessment vector into the cross-modal feature fusion layer as the initial bias weights for vibration feature vectors, acoustic feature vectors, and visual feature vectors in the multi-head attention mechanism calculation. This dynamically adjusts the fusion contribution ratios of the vibration feature vectors, acoustic feature vectors, and visual feature vectors, resulting in the adjusted fusion contribution ratios for the vibration feature vectors, acoustic feature vectors, and visual feature vectors. Specifically, this includes: first, clarifying the composition of the multimodal data reliability assessment vector, whose three components correspond one-to-one with the vibration feature vector, acoustic feature vector, and visual feature vector, respectively. The value of each component directly reflects the data reliability level of the corresponding feature vector; inputting this assessment vector into the cross-modal feature fusion layer, where the three components serve as the initial bias weights for the three feature vectors in the multi-head attention mechanism calculation; during the multi-head attention mechanism calculation, configuring three sets of learnable linear projection weight matrices (query weight matrix, key weight matrix, and value weight matrix) for each feature vector. All weight matrices are arranged in row-major order, and the elements in the matrix are generated using the Xavier initialization method, meaning that the element values follow a uniform distribution within the range [-]. , The initialization method is as follows: D is the feature vector dimension, and H is the hidden layer dimension of the attention mechanism. This initialization method can match the variance of the weight matrix with the input and output dimensions, ensuring signal stability during linear projection. The dimension of the weight matrix is determined based on the dimension of the feature vector and the hidden layer dimension of the attention mechanism. For example, if the feature vector dimension is D and the hidden layer dimension is H, then the dimensions of the query weight matrix and the key weight matrix are D×H, and the dimension of the value weight matrix is D×H. The preset bias term is a vector with the same dimension as the hidden layer dimension H, and all elements are 0.01. This value is based on the empirical rules for initializing deep learning model parameters to avoid the initial bias being too large and affecting the linear projection effect of the feature vector. The vibration feature vector is multiplied by the corresponding three sets of weight matrices, and then the preset bias term is added to obtain the query vector, key vector, and value vector corresponding to the vibration feature vector. Using the same calculation method, the acoustic feature vector and the visual feature vector are multiplied by the corresponding three sets of weight matrices and the bias term is added to obtain the query vector, key vector, and value vector corresponding to the acoustic feature vector and the visual feature vector, respectively.
[0073] Subsequently, attention scores are calculated based on the initial bias weights. A dot product similarity method is used, where the query vector of each feature vector is dot-producted with the key vectors of all three feature vectors, and then scaled by the square root of the key vector dimension to obtain the initial attention score for each feature vector relative to the others. The initial bias weights corresponding to the multimodal data reliability assessment vector are then added to the initial attention score of each feature vector, thus correcting the attention score based on data credibility. This strengthens the attention score of high-credibility feature vectors and suppresses the attention score of low-credibility feature vectors. Next, the corrected attention scores are subjected to softmax normalization. The corrected attention score for each feature vector is input into the softmax function, and the proportion of each score to the total attention scores of that feature vector is calculated. The attention weights of each feature vector relative to other feature vectors are obtained. These weights range from 0 to 1, and the sum of all attention weights is 1. Finally, the multi-head attention weights are averaged and fused. The number of heads in the multi-head attention mechanism is set to 8. An independent linear projection weight matrix is configured for each head (the configuration rules are the same as for single heads, including the arrangement, initialization method, and dimension setting). The above process of generating query vectors, key vectors, value vectors, and calculating attention weights is repeated to obtain 8 sets of attention weights. The same fusion weight is assigned to each set of attention weights, with a value of 0.125 (the sum of the fusion weights is 1, determined by the reciprocal of the number of heads). The 8 sets of attention weights are multiplied by their corresponding fusion weights and then added to obtain the final fusion contribution ratio of each feature vector. This ratio is dynamically matched with the data credibility level of the corresponding feature vector to complete the adjustment of the fusion contribution ratio.
[0074] Step 601 involves re-weighting and fusing the vibration feature vector, acoustic feature vector, and visual feature vector using the adjusted fusion contribution ratio to generate a new multimodal joint feature vector. Specifically, this includes: re-weighting and fusing the vibration feature vector, acoustic feature vector, and visual feature vector based on the adjusted fusion contribution ratio obtained in step 600. The specific calculation process is as follows: multiply each element in the vibration feature vector by its corresponding fusion contribution ratio to obtain the weighted feature of the vibration feature vector; multiply each element in the acoustic feature vector by its corresponding fusion contribution ratio to obtain the weighted feature of the acoustic feature vector; multiply each element in the visual feature vector by its corresponding fusion contribution ratio to obtain the weighted feature of the visual feature vector; and perform element-wise addition of the above three weighted features according to their corresponding dimensions, i.e., adding the three weighted feature elements of the same dimension to obtain the fused multimodal joint feature vector. This vector fully integrates the effective information of the three feature vectors, and the fusion weights are consistent with the data reliability of each feature vector.
[0075] Step 602a: Input the new multimodal joint feature vector into the fault diagnosis classifier. The fault diagnosis classifier is based on a classification model trained using vibration feature vectors, acoustic feature vectors, visual feature vectors, and fault type labels corresponding to the feature vectors collected during the elevator's historical operation. It outputs a fault probability vector containing multiple dimensions, where each dimension corresponds to one of the following fault modes: traction machine bearing wear, abnormal brake shoe clearance, guide rail alignment deviation, deterioration of car vertical stability, and inaccurate synchronization of landing door opening and closing. The fault mode corresponding to the dimension whose probability value exceeds a preset threshold in the fault probability vector is selected as the final fault mode identification result. Specifically, this includes: inputting the new multimodal joint feature vector generated in step 601 into the fault diagnosis classifier. The multimodal joint feature vector is input to the fault diagnosis classifier, which is built based on a deep learning network. The specific construction process is as follows: the network architecture includes an input layer, a feature extraction layer, a fully connected layer, an activation function layer, and an output layer. The input layer receives the multimodal joint feature vector. The feature extraction layer uses a three-layer convolutional neural network, each layer containing 64 convolutional kernels of size 3×3 with a stride of 1 and uniform padding, used to extract deep fault features from the joint features. The fully connected layer consists of two layers: the first fully connected layer has 256 neurons, and the second fully connected layer has 128 neurons. The activation function layer uses a linear rectified activation function to introduce non-linear mapping capabilities and improve performance. The model's feature representation capability is enhanced; the output layer has 5 neurons, corresponding to five preset fault modes; the classifier's training process is based on sample data collected during the elevator's historical operation. This sample data includes a large number of vibration feature vectors, acoustic feature vectors, visual feature vectors, and fault type labels corresponding to each set of feature vectors. The fault type labels cover five core elevator system fault modes: traction machine bearing wear (referring to the wear and lubrication failure of the balls and raceways in the traction machine's drive or non-drive end bearings due to long-term operation, leading to increased rotational resistance and intensified vibration), and abnormal brake shoe clearance (referring to the clearance between the brake shoe and brake wheel exceeding the preset standard range of 0.1 mm to 0.3 mm). This can lead to several problems, including: delayed braking response, insufficient braking force, or abnormal braking noise; guide rail misalignment (meaning the installation position of the elevator car guide rail or counterweight guide rail deviates from the design benchmark, with a parallelism deviation exceeding 0.5 mm / m and a verticality deviation exceeding 0.3 mm / m, resulting in increased friction and vibration during car operation); deterioration of vertical stability of car operation (meaning that due to uneven tension of the traction steel wire rope, failure of the vibration damper, or poor lubrication of the guide rail, the horizontal vibration acceleration of the car during vertical operation exceeds 0.15 m / s², resulting in bumps, swaying, and stability that does not meet the operating standards); and inaccurate synchronization of landing door opening and closing (meaning that due to a failure of the linkage mechanism, door lock device, or drive motor of the elevator landing door, the time difference between the opening and closing of the left and right doors exceeds 0.(A fault that causes asynchronous door movement, abnormal opening and closing speed, or incomplete closure within 2 seconds) is identified. One-hot encoding is used to label the fault type. During training, the cross-entropy loss function is used to calculate the error between the predicted value and the true label. An adaptive moment estimation optimizer is used to iteratively update the network parameters until the model's diagnostic accuracy on the validation set no longer improves, thus completing model training.
[0076] The multimodal joint feature vector enters the feature extraction layer through the input layer, where deep fault features are extracted through convolution and pooling operations. These deep fault features are then input to the first fully connected layer, where matrix multiplication and ReLU activation are performed to achieve feature dimension transformation and nonlinear mapping. The transformed features are then input to the second fully connected layer, where they undergo matrix multiplication and ReLU activation again to further optimize the feature representation. Finally, the output of the second fully connected layer is input to the output layer, where the Softmax activation function is used to map the output values to a probability distribution. The final output is a fault probability vector with five dimensions, each corresponding to a preset fault mode; for example, the first dimension corresponds to traction. The five fault modes are: bearing wear, brake shoe clearance abnormality, guide rail alignment deviation, car vertical stability deterioration, and landing door opening / closing synchronization inaccuracy. The values for each dimension represent the probability that the current elevator operation belongs to that fault mode. A preset threshold of 0.7 is set. This threshold is determined based on the accuracy verification results of historical fault diagnosis data, achieving an optimal balance between fault identification accuracy and false negative rate. Fault modes corresponding to dimensions with values exceeding 0.7 in the fault probability vector are selected as the final fault mode identification result. If the values of all dimensions do not exceed 0.7, it is determined that the elevator currently does not have any of the above five fault modes.
[0077] Step 602b involves inputting the new multimodal joint feature vector into the remaining useful life regressor. This regressor is based on a regression model trained using the multimodal joint feature vectors collected during the elevator's historical operation and the corresponding remaining useful life labels. It calculates and outputs a remaining useful life value in days as a lifespan prediction. Specifically, this includes inputting the new multimodal joint feature vector generated in step 601 into the remaining useful life regressor. This regressor is built using a deep learning network, specifically with an input layer, feature extraction layer, regression prediction layer, and output layer. The input layer receives the multimodal joint feature vector, and the vector dimension is the same as the joint feature dimension output in step 601. To maintain consistency and ensure compatibility of feature transmission, the feature extraction layer employs a two-layer Long Short-Term Memory (LSTM) network, with each layer containing 128 hidden units. A random inactivation rate of 0.2 is set between the two layers to suppress overfitting. Simultaneously, the LSM network's gating mechanism accurately captures the temporal dependencies in joint features, such as the gradual changes in vibration intensity, noise frequency, and component wear during elevator operation. The regression prediction layer consists of two fully connected layers: the first fully connected layer has 64 neurons, and the second fully connected layer has 32 neurons. This progressively compresses feature dimensions and strengthens the mapping between features and remaining lifetime. The output layer is a single neuron, employing a linear activation function (mathematically expressed as...). =wx+b, where w is the output layer weight vector with a dimension of 32×1. The initial value is generated using the Xavier initialization method, follows a uniform distribution, and takes values in the range [−]. , ], i.e. [-0.436, 0.436], is iteratively updated through the adaptive moment estimator during training; b is the output layer bias value, initially set to 0.01, and iteratively optimized synchronously during training; x is the 32-dimensional optimized feature vector from input to output layer, with each element taking values in the range of [0, 1] (because the features have been preprocessed by standardization). The output of the linear activation function is the intermediate value of the remaining life prediction after logarithmic transformation (the value range converges to the interval [-3, 5] as the training process progresses). It directly outputs the remaining life prediction value in days.
[0078] The regressor training process is based on sample data collected during the elevator's historical operation. Specifically, the sample data comes from elevators of different models and years of operation, covering the entire lifecycle states including normal operation, minor aging, moderate failure, and near-retirement. The sample data includes a large number of multimodal joint feature vectors and a remaining life label corresponding to each joint feature vector. The remaining life label represents the actual number of days the elevator operated from the time the joint feature vector was collected until an irreversible failure or mandatory retirement standard was reached. The label data was manually verified and cross-validated with equipment operation logs to ensure accuracy. The multimodal joint feature vectors were standardized, mapping all feature dimensions to the [0, 1] interval to eliminate the impact of dimensional differences on model training. The remaining life label underwent a natural logarithmic transformation (specifically, y'=ln( +1), where y' is the original remaining life label, with a value range of [1, 3650] days (covering the common 10-year lifespan of elevators). y' is the transformed label value, with a value range of [ln(1+1), ln(3650+1)], i.e., [0, 8.2]. Adding 1 is to avoid the original label... When the logarithmic value is 0, logarithmic operations are meaningless. This transformation can effectively compress the dynamic range of label values, reduce the interference of extreme large values on model training, and make the model more likely to learn the mapping relationship between features and remaining lifetime. During training, the mean squared error loss function is used to calculate the error between the predicted value and the true remaining lifetime label. The loss function is calculated as the average of the squares of the differences between the predicted value and the true label. An adaptive moment estimation optimizer is used to iteratively update the network parameters. The initial learning rate of the optimizer is set to 0.001, and the learning rate adopts an exponential decay strategy, decaying to 0.9 times the current value every 10 iterations. The total number of iterations is set to 50. At the same time, an early stopping mechanism is introduced. If the mean absolute error on the validation set does not decrease for 5 consecutive iterations, the training is terminated to avoid model overfitting. Finally, the model training is completed.
[0079] The multimodal joint feature vector enters the feature extraction layer through the input layer. The first-layer long short-term memory network processes features collaboratively through a three-gating mechanism. The input gate combines the current feature vector with the result of the previous hidden state to filter new feature information and incorporate it into the current cell state. The forget gate outputs a weight value in the range of 0 to 1 through the Sigmoid activation function to filter out meaningless historical features such as transient anomalies. The cell state is dynamically updated based on the effective information from the input gate and the filtering result from the forget gate, retaining core temporal features (i.e., key features that continuously change over time during elevator operation and are strongly correlated with the remaining lifespan, such as the slow increasing trend of vibration peak with running time, the gradual increase in the proportion of abnormal frequency components in operating noise, the continuous cumulative effect of temperature rise in traction machine bearings, and the gradual increase in the coefficient of friction of the guide rail). The output gate generates the preliminary extracted features of the first-layer network based on the current cell state and the previous hidden state. This feature is fed into the second-layer long short-term memory network, and the above gated collaborative operation is repeated to further deepen the temporal features. The process involves several steps: First, a 128-dimensional core feature vector is extracted and output (the dimension matches the number of hidden units in the Long Short-Term Memory network), accurately depicting the aging state and fault development trend of elevators. Then, the core feature vector is input into the first fully connected layer of the regression prediction layer. A matrix multiplication operation is performed between the core feature vector and the weight matrix of this layer (the weight matrix has a dimension of 128×64), followed by the addition of a 64-dimensional bias vector (with initial values of 0.01 for each bias vector), resulting in a 64-dimensional intermediate feature vector. This intermediate feature vector is then processed by a linear rectified activation function, setting values less than 0 to 0 while retaining effective feature information. This achieves feature dimension compression and nonlinear transformation. The transformed 64-dimensional feature vector is then input into the second fully connected layer, where it undergoes matrix multiplication with the weight matrix (dimension 64×32), followed by the addition of a 32-dimensional bias vector (with initial values of 0.01 for each bias vector), resulting in a 32-dimensional optimized feature vector. This optimized feature vector is then processed again by a linear rectified activation function to further strengthen the nonlinear mapping relationship between features and remaining lifespan.
[0080] Finally, the 32-dimensional optimized feature vector output from the second fully connected layer is input into the output layer and multiplied by matrix multiplication with the output layer's weight vector w (32×1, converged to the interval [-0.3, 0.3] after training). A single bias value b (converged to the interval [-0.1, 0.1] after training) is then added, and a continuous numerical value Y (i.e., the label scale after logarithmic transformation, with a value range of [-3, 5]) is calculated using a linear activation function. This value is then subjected to an inverse natural exponential transformation (specifically, Y = ...). -1, where The output layer's calculation result has a value range of [-3, 5]; e is the natural constant; Y is the actual remaining useful life value after restoration, with a value range of [ 1, -1], that is, [0, 147] days, which is consistent with the actual scenario of the remaining life before the elevator failure), and finally the remaining life prediction value in days is obtained.
[0081] Step 602c involves combining the fault mode identification results with the remaining service life value to form a state assessment tuple containing specific fault types and quantified remaining service life. Specifically, this includes combining the fault mode identification results obtained in step 602a with the remaining service life value obtained in step 602b to form a state assessment tuple. The structure of this tuple is (fault mode identification result, remaining service life value). The fault mode identification result clearly identifies the specific fault types currently existing in the elevator, including five core system faults: traction machine bearing wear, abnormal brake shoe clearance, guide rail alignment deviation, deterioration of car running vertical stability, and inaccurate door opening and closing synchronization. The identification result may be a single fault mode or a combination of multiple fault modes. The remaining service life value quantifies the remaining operating days of the elevator in its current state, reflecting the elevator's remaining operating time from its current operating state until an irreversible fault occurs or it reaches the mandatory scrapping standard. The tuple fully presents the elevator's fault state and lifespan state.
[0082] Step 603: Based on the fault mode identification results and lifespan prediction values, and combined with predefined fault level and maintenance measure mapping rules, dynamically generate a maintenance decision work order containing early warning levels, fault location information, and maintenance suggestions. Specifically, this includes: First, predefining fault level and maintenance measure mapping rules. These rules are comprehensively formulated based on the severity, impact range, and development trend of the fault mode, combined with the interval division of the remaining service life value. Specifically, the rules are as follows: First, clarifying the remaining service life value intervals, based on the elevator's average maintenance cycle and fault development speed, dividing the remaining service life into three intervals: [0, 30 days], [31, 90 days], and [over 91 days]. Second, classifying the fault severity levels, specifically into three levels: Minor faults, affecting only a single component, with no safety hazards and a slow development trend, including guide rail alignment deviation and door opening / closing synchronization inaccuracy; and Moderate faults, affecting a local system, with potential safety hazards and a moderate development trend. Specifically, the issues include: deterioration of the vertical stability of the car operation; serious malfunctions affecting the core system, posing direct safety hazards, and rapidly developing, encompassing two types of malfunctions: traction machine bearing wear and abnormal brake shoe clearance; and the establishment of a gradient correspondence between warning levels, with warning levels divided into four levels from high to low: Level 1 (emergency), Level 2 (high risk), Level 3 (general), and Level 4 (advancement). The correspondence between each warning level and the fault type and remaining service life range is clear: Level 1 warning corresponds to serious malfunctions with a remaining service life of [0, 30 days]; Level 2 warning corresponds to serious malfunctions with a remaining service life of [31, 90 days] and moderate malfunctions with a remaining service life of [0, 30 days]; Level 3 warning corresponds to moderate malfunctions with a remaining service life of [31, 90 days] and minor malfunctions with a remaining service life of [0, 30 days]; and Level 4 warning corresponds to minor malfunctions with a remaining service life of [31, 90 days] and all fault types with a remaining service life of [more than 91 days].
[0083] Fourth, refine fault location information, specifying the exact location for each type of fault. For example, traction machine bearing wear is located in the elevator traction machine bearing assembly, specifically at the connection between the traction machine motor output shaft and the gearbox; abnormal brake shoe clearance is located in the elevator brake shoe and brake wheel clearance adjustment mechanism, specifically at the connection between the brake arm and the brake shoe; guide rail alignment deviation is located in the elevator guide rail installation and fixing structure, specifically at the connection between the guide rail bracket and the guide rail body; deterioration of car running vertical stability is located in the elevator car suspension system and guide mechanism, specifically at the fixed end of the car guide rail slider and the traction steel wire rope; misalignment of landing door opening and closing synchronization is located in the elevator landing door transmission mechanism, specifically at the connection between the landing door linkage rod and the door lock device. Fifth, develop supporting maintenance recommendations, covering four aspects: maintenance timing, maintenance process, required tools, and required spare parts. Regarding maintenance timing, Level 1 warnings must be handled within 24 hours, Level 2 warnings within 72 hours, and Level 3 warnings within 7 days. Level 1 and 2 warnings must be handled within 30 days. In terms of maintenance procedures, Level 1 and 2 warnings require immediate shutdown for inspection, disassembly of faulty components for testing or replacement, and then system debugging and trial operation. Level 3 and 4 warnings can be addressed during normal elevator downtime, without requiring prolonged shutdown. Required tools include torque wrenches, guide rail straighteners, bearing testers, clearance measuring rulers, and elevator debugging instruments. Required spare parts include dedicated spare parts for the corresponding fault modes, such as traction machine bearings, brake shoes, guide rail straightening shims, car sliders, and landing door linkage rods. In practical application, the fault mode identification results obtained in step 602a and the remaining service life prediction values obtained in step 602b are input into the above mapping rules. The system will automatically match the corresponding warning level, fault location information, and maintenance suggestions. Subsequently, it integrates all information according to a preset format and dynamically generates a maintenance decision work order containing the warning level, fault location information, and maintenance suggestions, providing clear and executable operational guidelines for elevator maintenance work.
[0084] This embodiment achieves dynamic adjustment of the fusion contribution ratio by using the multimodal data reliability assessment vector as the initial bias weight. This enables the feature fusion process to accurately match the credibility of each modality of data, effectively reducing the interference of low-credibility data on the fusion result and improving the information effectiveness and accuracy of the multimodal joint feature vector. The reweighted fusion method integrates the effective information of the three feature vectors based on the dynamically adjusted contribution ratio, enabling the joint feature vector to comprehensively and prominently reflect the key features of the elevator's operating status. Both the fault diagnosis classifier and the remaining service life regressor are trained based on historical elevator operating sample data, enabling them to fully learn features and fault modes, remaining service life, and other relevant parameters. The mapping relationship between lifespan ensures the accuracy of fault mode identification results and the reliability of remaining service life prediction values, enabling a comprehensive quantitative assessment of elevator status. Status assessment tuples combine specific fault types with quantified remaining service life, making elevator status information more intuitive and complete, facilitating quick understanding of the elevator's core status by staff. Maintenance decision work orders are dynamically generated based on predefined mapping rules, transforming fault diagnosis and lifespan prediction results into directly executable maintenance plans. These plans clearly define warning levels, fault locations, and maintenance recommendations, achieving seamless integration from status assessment to maintenance execution, reducing the risk of fault escalation, extending elevator lifespan, and improving the safety and stability of elevator operation.
[0085] like Figure 2 As shown, embodiments of the present invention also provide an intelligent elevator operation and maintenance management system based on a multimodal neural network, including:
[0086] The acquisition module is used to acquire multi-source time-series monitoring data and construct a multimodal neural network. The multi-source time-series monitoring data is input into the multimodal neural network to obtain vibration feature vectors, acoustic feature vectors and visual feature vectors.
[0087] The fusion module is used to input the vibration feature vector, acoustic feature vector and visual feature vector into the cross-modal feature fusion layer in the multimodal neural network for fusion to generate a multimodal joint feature vector;
[0088] The module is used to calculate a steady-state reference feature vector in the feature space based on the multimodal joint feature vector as the benchmark feature anchor point, and to define five key health status characterization vectors: traction machine bearing wear, brake shoe clearance, guide rail alignment, car vertical stability, and landing door opening and closing synchronization. Using the benchmark feature anchor point as the vertex, the five key health status characterization vectors are synthesized as edge vectors to construct a multidimensional health status characterization polyhedron.
[0089] The quantization module is used to map the multidimensional health status representation polyhedron to a reference space defined by a preset regular polyhedron, and to calculate the vertex deviation, volume overlap ratio and surface area difference coefficient.
[0090] The evaluation module is used to calculate the reliability evaluation vector of multimodal data based on the vertex deviation vector, volume overlap ratio, and surface area difference coefficient.
[0091] The decision module is used to dynamically adjust the feature weights in cross-modal fusion using multimodal data reliability assessment vectors, and to make fault diagnosis and maintenance decisions based on the fusion features generated after weight adjustment.
[0092] It should be noted that this system is a system corresponding to the above method. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effect.
[0093] Embodiments of the present invention also provide a computer-readable storage medium storing instructions that, when executed on a computer, cause the computer to perform the method described above. All implementations in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0094] The above description represents the preferred embodiments of the present invention. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.
Claims
1. An intelligent elevator operation and maintenance management method based on multimodal neural networks, characterized in that, The method includes: Step 1: Collect multi-source time-series monitoring data and construct a multimodal neural network. Input the multi-source time-series monitoring data into the multimodal neural network to obtain vibration feature vector, acoustic feature vector and visual feature vector; Step 2: Input the vibration feature vector, acoustic feature vector and visual feature vector into the cross-modal feature fusion layer in the multimodal neural network for fusion to generate a multimodal joint feature vector; Step 3: Based on the multimodal joint feature vector, a steady-state reference feature vector is calculated in the feature space as the baseline feature anchor point. Based on the preset feature dimension-health state association rule, from the difference between the current multimodal joint feature vector and the steady-state reference feature vector, the following five key health state representation vectors are extracted: a first state vector representing the wear degree of the traction machine bearing, a second state vector representing the brake shoe clearance, a third state vector representing the guide rail alignment, a fourth state vector representing the vertical stability of the car operation, and a fifth state vector representing the synchronicity of the landing door opening and closing. These five key health state representation vectors are then synthesized using the baseline feature anchor point as the vertex and the five key health state representation vectors as edge vectors to construct a multidimensional health state representation polyhedron. Step 4: Map the multidimensional health status representation polyhedron to a reference space defined by a pre-defined regular polyhedron, and calculate the vertex deviation, volume overlap ratio, and surface area difference coefficient. Step 5: Calculate the multimodal data reliability assessment vector based on the vertex deviation vector, volume overlap ratio, and surface area difference coefficient. Step 6: Dynamically adjust the feature weights in cross-modal fusion using the multimodal data reliability assessment vector, and make fault diagnosis and maintenance decisions based on the fusion features generated after weight adjustment.
2. The elevator intelligent operation and maintenance management method based on multimodal neural networks according to claim 1, characterized in that, Step 1 includes: Multi-source time-series monitoring data are collected synchronously by a heterogeneous sensor array deployed in key parts of the elevator. The heterogeneous sensor array includes a vibration acceleration sensor installed in the traction machine bearing housing, a broadband acoustic sensor installed on the top of the car, and an image sensing unit deployed in the hoistway and landing door areas. A multimodal neural network is constructed, comprising a parallel vibration feature extraction network, an acoustic feature extraction network, and a visual feature extraction network. Time-series data collected by a vibration acceleration sensor is input into the vibration feature extraction network, which outputs a vibration feature vector. Time-series data collected by a broadband acoustic sensor is input into the acoustic feature extraction network, which outputs an acoustic feature vector. Image sequences collected by an image sensing unit are input into the visual feature extraction network, which outputs a visual feature vector.
3. The elevator intelligent operation and maintenance management method based on multimodal neural networks according to claim 2, characterized in that, Step 2 includes: Vibration feature vector, acoustic feature vector and visual feature vector are simultaneously input into the cross-modal feature fusion layer. In the cross-modal feature fusion layer, the pairwise correlation weights between vibration feature vector, acoustic feature vector and visual feature vector are calculated by a multi-head attention mechanism. Based on the pairwise correlation weights, the vibration feature vector, acoustic feature vector, and visual feature vector are weighted, concatenated, and linearly transformed to generate a unified multimodal joint feature vector.
4. The elevator intelligent operation and maintenance management method based on multimodal neural networks according to claim 3, characterized in that, Step 3 includes: Based on a predefined sliding time window, the mean of multiple multimodal joint feature vectors generated continuously in history within the window is calculated to obtain a steady-state reference feature vector that represents the typical operating state within the window, and the steady-state reference feature vector is used as the benchmark feature anchor point. From the multimodal joint feature vector at the current moment, extract the first state vector representing the wear of the traction machine bearing, the second state vector representing the brake shoe clearance, the third state vector representing the guide rail alignment, the fourth state vector representing the vertical stability of the car operation, and the fifth state vector representing the synchronicity of the landing door opening and closing, as five key health state representation vectors. Using the steady-state reference feature vector as the spatial starting point, the first state vector, the second state vector, the third state vector, the fourth state vector, and the fifth state vector are used as spatial edge vectors. The five vertices of the multidimensional health state representation polyhedron are determined sequentially by vector addition. The steady-state reference feature vector is sequentially connected to the five vertices of the multidimensional health state representation polyhedron, and edges are sequentially connected between the five vertices to construct a closed multidimensional health state representation polyhedron.
5. The elevator intelligent operation and maintenance management method based on multimodal neural networks according to claim 4, characterized in that, Step 4 includes: Perform a spatial affine coordinate transformation on each vertex of the multidimensional health state representation polyhedron to map the multidimensional health state representation polyhedron to a reference space defined by a predefined regular polyhedron; In the reference space, calculate the Euclidean distance from each vertex of the multidimensional health state characterization polyhedron to the corresponding vertex of the regular polyhedron. Based on all Euclidean distances, calculate the distribution variance and mean of the Euclidean distances, and generate the vertex deviation vector based on the distribution variance and mean. The ratio of the intersection and union of the spatial volume occupied by the multidimensional health status representation polyhedron after mapping to the volume of the regular polyhedron is calculated to obtain the volume overlap ratio. The surface area difference coefficient is obtained by calculating the ratio of the absolute difference between the total surface area after the multidimensional health status representation polyhedron mapping and the surface area of the regular polyhedron to the surface area of the regular polyhedron.
6. The elevator intelligent operation and maintenance management method based on multimodal neural networks according to claim 5, characterized in that, Step 5 includes: The vertex deviation vector is normalized to obtain the normalized vertex deviation vector; the volume overlap ratio and the surface area difference coefficient are weighted and summed to obtain the geometric consistency index. Multiplying each component of the normalized vertex deviation vector with the geometric consistency index yields a multimodal data reliability assessment vector that reflects the current data credibility of the vibration feature vector, acoustic feature vector, and visual feature vector.
7. The elevator intelligent operation and maintenance management method based on multimodal neural networks according to claim 6, characterized in that, Step 6 includes: The multimodal data reliability assessment vector is input into the cross-modal feature fusion layer as the initial bias weights of vibration feature vector, acoustic feature vector and visual feature vector in the multi-head attention mechanism calculation, so as to dynamically adjust the fusion contribution ratio of vibration feature vector, acoustic feature vector and visual feature vector, and obtain the adjusted fusion contribution ratio of vibration feature vector, acoustic feature vector and visual feature vector. The vibration feature vector, acoustic feature vector, and visual feature vector are reweighted and fused using the adjusted fusion contribution ratio to generate a new multimodal joint feature vector. The new multimodal joint feature vector is input into a preset fault diagnosis classifier to obtain the fault mode recognition result, and the new multimodal joint feature vector is input into a preset remaining service life regressor to obtain the service life prediction value. Based on the failure mode identification results and life prediction values, and combined with predefined failure level and maintenance measure mapping rules, a maintenance decision work order containing early warning level, failure location information and maintenance suggestions is dynamically generated.
8. The elevator intelligent operation and maintenance management method based on multimodal neural networks according to claim 7, characterized in that, The new multimodal joint feature vector is input into a preset fault diagnosis classifier to obtain fault mode recognition results, and the new multimodal joint feature vector is input into a preset remaining useful life regressor to obtain life prediction values, including: The new multimodal joint feature vector is input into the fault diagnosis classifier. The fault diagnosis classifier is based on a classification model trained by collecting vibration feature vectors, acoustic feature vectors, visual feature vectors and fault type labels corresponding to the feature vectors during the elevator's historical operation. It outputs a fault probability vector containing multiple dimensions, where each dimension corresponds to one of the following fault modes: traction machine bearing wear, abnormal brake shoe clearance, guide rail centering deviation, deterioration of vertical stability of car operation, and inaccuracy of landing door opening and closing synchronization. The fault mode corresponding to the dimension whose probability value in the fault probability vector exceeds a preset threshold is selected as the final fault mode identification result. The new multimodal joint feature vector is input into the remaining useful life regressor. The remaining useful life regressor is a regression model trained based on the multimodal joint feature vector collected during the elevator's historical operation and the remaining useful life label corresponding to the multimodal joint feature vector. It calculates and outputs a remaining useful life value in days as a life prediction value. The failure mode identification results are combined with the remaining useful life value to form a state assessment tuple containing the specific failure type and the quantified remaining useful life.
9. An intelligent elevator operation and maintenance management system based on a multimodal neural network, wherein the system implements the method as described in any one of claims 1 to 8, characterized in that, include: The acquisition module is used to acquire multi-source time-series monitoring data and construct a multimodal neural network. The multi-source time-series monitoring data is input into the multimodal neural network to obtain vibration feature vectors, acoustic feature vectors and visual feature vectors. The fusion module is used to input the vibration feature vector, acoustic feature vector and visual feature vector into the cross-modal feature fusion layer in the multimodal neural network for fusion to generate a multimodal joint feature vector; The module is used to calculate a steady-state reference feature vector in the feature space based on the multimodal joint feature vector as the benchmark feature anchor point, and to define five key health status characterization vectors: traction machine bearing wear, brake shoe clearance, guide rail alignment, car vertical stability, and landing door opening and closing synchronization. Using the benchmark feature anchor point as the vertex, the five key health status characterization vectors are synthesized as edge vectors to construct a multidimensional health status characterization polyhedron. The quantization module is used to map the multidimensional health status representation polyhedron to a reference space defined by a preset regular polyhedron, and to calculate the vertex deviation, volume overlap ratio and surface area difference coefficient. The evaluation module is used to calculate the reliability evaluation vector of multimodal data based on the vertex deviation vector, volume overlap ratio, and surface area difference coefficient. The decision module is used to dynamically adjust the feature weights in cross-modal fusion using multimodal data reliability assessment vectors, and to make fault diagnosis and maintenance decisions based on the fusion features generated after weight adjustment.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that, when executed by a processor, implements the method as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Elevator fault mode recognition system based on big data
CN115783923A
Noise cause analysis system and operation method after elevator installation based on AI
KR102776631B1