Elevator state monitoring method and device and computer readable storage medium

By combining video and audio data collected by the elevator's built-in camera, vibration and abnormal noise analysis is performed, and a deep learning model is used to assess the elevator's health status. This solves the problem of multi-angle synchronous monitoring in traditional elevator monitoring, improving monitoring accuracy and reducing costs.

CN120829099APending Publication Date: 2025-10-24BEIJING URBAN CONSTR INTELLIGENT CONTROL TECH CO LTD

Patent Information

Application Number
CN202511202346.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-10-24

AI Technical Summary

Technical Problem

Traditional elevator health status monitoring relies on a single sensor, making it difficult to simultaneously monitor faults from multiple angles, resulting in inaccurate monitoring results and additional hardware costs.

Method used

Video and audio data are acquired using image acquisition equipment, time-series data are generated through vibration and abnormal noise analysis, weighted fusion is performed using an attention mechanism, and a deep learning model is used for state assessment to generate a time-series health score.

Benefits of technology

This technology improves the accuracy and efficiency of elevator health status monitoring without requiring additional hardware costs, avoiding the shortcomings of monitoring caused by a single sensor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120829099A_ABST
    Figure CN120829099A_ABST
Patent Text Reader

Abstract

The invention discloses an elevator state monitoring method and device and a computer readable storage medium. The method comprises the steps that video data and audio data collected by image collection equipment in the elevator running process are obtained; performing vibration analysis on the video data to obtain time sequence vibration data, and performing abnormal sound analysis on the audio data to obtain time sequence abnormal sound data; performing weighted fusion on the time sequence vibration data and the time sequence abnormal sound data by using an attention mechanism to obtain time sequence fusion data; and the time sequence fusion data are input into a state monitoring model, the state monitoring model is used for conducting state evaluation on the time sequence fusion data, and a time sequence health score of the elevator in the running process is obtained. The technical problems that in the related technology, due to the fact that a single sensor is adopted, it is difficult to synchronously monitor possible faults of an elevator from multiple angles, the monitoring result is inaccurate, and extra hardware cost is likely to be increased are solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent monitoring, in particular to a method and device for monitoring the state of an elevator, and a computer readable storage medium. BACKGROUND

[0002] Traditional methods for monitoring the health state of an elevator mainly rely on special sensors installed on the elevator, such as accelerometers, vibration sensors, noise sensors, etc. These sensors can monitor abnormal signals such as vibration and abnormal noise during the operation of the elevator. However, the installation of additional sensors significantly increases the hardware cost and the cost of subsequent maintenance, and the data collected by a single sensor only reflects the state of the elevator in one aspect, which may lead to inaccurate monitoring results.

[0003] To address the above technical problems in the related art, i.e., the single sensor cannot synchronously monitor possible faults of the elevator from multiple angles, leading to inaccurate monitoring results and increasing the cost of additional hardware, no effective solutions have been proposed. SUMMARY

[0004] The embodiments of the present application provide a method and device for monitoring the state of an elevator, and a computer readable storage medium, to at least solve the technical problem in the related art that the use of a single sensor cannot synchronously monitor possible faults of the elevator from multiple angles, leading to inaccurate monitoring results and increasing the cost of additional hardware.

[0005] According to an aspect of the embodiments of the present application, a method for monitoring the state of an elevator is provided, including: acquiring video data and audio data collected by an image acquisition device during the operation of the elevator; performing vibration analysis on the video data to obtain time-series vibration data, and performing abnormal noise analysis on the audio data to obtain time-series abnormal noise data, wherein the time-series vibration data includes vibration data of each timestamp during the operation of the elevator, and the time-series abnormal noise data includes abnormal noise data of each timestamp during the operation of the elevator; using an attention mechanism to perform weighted fusion on the time-series vibration data and the time-series abnormal noise data to obtain time-series fusion data; inputting the time-series fusion data into a state monitoring model to perform state evaluation on the time-series fusion data using the state monitoring model to obtain a time-series health score of the elevator during the operation, wherein the state monitoring model is obtained by training a plurality of sets of first training data using deep learning, and each set of the plurality of sets of first training data includes historical time-series fusion data and a historical time-series health score corresponding to the historical time-series fusion data, and the time-series health score includes a health score representing the health state of the elevator at each timestamp during the operation.

[0006] Optionally, the video data is subjected to vibration analysis to obtain time-series vibration data, including: intercepting video frames corresponding to a high-fault area in the video data to obtain target video data, wherein the high-fault area is at least one of the following: an elevator car area, a guide rail area; calculating motion vectors of the elevator in each timestamp in the target video data using an inter-frame optical flow algorithm to obtain a plurality of motion vectors, wherein each motion vector represents the motion state of the elevator in the corresponding timestamp; integrating a plurality of motion vectors in the time sequence of the target video data to obtain time-series data; and performing spectral analysis on the time-series data to quantify the motion vectors to obtain the time-series vibration data.

[0007] Optionally, the spectral analysis of the time-series data to quantify the motion vectors to obtain the time-series vibration data includes: performing Fourier transform or time-series wavelet transform on the time-series data to obtain a video spectrum graph; determining a peak value in the video spectrum graph as a vibration frequency of the elevator at each timestamp; determining a spectrum value or wavelet coefficient corresponding to the vibration frequency as a vibration amplitude of the elevator at the corresponding timestamp; determining the vibration frequency and the vibration amplitude of each timestamp as vibration data corresponding to the timestamp; and replacing the motion vector corresponding to the timestamp in the time-series data with the vibration data to obtain the time-series vibration data.

[0008] Optionally, after determining the vibration frequency and the vibration amplitude of each timestamp as vibration data corresponding to the timestamp, the method of monitoring the state of the elevator further includes: adding an abnormal vibration identifier to the vibration data with a vibration frequency greater than a preset frequency and / or a vibration amplitude greater than a preset amplitude to represent that the vibration data is abnormal vibration data; and adding a normal vibration identifier to the vibration data with a vibration frequency not greater than a preset frequency and a vibration amplitude not greater than a preset amplitude to represent that the vibration data is normal vibration data.

[0009] Optionally, the audio data is subjected to abnormal sound analysis to obtain time-series abnormal sound data, including: mapping the audio data to a frequency spectrum arranged according to a mel frequency scale to obtain an audio spectrum graph, wherein the mel frequency scale is a nonlinear frequency scale, and the audio spectrum graph is used to represent the audio frequency of the audio data at each timestamp; performing feature extraction on the audio spectrum graph to obtain a plurality of audio feature data; performing abnormal sound identification on each audio feature data using an abnormal sound identification model to obtain an abnormal sound type of each audio feature data, wherein the abnormal sound identification model is obtained by training a convolutional neural network using a plurality of second training data, and each of the plurality of second training data includes historical audio feature data and a historical abnormal sound type corresponding to the historical audio feature data; adding the abnormal sound type as an abnormal sound identifier to the audio feature data to obtain the abnormal sound data corresponding to the timestamp; and integrating the abnormal sound data according to the time sequence of the audio data to obtain the time-series abnormal sound data.

[0010] Optionally, the time-series vibration data and the time-series abnormal sound data are weighted and fused using an attention mechanism to obtain time-series fusion data, including: determining a first weight value and a second weight value corresponding to each timestamp using the attention mechanism, wherein the first weight value is a weight coefficient of the vibration data within the timestamp, and the second weight value is a weight coefficient of the abnormal sound data within the timestamp; weighting and fusing the vibration data and the abnormal sound data according to the first weight value and the second weight value corresponding to each timestamp to obtain fusion data corresponding to the timestamp; and integrating the fusion data according to the time sequence of the audio data or the video data to obtain the time-series fusion data.

[0011] Optionally, after obtaining the video data and the audio data collected by the image acquisition device during the operation of the elevator, the method for monitoring the state of the elevator further includes: performing illumination detection on the video data to obtain an illumination index of the video data, and performing noise detection on the audio data to obtain a noise index of the audio data; in the case where the illumination index is lower than a preset illumination threshold, performing image enhancement processing on the video data to make the illumination index of the video data higher than the preset illumination threshold; and in the case where the noise index is higher than a preset noise threshold, performing noise reduction processing on the audio data to make the noise index of the audio data lower than the preset noise threshold.

[0012] Optionally, after the time sequence fusion data is input into a state monitoring model to perform state evaluation on the time sequence fusion data by using the state monitoring model to obtain a time sequence health score of the elevator during the running process, the elevator state monitoring method further comprises: comparing each health score in the time sequence health score with a preset score threshold to obtain a comparison result; in a case where the comparison result indicates that at least one health score is lower than the preset score threshold, determining that fusion data corresponding to the low health score in the time sequence fusion data as abnormal fusion data, wherein the low health score is the health score lower than the preset score threshold; obtaining a vibration identifier of the vibration data and an abnormal sound identifier of the abnormal sound data in the abnormal fusion data, wherein the vibration identifier is a normal vibration identifier or an abnormal vibration identifier, and the abnormal sound identifier is an abnormal sound type of the abnormal sound data; determining a fault type corresponding to the vibration identifier and the abnormal sound identifier in a fault graph as an abnormal fault existing in the elevator, wherein the fault graph is used to record a mapping relationship between the vibration identifier, the abnormal sound identifier and the fault type; issuing an abnormal alarm and sending the abnormal fault to a terminal device to prompt a target object to maintain the elevator.

[0013] According to another aspect of the embodiment of the present application, an elevator state monitoring device is further provided, comprising: a first acquisition unit configured to acquire video data and audio data collected by an image collection device during an elevator running process; an analysis unit configured to perform vibration analysis on the video data to obtain time sequence vibration data, and perform abnormal sound analysis on the audio data to obtain time sequence abnormal sound data, wherein the time sequence vibration data comprises vibration data of each timestamp during the running process of the elevator, and the time sequence abnormal sound data comprises abnormal sound data of each timestamp during the running process of the elevator; a fusion unit configured to perform weighted fusion on the time sequence vibration data and the time sequence abnormal sound data by using an attention mechanism to obtain time sequence fusion data; and an evaluation unit configured to input the time sequence fusion data into a state monitoring model to perform state evaluation on the time sequence fusion data by using the state monitoring model to obtain a time sequence health score of the elevator during the running process, wherein the state monitoring model is obtained by training a plurality of sets of first training data by deep learning, and each set of the plurality of sets of first training data comprises historical time sequence fusion data and a historical time sequence health score corresponding to the historical time sequence fusion data, and the time sequence health score comprises a health score representing a health state of each timestamp during the running process of the elevator.

[0014] Optionally, the analysis unit comprises: a clipping module configured to clip video frames corresponding to a high-fault area in the video data to obtain target video data, wherein the high-fault area is at least one of the following: an elevator car area, a guide rail area; a calculation module configured to calculate motion vectors of the elevator in each timestamp in the target video data using an inter-frame optical flow algorithm to obtain a plurality of motion vectors, wherein each motion vector represents a motion state of the elevator in the corresponding timestamp; an integration module configured to integrate a plurality of motion vectors in the time sequence of the target video data to obtain time series data; and an analysis module configured to perform spectral analysis on the time series data to quantify the motion vectors to obtain the time series vibration data.

[0015] Optionally, the analysis module comprises: a first acquisition sub-module configured to perform Fourier transform or time series wavelet transform on the time series data to obtain a video spectrum graph; a first determination sub-module configured to determine a peak value in the video spectrum graph as a vibration frequency of the elevator in each timestamp; a second determination sub-module configured to determine a spectrum value or wavelet coefficient corresponding to the vibration frequency as a vibration amplitude of the elevator corresponding to the timestamp; a third determination sub-module configured to determine the vibration frequency and the vibration amplitude of each timestamp as vibration data corresponding to the timestamp; and a second acquisition sub-module configured to replace the motion vector corresponding to the timestamp in the time series data with the vibration data to obtain the time series vibration data.

[0016] Optionally, the elevator state monitoring device further comprises: an addition sub-module configured to, after determining the vibration frequency and the vibration amplitude of each timestamp as vibration data corresponding to the timestamp, add an abnormal vibration identifier to the vibration data with a vibration frequency greater than a preset frequency and / or a vibration amplitude greater than a preset amplitude to represent that the vibration data is abnormal vibration data; and add a normal vibration identifier to the vibration data with a vibration frequency not greater than a preset frequency and a vibration amplitude not greater than a preset amplitude to represent that the vibration data is normal vibration data.

[0017] Optionally, the analysis unit comprises: a first obtaining module configured to map the audio data to a frequency spectrum arranged according to a mel frequency scale to obtain an audio spectrum graph, wherein the mel frequency scale is a nonlinear frequency scale, and the audio spectrum graph is used to represent the audio frequency of the audio data at each timestamp; a second obtaining module configured to perform feature extraction on the audio spectrum graph to obtain a plurality of pieces of audio feature data; a third obtaining module configured to perform abnormal sound identification on each piece of the audio feature data by using an abnormal sound identification model to obtain an abnormal sound type of each piece of the audio feature data, wherein the abnormal sound identification model is obtained by training a convolutional neural network by using a plurality of sets of second training data, and each set of the plurality of sets of second training data comprises historical audio feature data and a historical abnormal sound type corresponding to the historical audio feature data; a fourth obtaining module configured to add the abnormal sound type as an abnormal sound identifier to the audio feature data to obtain the abnormal sound data corresponding to the timestamp; and a fifth obtaining module configured to integrate the abnormal sound data according to a time sequence of the audio data to obtain the time-series abnormal sound data.

[0018] Optionally, the fusion unit comprises: a determination module configured to determine a first weight value and a second weight value corresponding to each timestamp by using the attention mechanism, wherein the first weight value is a weight coefficient of the vibration data in the timestamp, and the second weight value is a weight coefficient of the abnormal sound data in the timestamp; a fusion module configured to perform weighted fusion on the vibration data and the abnormal sound data according to the first weight value and the second weight value corresponding to each timestamp to obtain fusion data corresponding to the timestamp; and a sixth obtaining module configured to integrate the fusion data according to a time sequence of the audio data or the video data to obtain the time-series abnormal sound data.

[0019] Optionally, the elevator state monitoring device further comprises: a detection unit configured to, after obtaining the video data and the audio data collected by the image acquisition device during the operation of the elevator, perform illumination detection on the video data to obtain an illumination index of the video data, and perform noise detection on the audio data to obtain a noise index of the audio data; a first processing unit configured to, in a case where the illumination index is lower than a preset illumination threshold, perform image enhancement processing on the video data to make the illumination index of the video data higher than the preset illumination threshold; and a second processing unit configured to, in a case where the noise index is higher than a preset noise threshold, perform noise reduction processing on the audio data to make the noise index of the audio data lower than the preset noise threshold.

[0020] Optionally, the monitoring device of the elevator state further comprises: a comparison unit, configured to compare each of the time sequence health scores with a preset score threshold after inputting the time sequence fusion data into a state monitoring model to perform state evaluation on the time sequence fusion data by using the state monitoring model to obtain the time sequence health scores of the elevator in the running process, and obtain a comparison result; a first determination unit, configured to determine, in a case where the comparison result indicates that at least one of the health scores is lower than the preset score threshold, that fusion data corresponding to a low health score in the time sequence fusion data as abnormal fusion data, wherein the low health score is the health score lower than the preset score threshold; a second acquisition unit, configured to acquire a vibration identifier of the vibration data and an abnormal sound identifier of the abnormal sound data in the abnormal fusion data, wherein the vibration identifier is a normal vibration identifier or an abnormal vibration identifier, and the abnormal sound identifier is an abnormal sound type of the abnormal sound data; a second determination unit, configured to determine a fault type corresponding to the vibration identifier and the abnormal sound identifier in a fault graph as an abnormal fault existing in the elevator, wherein the fault graph is used to record a mapping relationship between the vibration identifier, the abnormal sound identifier and the fault type; and an early warning unit, configured to issue an abnormal alarm and send the abnormal fault to a terminal device to prompt a target object to maintain the elevator.

[0021] According to another aspect of the embodiments of the present application, a computer readable storage medium is further provided, the computer readable storage medium comprising a stored program, wherein the program performs any of the above-mentioned elevator state monitoring methods.

[0022] According to another aspect of the embodiments of the present application, a computer program product is further provided, comprising computer instructions, which, when executed by a processor, perform any of the above-mentioned elevator state monitoring methods.

[0023] In the embodiment of the present application, the video data and audio data collected by the image acquisition device during the operation of the elevator are first obtained; then the video data is subjected to vibration analysis to analyze the jitter of the image, and the vibration data in each time stamp of the elevator during the operation is obtained, and the data are integrated in time sequence to obtain the corresponding time series vibration data; the audio data is subjected to abnormal sound analysis, that is, the frequency of the audio signal is analyzed to obtain the abnormal sound data in each time stamp of the elevator during the operation, and the data are integrated in time sequence to obtain the corresponding time series abnormal sound data; then the attention mechanism can be used to perform weighted fusion on the analyzed time series vibration data and time series abnormal sound data to obtain the corresponding time series abnormal sound data. time series fusion data; finally, the time series fusion data can be input into the state monitoring model, so as to use the state monitoring model to perform state evaluation on the time series fusion data and obtain the time series health score of the elevator during operation, thereby achieving the purpose of monitoring the health status of the elevator according to the time series health score, and realizing the technical effect of monitoring the health status of the elevator operation by combining the multimodal fusion data collected by the camera installed in the elevator itself, avoiding additional hardware costs and maintenance costs, and improving the accuracy of elevator health detection, thereby solving the technical problems in related technologies that a single sensor is difficult to synchronously monitor possible faults of the elevator from multiple angles, resulting in inaccurate monitoring results and easily increasing additional hardware costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0025] Figure 1 This is a hardware structure block diagram of a mobile terminal for monitoring an elevator status according to an embodiment of the present application;

[0026] Figure 2 is a flow chart of a method for monitoring an elevator status according to an embodiment of the present application;

[0027] Figure 3 2 is a schematic diagram of an elevator status monitoring device according to an embodiment of the present application.

[0028] The above drawings include the following reference numerals:

[0029] 102. Processor; 104. Memory; 106. Transmission device; 108. Input / output device. DETAILED DESCRIPTION

[0030] In order for those skilled in the art to better understand the scheme of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.

[0031] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily have to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device including a series of steps or units does not have to be limited to those steps or units clearly listed, but can include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0032] For the convenience of description, the following explains some nouns or terms that may be involved in the embodiments of the present application:

[0033] CNN: a kind of feedforward neural network with deep structure, is one of the representative algorithms of deep learning.

[0034] LSTM: used to solve the long-term dependence problem commonly existing in general recurrent neural networks, using LSTM can effectively pass and express information in long time series and will not cause long time ago useful information to be ignored (forgotten). At the same time, LSTM can also solve the gradient vanishing / explosion problem in RNN.

[0035] Zero-DCE: an algorithm for low-light image enhancement, which can improve the brightness and clarity of the image without reference image through deep learning technology.

[0036] As introduced in the background, the single sensor used in the related art has difficulty in synchronously monitoring possible faults of the elevator from multiple angles, resulting in inaccurate monitoring results and easily increasing additional hardware cost. In view of the above defects, a kind of elevator state monitoring method and device, computer readable storage medium are provided in the embodiments of the present application.

[0037] The technical solutions in the embodiments of the present application will be described clearly and completely in conjunction with the drawings in the embodiments of the present application.

[0038] The method embodiments provided in the embodiments of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 This is a hardware structure block diagram of a mobile terminal of an elevator status monitoring method according to an embodiment of the present application. Figure 1 As shown, the mobile terminal may include one or more ( Figure 1 Only one is shown) a processor 102 (the processor 102 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA and other processing devices) and a memory 104 for storing data, wherein the above mobile terminal may also include a transmission device 106 and an input and output device 108 for communication functions. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the mobile terminal. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0039] The memory 104 can be used to store computer programs, such as software programs and modules of application software, such as the computer program corresponding to the elevator status monitoring method in the embodiments of the present application. The processor 102 executes the computer programs stored in the memory 104 to execute various functional applications and data processing, thereby implementing the above-mentioned method. The memory 104 may include high-speed random access memory and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 may further include a memory remotely located relative to the processor 102, and these remote memories can be connected to the mobile terminal via a network. Examples of such networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof. The transmission device 106 is used to receive or transmit data via a network. Specific examples of such networks may include a wireless network provided by the mobile terminal's telecommunications provider. In one example, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In one example, the transmission device 106 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0040] According to the embodiment of the present application, a method embodiment of the elevator state monitoring method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from here.

[0041] Figure 2 is a flowchart of the elevator state monitoring method according to the embodiment of the present application, as Figure 2 shown, the method comprises the following steps:

[0042] Step S202, acquiring video data and audio data collected by the image acquisition device during the operation of the elevator.

[0043] Optionally, the above-mentioned image acquisition device can include but is not limited to a camera, a multifunctional sensor integrating a camera and a microphone, and other devices that can collect both video data and audio data.

[0044] In this embodiment, the camera originally deployed in the elevator for security monitoring can be used as the image acquisition device to collect the video data and audio data of the elevator during operation, so as to avoid adding additional monitoring sensors and other hardware devices, and to avoid additional hardware costs and maintenance costs, thereby realizing non-intrusive monitoring.

[0045] Step S204, performing vibration analysis on the video data to obtain time-series vibration data, and performing abnormal sound analysis on the audio data to obtain time-series abnormal sound data, wherein the time-series vibration data includes vibration data of each timestamp during the operation of the elevator, and the time-series abnormal sound data includes abnormal sound data of each timestamp during the operation of the elevator.

[0046] In this embodiment, the acquired video data can be subjected to vibration analysis, and the acquired audio data can be subjected to abnormal sound detection and abnormal sound analysis, wherein the video data can be used to quantify the vibration condition during the operation of the elevator, so that the health status of the elevator can be analyzed from the vibration condition and possible abnormal sound during the operation of the elevator.

[0047] It should be noted that the time-series vibration data includes vibration data of each timestamp during the operation of the elevator, and the time-series abnormal sound data includes abnormal sound data of each timestamp during the operation of the elevator; that is, the corresponding vibration data and abnormal sound data in multiple timestamps during the operation of the elevator form a time series data, thereby forming the corresponding time-series vibration data and time-series abnormal sound data.

[0048] Step S206, weighting and fusing the time-series vibration data and the time-series abnormal sound data by using an attention mechanism to obtain time-series fusion data.

[0049] In this embodiment, the acquired time-series vibration data and time-series abnormal sound data can be weighted and fused by using an attention mechanism to obtain corresponding time-series fusion data, that is, the vibration data and abnormal sound data are weighted and fused dynamically by combining the weights of the vibration data and abnormal sound data at different time points to improve the accuracy of the evaluation of the health status of the elevator.

[0050] In step S208, the time-series fusion data is input into the state monitoring model to perform state evaluation on the time-series fusion data by using the state monitoring model to obtain a time-series health score of the elevator during operation, wherein the state monitoring model is obtained by training a plurality of sets of first training data by deep learning, and each set of the plurality of sets of first training data includes historical time-series fusion data, a historical time-series health score corresponding to the historical time-series fusion data, and the time-series health score includes a health score representing the health status of the elevator at each timestamp during operation.

[0051] In this embodiment, the state monitoring model trained by using the LSTM model to learn the time-series dependency relationship through historical data can be used to perform state evaluation on the input time-series fusion data to obtain a time-series health score of the elevator during operation.

[0052] It should be noted that the weighted fusion of the time-series vibration data and the time-series abnormal sound data can be realized before inputting the LSTM model or realized inside the LSTM model, which is not specifically limited here. The former can first fuse the features (including the time-series vibration data and the time-series abnormal sound data) of different modalities to form a multi-dimensional feature vector before inputting the features into the LSTM model; this can be realized by simple mathematical operations (such as connection, averaging) or more complex algorithms (such as the first few layers of a deep neural network), and the fused feature vector is then input into the LSTM. The latter can dynamically adjust the feature weights by introducing an attention mechanism, which can more effectively handle time-series dependencies and give different modalities appropriate importance according to specific circumstances, thereby more accurately evaluating the health status of the elevator.

[0053] As can be seen from the above, through the technical solution provided by the above embodiments of the present application, the video data and audio data collected by the image acquisition device during the operation of the elevator can be first obtained; then the video data is subjected to vibration analysis to analyze the jitter of its picture, and the vibration data in each time stamp of the elevator during the operation is obtained, and the data are integrated in chronological order to obtain the corresponding time series vibration data; the audio data is subjected to abnormal sound analysis, that is, the frequency of the audio signal is analyzed to obtain the abnormal sound data in each time stamp of the elevator during the operation, and the data are integrated in chronological order to obtain the corresponding time series abnormal sound data; then the attention mechanism can be used to perform weighted fusion on the analyzed time series vibration data and time series abnormal sound data The corresponding time series fusion data can be obtained by combining the time series fusion data. Finally, the time series fusion data can be input into the state monitoring model, so as to use the state monitoring model to perform state evaluation on the time series fusion data and obtain the time series health score of the elevator during operation, thereby achieving the purpose of monitoring the health status of the elevator according to the time series health score, and realizing the technical effect of monitoring the health status of the elevator operation by combining the multimodal fusion data collected by the camera installed in the elevator itself, avoiding additional hardware costs and maintenance costs, and improving the accuracy of elevator health detection, thereby solving the technical problems in related technologies that a single sensor is difficult to synchronously monitor possible faults of the elevator from multiple angles, resulting in inaccurate monitoring results and easily increasing additional hardware costs.

[0054] According to the above-mentioned embodiment of the present application, vibration analysis is performed on video data to obtain time-series vibration data, including: intercepting video frames corresponding to high-fault-incidence areas in the video data to obtain target video data, wherein the high-fault-incidence areas are at least one of the following: elevator car area, guide rail area; using an inter-frame optical flow algorithm to calculate the motion vector of the elevator in each time stamp in the target video data to obtain multiple motion vectors, wherein each motion vector represents the motion state of the elevator in the corresponding time stamp; integrating the multiple motion vectors according to the time sequence of the target video data to obtain time series data; performing spectral analysis on the time series data to quantify the motion vector to obtain time series vibration data.

[0055] In this embodiment, prior knowledge or target detection algorithms (such as YOLO, SSD) can be used to automatically locate high-fault areas such as the elevator car area and guide rail area of ​​the elevator, and the video frames corresponding to the high-fault areas can be intercepted to obtain target video data; rather than analyzing the entire video data, the amount of calculation can be greatly reduced and the processing speed can be improved; then, vibration analysis is performed on the target video data, that is, the inter-frame optical flow algorithm is used to extract the movement of pixels in continuous video frames within each timestamp to obtain the motion vector of the elevator within each timestamp; then, multiple motion vectors are integrated according to the time sequence of the target video data to obtain corresponding time series data; finally, the time series data can be spectrally analyzed to quantify the vibration frequency and vibration amplitude of the elevator according to the motion vector to obtain time series vibration data.

[0056] Specifically, when analyzing target video data, computer vision technology can also be used to amplify, visualize and accurately measure tiny mechanical vibrations (micro-vibrations) captured by the camera that are almost or completely imperceptible to the naked eye and are related to potential faults, that is, to amplify tiny movements or color changes of specific time-frequency characteristics in the video; that is, the input target video data can be regarded as a signal in three-dimensional space and time, and each frame of the image is spatially downsampled to obtain image sequences of different spatial scales. The time series of each spatial position (pixel or small block) at different scales are time-domain filtered, and the filtered signal (representing a specific frequency) is converted into a signal. The method multiplies the time-varying signal (a small change in the rate) by an amplification factor α, adds the amplified time-varying signal back to the original spatial scale signal, and upsamples and fuses the processed images at each scale to reconstruct the final amplified video sequence. The inter-frame optical flow algorithm can then be used on the amplified video sequence to analyze the elevator's motion vector at each timestamp. By amplifying the tiny movements in this way, the optical flow calculation in these areas can be made more stable and accurate, achieving precise quantification, going beyond simple visualization, and realizing accurate and automated measurement of key micro-vibration parameters (amplitude, frequency, and direction), providing structured data for intelligent diagnosis.

[0057] In a specific embodiment of the present application, spectral analysis is performed on time series data to quantify motion vectors and obtain time series vibration data, including: performing Fourier transform or time series wavelet transform on the time series data to obtain a video spectrum graph; determining the peak in the video spectrum graph as the vibration frequency of the elevator at each time stamp; determining the spectral value or wavelet coefficient corresponding to the vibration frequency as the vibration amplitude of the elevator at the corresponding time stamp; determining the vibration frequency and vibration amplitude of each time stamp as the vibration data of the corresponding time stamp; and using the vibration data to replace the motion vector of the corresponding time stamp in the time series data to obtain the time series vibration data.

[0058] Specifically, the time series data can be subjected to frequency spectrum analysis by using Fourier transform or time series wavelet transform, the frequency spectrum analysis is used to extract the frequency components of the signal, so as to identify the vibration frequency and vibration amplitude of the elevator; if Fourier transform is used, the time domain data can be converted into frequency domain data, i.e. the motion intensity data of the entire time series is converted into a frequency spectrum, from which the dominant vibration frequency and its corresponding amplitude can be identified; if wavelet transform is used, the time domain and frequency domain characteristics of the signal can be analyzed simultaneously, i.e. through multi-resolution analysis of the time series, vibration components of different frequencies can be separated; the output of the wavelet transform is wavelet coefficients, which reflect the local characteristics of the signal at each scale, by analyzing the distribution of the modulus of the wavelet coefficients at different scales, the frequency and amplitude of the vibration can be identified; in the result of the frequency spectrum analysis, the vibration frequency can be determined by finding the peak position in the frequency spectrum, wherein the frequency at which the peak is located is the main frequency of the vibration (i.e. the vibration frequency), the vibration amplitude is quantified by the size of the frequency spectrum value or the wavelet coefficient at the corresponding frequency, the larger the frequency spectrum value or the wavelet coefficient, the stronger the amplitude of the vibration; finally, the motion vector within the corresponding time stamp is replaced by the vibration frequency and the vibration amplitude, so that the corresponding time series vibration data is obtained.

[0059] In an optional embodiment of the present application, after determining the vibration frequency and the vibration amplitude of each time stamp as the vibration data of the corresponding time stamp, the elevator state monitoring method further comprises: adding an abnormal vibration identifier to the vibration data with a vibration frequency greater than a preset frequency and / or a vibration amplitude greater than a preset amplitude, to represent that the vibration data is abnormal vibration data; adding a normal vibration identifier to the vibration data with a vibration frequency not greater than the preset frequency and a vibration amplitude not greater than the preset amplitude, to represent that the vibration data is normal vibration data.

[0060] Specifically, after obtaining each vibration data, a threshold can be determined to distinguish between normal and abnormal vibration characteristics based on known normal operation data distribution, i.e. a corresponding threshold (including a preset frequency and a preset amplitude) can be set for the vibration frequency and the vibration amplitude, the threshold here can be obtained based on historical data analysis, then each vibration frequency and vibration amplitude is compared with the corresponding threshold, if the vibration frequency of a certain vibration data is greater than the preset frequency and / or the vibration amplitude is greater than the preset amplitude, it can be considered that the current vibration data is abnormal, and an abnormal vibration identifier can be added; otherwise, it is considered that the vibration data is normal vibration data, and a normal vibration identifier can be added, so as to discover potential faults and lay a certain foundation for subsequent comprehensive analysis of fused data.

[0061] According to the above embodiment of the present application, the audio data is analyzed to obtain time sequence abnormal sound data, including: mapping the audio data to a spectrum arranged according to a mel frequency scale to obtain an audio spectrum graph, wherein the mel frequency scale is a non-linear frequency scale, and the audio spectrum graph is used to represent the audio frequency of the audio data at each timestamp; performing feature extraction on the audio spectrum graph to obtain a plurality of audio feature data; using an abnormal sound identification model to identify the abnormal sound of each audio feature data respectively to obtain the abnormal sound type of each audio feature data, wherein the abnormal sound identification model is obtained by training a convolutional neural network using a plurality of second training data, and each of the plurality of second training data includes: historical audio feature data, and a historical abnormal sound type corresponding to the historical audio feature data; adding the abnormal sound type as an abnormal sound identifier to the audio feature data to obtain abnormal sound data corresponding to the timestamp; and integrating the abnormal sound data according to the time sequence of the audio data to obtain time sequence abnormal sound data.

[0062] In this embodiment, first, the collected audio data can be converted into a mel spectrum graph, which is a frequency spectrum representation method based on human auditory perception. The audio data can be mapped to a spectrum arranged according to a mel frequency scale. Compared with linear spectrum, the mel spectrum graph can better simulate the sensitivity difference of human ears in different frequency ranges, thereby more efficiently capturing and expressing acoustic features. Then, a pre-trained CNN model (i.e., the abnormal sound identification model in the present application) can be used to analyze the mel spectrum graph. CNN model is particularly suitable for processing data with spatial structure, such as images and spectrum graphs. In acoustic anomaly detection, CNN can automatically extract features from the mel spectrum graph to identify abnormal sound patterns related to the health status of the elevator. CNN extracts features from local regions of the mel spectrum graph through convolution layers, captures the spatial and temporal structure of the audio signal, and uses activation functions (such as ReLU) for nonlinear transformation to increase the expression ability of the model. The pooling layer is used to reduce the spatial dimension, and also plays a role in feature dimension reduction and noise suppression. The fully connected layer is used to integrate these features for final classification or anomaly detection, so that the abnormal sound type (such as bearing friction sound and high-frequency noise before steel wire rope breakage) corresponding to the audio data can be obtained. Then, the abnormal sound data of each timestamp is integrated in time sequence to obtain the corresponding time sequence abnormal sound data. Of course, if the frequency of a certain audio signal is within the normal range, its abnormal sound type can be marked as normal.

[0063] According to the above embodiment of the present application, the time sequence vibration data and the time sequence abnormal sound data are weighted and fused by using the attention mechanism to obtain time sequence fusion data, including: determining the first weight value and the second weight value corresponding to each timestamp by using the attention mechanism, wherein the first weight value is the weight coefficient of the vibration data in the corresponding timestamp, and the second weight value is the weight coefficient of the abnormal sound data in the corresponding timestamp; weighting and fusing the vibration data and the abnormal sound data according to the first weight value and the second weight value corresponding to each timestamp to obtain the fusion data in the corresponding timestamp; and integrating the fusion data according to the time sequence of the audio data or the video data to obtain the time sequence abnormal sound data.

[0064] In this embodiment, since the importance of the information of the two modalities of the time sequence vibration data reflecting the vibration of the elevator and the time sequence abnormal sound data reflecting the abnormal sound of the elevator may be different in different scenarios, for example, when the elevator is stationary or running at low speed, the acoustic signal may be more valuable for diagnosis, while when running at high speed or in some specific fault conditions, the visual vibration data may be more critical; therefore, the attention mechanism can be used to dynamically determine the weight of the visual feature (equivalent to the vibration data) and the acoustic feature (equivalent to the abnormal sound data) in each timestamp, and the feature with a larger weight will have a greater impact on the final result, then the vibration data and the abnormal sound data can be weighted and fused based on the assigned weight to obtain the corresponding fusion data, and then the fusion data of each timestamp is integrated according to the time sequence, i.e. the time sequence abnormal sound data is obtained; thus, the attention mechanism is introduced to adaptively adjust the weight distribution according to the characteristics of the input data, which can improve the accuracy and flexibility of the fusion result, and ensure accurate capture and efficient use of key information.

[0065] In another optional embodiment of the present application, after obtaining the video data and the audio data collected by the image acquisition device during the operation of the elevator, the monitoring method of the elevator state further includes: performing illumination detection on the video data to obtain an illumination index of the video data, and performing noise detection on the audio data to obtain a noise index of the audio data; in the case that the illumination index is lower than a preset illumination threshold, performing image enhancement processing on the video data to make the illumination index of the video data higher than the preset illumination threshold; in the case that the noise index is higher than a preset noise threshold, performing noise reduction processing on the audio data to make the noise index of the audio data lower than the preset noise threshold.

[0066] In this embodiment, after obtaining the video data and the audio data collected by the image acquisition device, it can be preliminarily detected whether the illumination index of the video data meets the requirements (i.e. whether it is not lower than the preset illumination threshold) and whether the audio data meets the requirements (i.e. whether it is not higher than the preset noise threshold), if not, it is considered that it may affect the accuracy of the subsequent analysis result, and it can be processed in advance.

[0067] Specifically, the illumination index of the video data is lower than the preset illumination threshold, which may be due to the insufficient light in the running environment of the elevator, especially in the scene of insufficient light such as night or underground parking lot, the image captured by the camera may appear blurred and low contrast, and the details in the image are not easy to identify; for such low-illumination scene, a low-illumination image enhancement network (such as Zero-DCE) can be used to process the image to improve the signal-to-noise ratio of the vibration signal, that is, the proportion of effective signal (such as pixel change caused by elevator vibration) and noise (such as image noise, artifact) in the image is increased, thereby enhancing the image quality, so that it can more clearly show the vibration details in the running of the elevator, so that the visual vibration detection technology such as optical flow analysis is more accurate, and the adverse effects of image noise caused by low light on vibration signal detection are reduced.

[0068] Specifically, the noise index of the audio data is higher than the preset noise threshold, which may be due to the abnormal sound generated by potential faults during normal operation of the elevator and the surrounding environmental noise (such as background noise inside the building, human voice outside the elevator, equipment sound, etc.). For such high-noise scene, an ambient noise suppression algorithm can be introduced, and the microphone array beamforming technology can be used to directionally focus the elevator shaft sound source. The microphone array can suppress noise, and beamforming can form a "beam" pointing to a specific sound source direction by weighting and phase adjusting the signals received by the microphone array, so as to concentrate and amplify the sound signals in the elevator shaft, while minimizing the noise in other directions. In this way, the quality of the audio signal can be improved, the influence of environmental noise can be reduced, and the subsequent abnormal sound feature recognition can be more accurate and reliable.

[0069] In another optional embodiment of the present application, after the time sequence fusion data is input into the state monitoring model to perform state evaluation on the time sequence fusion data by using the state monitoring model to obtain the time sequence health score of the elevator in the running process, the elevator state monitoring method further comprises: comparing each health score in the time sequence health score with a preset score threshold to obtain a comparison result; in the case that the comparison result indicates that at least one health score is lower than the preset score threshold, determining that the fusion data corresponding to the low health score in the time sequence fusion data as abnormal fusion data, wherein the low health score is a health score lower than the preset score threshold; obtaining the vibration identifier of the vibration data and the abnormal sound identifier of the abnormal sound data in the abnormal fusion data, wherein the vibration identifier is a normal vibration identifier or an abnormal vibration identifier, and the abnormal sound identifier is an abnormal sound type of the abnormal sound data; determining that the fault type corresponding to the vibration identifier and the abnormal sound identifier in the fault graph as an abnormal fault existing in the elevator, wherein the fault graph is used to record the mapping relationship between the vibration identifier, the abnormal sound identifier and the fault type; issuing an abnormal alarm and sending the abnormal fault to a terminal device to prompt a target object to maintain the elevator.

[0070] Optionally, the abnormality alarm can include, but is not limited to, an audible and light alarm, sending a warning message, etc.

[0071] Optionally, the terminal device can include, but is not limited to, a mobile phone, a tablet, a computer, etc.

[0072] In this embodiment, after obtaining the health scores of the elevator at each time point in the running process by processing the time series fusion data using the state monitoring model, the obtained health scores can be compared with the preset score threshold to determine whether there is a fault. If a certain health score is lower than the preset score threshold, it can be considered that the elevator may have a certain fault within the corresponding time stamp. Then, the corresponding vibration data and abnormal sound data can be analyzed, and the fault spectrum recording the mapping relationship between the vibration identifier, the abnormal sound identifier and the fault type can be used to determine the specific fault. For example, assuming that the vibration data within the time stamp corresponding to a certain lower health score carries an abnormal vibration identifier, and the abnormal sound data carries an abnormal sound identifier of metal scraping sound, it can be known from the fault spectrum that the corresponding fault is the traction sheave eccentricity. Then, an abnormality alarm can be issued, and the identified fault type (such as traction sheave eccentricity) can be sent to the terminal device of the relevant maintenance personnel to prompt the maintenance personnel to maintain the elevator and make maintenance preparations in advance for the fault.

[0073] It should be noted that for some health scores higher than the preset score threshold, but the vibration data in the corresponding fusion data may carry an abnormal vibration identifier or the abnormal sound data may carry an abnormal sound identifier of a certain abnormal sound type, it is possible that these abnormal targets do not affect the normal operation of the elevator. In order to avoid the possible impact on the safe operation of the elevator, these vibration identifiers and abnormal sound identifiers can also be sent to the terminal device of the maintenance personnel, so that they can make some maintenance in advance to avoid the occurrence of faults, thereby improving the safety of the elevator operation.

[0074] It should be noted that for the foregoing method embodiments, in order to simply describe, they are all expressed as a series of action combinations, but those skilled in the art should know that the present application is not limited by the action sequence described, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily required by the present application.

[0075] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be realized by means of software on a general hardware platform as necessary, and of course can also be realized by hardware, but in many cases the former is a better embodiment. Based on such understanding, the technical solutions of the present application can be embodied in the form of a software product in essence or in the form of a part that contributes to the prior art, and the computer software product is stored in a storage medium (such as a ROM / RAM, a magnetic disk, or an optical disk), and includes a plurality of instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device) to execute the method described in each embodiment of the present application.

[0076] According to the embodiments of the present application, a monitoring device for monitoring the state of an elevator is also provided, Figure 3 is a schematic diagram of the monitoring device for monitoring the state of an elevator according to the embodiments of the present application, as Figure 3 shown, the device includes a first acquisition unit 31, an analysis unit 33, a fusion unit 35, and an evaluation unit 37. The monitoring device for monitoring the state of an elevator will be described in detail below.

[0077] The first acquisition unit 31 is configured to acquire video data and audio data collected by an image collection device during operation of the elevator.

[0078] The analysis unit 33 is configured to perform vibration analysis on the video data to obtain time-series vibration data, and perform abnormal sound analysis on the audio data to obtain time-series abnormal sound data, wherein the time-series vibration data includes vibration data of each timestamp during operation of the elevator, and the time-series abnormal sound data includes abnormal sound data of each timestamp during operation of the elevator.

[0079] The fusion unit 35 is configured to perform weighted fusion on the time-series vibration data and the time-series abnormal sound data by using an attention mechanism to obtain time-series fusion data.

[0080] The evaluation unit 37 is configured to input the time-series fusion data into a state monitoring model to perform state evaluation on the time-series fusion data by using the state monitoring model to obtain a time-series health score of the elevator during operation, wherein the state monitoring model is obtained by training a plurality of sets of first training data by deep learning, and each set of the plurality of sets of first training data includes historical time-series fusion data and a historical time-series health score corresponding to the historical time-series fusion data, and the time-series health score includes a health score representing a health state of each timestamp during operation of the elevator.

[0081] It should be noted that the first acquisition unit 31, the analysis unit 33, the fusion unit 35 and the evaluation unit 37 correspond to steps S202 to S208 in the above embodiment, and the four units have the same instances and application scenarios as the corresponding steps, but are not limited to the content disclosed in the above embodiment.

[0082] As can be seen from the above, in the scheme described in the above embodiment of the present application, the first acquisition unit can be used to acquire video data and audio data collected by the image acquisition device during the operation of the elevator; then the analysis unit is used to analyze the video data to analyze the shaking of the picture, obtain the vibration data of the elevator in each time stamp during the operation, and integrate the vibration data in time sequence to obtain the corresponding time sequence vibration data, and analyze the audio data, that is, analyze the frequency of the audio signal to obtain the abnormal sound data of the elevator in each time stamp during the operation, and integrate the abnormal sound data in time sequence to obtain the corresponding time sequence abnormal sound data; then the fusion unit is used to weight and fuse the time sequence vibration data and the time sequence abnormal sound data by using the attention mechanism to obtain the time sequence fusion data; finally, the evaluation unit is used to input the time sequence fusion data into the state monitoring model to evaluate the state of the time sequence fusion data by using the state monitoring model to obtain the time sequence health score of the elevator during the operation, so as to achieve the purpose of monitoring the health state of the elevator according to the time sequence health score, realize the technical effect of monitoring the health state of the elevator by combining the multi-modal fusion data collected by the camera arranged in the elevator itself, avoid additional hardware cost and maintenance cost, improve the accuracy of elevator health detection, and solve the technical problems in the related art that the single sensor is difficult to monitor the possible faults of the elevator from multiple angles at the same time, the monitoring result is inaccurate, and the additional hardware cost is increased.

[0083] In an optional embodiment of the present application, the analysis unit comprises: an intercepting module configured to intercept video frames corresponding to a high-fault area in the video data to obtain target video data, wherein the high-fault area is at least one of the following: an elevator car area, a guide rail area; a calculation module configured to calculate motion vectors of the elevator in each time stamp in the target video data by using an inter-frame optical flow algorithm to obtain a plurality of motion vectors, wherein each motion vector represents the motion state of the elevator in the corresponding time stamp; an integration module configured to integrate the plurality of motion vectors in time sequence of the target video data to obtain time sequence data; and an analysis module configured to perform frequency spectrum analysis on the time sequence data to quantize the motion vectors to obtain time sequence vibration data.

[0084] In an optional embodiment of the present application, the analysis module comprises: a first obtaining sub-module, configured to perform Fourier transform or time series wavelet transform on the time series data to obtain a video spectrum diagram; a first determining sub-module, configured to determine a peak value in the video spectrum diagram as a vibration frequency of the elevator at each timestamp; a second determining sub-module, configured to determine a spectrum value or a wavelet coefficient corresponding to the vibration frequency as a vibration amplitude of the elevator at the corresponding timestamp; a third determining sub-module, configured to determine the vibration frequency and the vibration amplitude at each timestamp as vibration data of the corresponding timestamp; and a second obtaining sub-module, configured to replace a motion vector at the corresponding timestamp in the time series data with the vibration data to obtain time series vibration data.

[0085] In an optional embodiment of the present application, the monitoring device of the elevator state further comprises: an adding sub-module, configured to, after determining the vibration frequency and the vibration amplitude at each timestamp as the vibration data of the corresponding timestamp, add an abnormal vibration identifier to the vibration data with the vibration frequency greater than a preset frequency and / or the vibration amplitude greater than a preset amplitude, so as to represent that the vibration data is abnormal vibration data; and add a normal vibration identifier to the vibration data with the vibration frequency not greater than the preset frequency and the vibration amplitude not greater than the preset amplitude, so as to represent that the vibration data is normal vibration data.

[0086] In an optional embodiment of the present application, the analysis unit comprises: a first obtaining module, configured to map the audio data to a spectrum arranged according to a mel frequency scale to obtain an audio spectrum diagram, wherein the mel frequency scale is a nonlinear frequency scale, and the audio spectrum diagram is used to represent an audio frequency of the audio data at each timestamp; a second obtaining module, configured to perform feature extraction on the audio spectrum diagram to obtain a plurality of pieces of audio feature data; a third obtaining module, configured to perform abnormal sound identification on each piece of audio feature data by using an abnormal sound identification model to obtain an abnormal sound type of each piece of audio feature data, wherein the abnormal sound identification model is obtained by training a convolutional neural network by using a plurality of sets of second training data, and each set of the plurality of sets of second training data comprises historical audio feature data and a historical abnormal sound type corresponding to the historical audio feature data; a fourth obtaining module, configured to add the abnormal sound type as an abnormal sound identifier to the audio feature data to obtain abnormal sound data of the corresponding timestamp; and a fifth obtaining module, configured to integrate the abnormal sound data according to a time sequence of the audio data to obtain time series abnormal sound data.

[0087] In an optional embodiment of the present application, the fusion unit comprises: a determination module configured to determine a first weight value and a second weight value corresponding to each timestamp by using an attention mechanism, wherein the first weight value is a weight coefficient of the vibration data in the corresponding timestamp, and the second weight value is a weight coefficient of the abnormal sound data in the corresponding timestamp; a fusion module configured to perform weighted fusion on the vibration data and the abnormal sound data according to the first weight value and the second weight value corresponding to each timestamp to obtain fusion data in the corresponding timestamp; and a sixth acquisition module configured to integrate the fusion data according to the time sequence of the audio data or the video data to obtain time-series abnormal sound data.

[0088] In an optional embodiment of the present application, the monitoring device of the elevator state further comprises: a detection unit configured to perform illumination detection on the video data to obtain an illumination index of the video data and noise detection on the audio data to obtain a noise index of the audio data after acquiring the video data and the audio data collected by the image acquisition device during the operation of the elevator; a first processing unit configured to perform image enhancement processing on the video data to make the illumination index of the video data higher than a preset illumination threshold value in the case where the illumination index is lower than the preset illumination threshold value; and a second processing unit configured to perform noise reduction processing on the audio data to make the noise index of the audio data lower than a preset noise threshold value in the case where the noise index is higher than the preset noise threshold value.

[0089] In an optional embodiment of the present application, the monitoring device of the elevator state further comprises: a comparison unit configured to compare each health score in the time-series health score with a preset score threshold value to obtain a comparison result after inputting the time-series fusion data into the state monitoring model to perform state evaluation on the time-series fusion data by using the state monitoring model to obtain the time-series health score of the elevator during the operation; a first determination unit configured to determine that the fusion data corresponding to the low health score in the time-series fusion data is abnormal fusion data in the case where the comparison result indicates that at least one health score is lower than the preset score threshold value, wherein the low health score is a health score lower than the preset score threshold value; a second acquisition unit configured to acquire a vibration identifier of the vibration data and an abnormal sound identifier of the abnormal sound data in the abnormal fusion data, wherein the vibration identifier is a normal vibration identifier or an abnormal vibration identifier, and the abnormal sound identifier is an abnormal sound type of the abnormal sound data; a second determination unit configured to determine that a fault type corresponding to the vibration identifier and the abnormal sound identifier in a fault atlas is an abnormal fault existing in the elevator, wherein the fault atlas is used to record the mapping relationship between the vibration identifier, the abnormal sound identifier and the fault type; and an early warning unit configured to issue an abnormal alarm and send the abnormal fault to a terminal device to prompt a target object to repair the elevator.

[0090] According to another aspect of the embodiments of the present application, the elevator state monitoring system is also provided, which uses any of the elevator state monitoring methods described above.

[0091] According to another aspect of the embodiments of the present application, the computer readable storage medium is also provided, which includes a stored program, wherein the program executes any of the elevator state monitoring methods described above.

[0092] Optionally, in the present embodiment, the computer readable storage medium can be located in any of the computer terminals in the computer terminal group in the computer network, or in any of the communication devices in the communication device group.

[0093] Optionally, in the present embodiment, the computer readable storage medium is configured to store program code for performing the following steps: obtaining video data and audio data collected by the image acquisition device during the operation of the elevator; performing vibration analysis on the video data to obtain time-series vibration data, and performing abnormal sound analysis on the audio data to obtain time-series abnormal sound data, wherein the time-series vibration data includes vibration data of each timestamp during the operation of the elevator, and the time-series abnormal sound data includes abnormal sound data of each timestamp during the operation of the elevator; using the attention mechanism to perform weighted fusion on the time-series vibration data and the time-series abnormal sound data to obtain time-series fusion data; inputting the time-series fusion data into the state monitoring model to perform state evaluation on the time-series fusion data by using the state monitoring model to obtain a time-series health score of the elevator during the operation, wherein the state monitoring model is trained by using a plurality of sets of first training data through deep learning, and each set of the plurality of sets of first training data includes historical time-series fusion data, a historical time-series health score corresponding to the historical time-series fusion data, and a health score representing the health state of the elevator at each timestamp during the operation.

[0094] Optionally, in the present embodiment, the computer readable storage medium is configured to store program code for performing the following steps: intercepting video frames corresponding to a high-failure area in the video data to obtain target video data, wherein the high-failure area is at least one of the following: an elevator car area, a guide rail area; calculating the motion vector of the elevator in each timestamp in the target video data by using an inter-frame optical flow algorithm to obtain a plurality of motion vectors, wherein each motion vector represents the motion state of the elevator in the corresponding timestamp; integrating the plurality of motion vectors according to the time sequence of the target video data to obtain time series data; performing frequency spectrum analysis on the time series data to quantize the motion vectors to obtain time-series vibration data.

[0095] Optionally, in the embodiment, the computer readable storage medium is configured to store program code for performing the following steps: performing Fourier transform or time series wavelet transform on the time series data to obtain a video spectrum map; determining a peak value in the video spectrum map as a vibration frequency of the elevator at each time stamp; determining a spectrum value or wavelet coefficient corresponding to the vibration frequency as a vibration amplitude of the elevator at the corresponding time stamp; determining the vibration frequency and the vibration amplitude at each time stamp as vibration data at the corresponding time stamp; and replacing the motion vector at the corresponding time stamp in the time series data with the vibration data to obtain time series vibration data.

[0096] Optionally, in the embodiment, the computer readable storage medium is configured to store program code for performing the following steps: adding an abnormal vibration identifier to the vibration data with a vibration frequency greater than a preset frequency and / or a vibration amplitude greater than a preset amplitude, to represent that the vibration data is abnormal vibration data; and adding a normal vibration identifier to the vibration data with a vibration frequency not greater than the preset frequency and a vibration amplitude not greater than the preset amplitude, to represent that the vibration data is normal vibration data.

[0097] Optionally, in the embodiment, the computer readable storage medium is configured to store program code for performing the following steps: mapping the audio data to a spectrum arranged according to a mel frequency scale to obtain an audio spectrum map, wherein the mel frequency scale is a non-linear frequency scale, and the audio spectrum map is used to represent an audio frequency of the audio data at each time stamp; performing feature extraction on the audio spectrum map to obtain a plurality of audio feature data; performing abnormal sound identification on each audio feature data by using an abnormal sound identification model to obtain an abnormal sound type of each audio feature data, wherein the abnormal sound identification model is obtained by training a convolutional neural network by using a plurality of second training data, and each of the plurality of second training data includes historical audio feature data and a historical abnormal sound type corresponding to the historical audio feature data; adding the abnormal sound type as an abnormal sound identifier to the audio feature data to obtain abnormal sound data at the corresponding time stamp; and integrating the abnormal sound data according to a time sequence of the audio data to obtain time series abnormal sound data.

[0098] Optionally, in the embodiment, the computer readable storage medium is configured to store program code for performing the following steps: determining a first weight value and a second weight value corresponding to each time stamp by using an attention mechanism, wherein the first weight value is a weight coefficient of the vibration data in the corresponding time stamp, and the second weight value is a weight coefficient of the abnormal sound data in the corresponding time stamp; and performing weighted fusion on the vibration data and the abnormal sound data according to the first weight value and the second weight value corresponding to each time stamp to obtain fusion data in the corresponding time stamp; and integrating the fusion data according to a time sequence of the audio data or the video data to obtain time series abnormal sound data.

[0099] Optionally, in the embodiment, the computer readable storage medium is configured to store program code for performing the following steps: performing illumination detection on the video data to obtain an illumination index of the video data, and performing noise detection on the audio data to obtain a noise index of the audio data; performing image enhancement processing on the video data to make the illumination index of the video data higher than a preset illumination threshold, in a case where the illumination index is lower than the preset illumination threshold; performing noise reduction processing on the audio data to make the noise index of the audio data lower than a preset noise threshold, in a case where the noise index is higher than the preset noise threshold.

[0100] Optionally, in the embodiment, the computer readable storage medium is configured to store program code for performing the following steps: comparing each health score in the time sequence health scores with a preset score threshold to obtain a comparison result; determining that fusion data corresponding to a low health score in the time sequence fusion data as abnormal fusion data, in a case where the comparison result indicates that at least one health score is lower than the preset score threshold, wherein the low health score is a health score lower than the preset score threshold; obtaining a vibration identifier of vibration data and an abnormal sound identifier of abnormal sound data in the abnormal fusion data, wherein the vibration identifier is a normal vibration identifier or an abnormal vibration identifier, and the abnormal sound identifier is an abnormal sound type of the abnormal sound data; determining that a fault type corresponding to the vibration identifier and the abnormal sound identifier in the fault atlas as an abnormal fault existing in the elevator, wherein the fault atlas is used to record a mapping relationship between the vibration identifier, the abnormal sound identifier and the fault type; issuing an abnormal alarm and sending the abnormal fault to a terminal device to prompt a target object to maintain the elevator.

[0101] According to another aspect of the embodiments of the present application, a processor is also provided, which is used to run a program, wherein the program performs any one of the monitoring methods of the state of the elevator when running.

[0102] According to another aspect of the embodiments of the present application, a computer program product is also provided, which is used to run a program, wherein the program performs any one of the monitoring methods of the state of the elevator when running.

[0103] The above sequence numbers of the embodiments of the present application are only for description, and do not represent the advantages and disadvantages of the embodiments.

[0104] In the above embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.

[0105] In several embodiments provided in the present application, it should be understood that the disclosed technology can be implemented by other ways. Among them, the above-described device embodiments are only schematic, for example, the division of the units can be a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, units or modules, and can be electrical or other forms.

[0106] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed to multiple units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.

[0107] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0108] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, including a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.

[0109] The above is only the preferred embodiment of the present application, and it should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.

Claims

1. A method of monitoring the state of an elevator, characterized by The method comprises: acquiring video data and audio data collected by an image collection device during elevator operation; performing vibration analysis on the video data to obtain time-series vibration data, and performing abnormal sound analysis on the audio data to obtain time-series abnormal sound data, wherein the time-series vibration data comprises vibration data of the elevator at each timestamp during the operation, and the time-series abnormal sound data comprises abnormal sound data of the elevator at each timestamp during the operation; performing weighted fusion of the time-series vibration data and the time-series abnormal sound data using an attention mechanism to obtain time-series fusion data; inputting the time-series fusion data into a state monitoring model to perform state evaluation on the time-series fusion data using the state monitoring model to obtain a time-series health score of the elevator during the operation, wherein the state monitoring model is obtained by training a plurality of sets of first training data using deep learning, and each set of the plurality of sets of first training data comprises historical time-series fusion data and a historical time-series health score corresponding to the historical time-series fusion data, wherein the time-series health score comprises a health score representing the health status of the elevator at each timestamp during the operation.

2. The method of monitoring the state of an elevator according to claim 1, characterized in that, The vibration analysis on the video data to obtain time-series vibration data comprises: extracting video frames corresponding to a high-failure area in the video data to obtain target video data, wherein the high-failure area is at least one of the following: an elevator car area, a guide rail area; calculating motion vectors of the elevator in each timestamp in the target video data using an inter-frame optical flow algorithm to obtain a plurality of motion vectors, wherein each motion vector represents the motion state of the elevator in the corresponding timestamp; integrating a plurality of motion vectors in the time order of the target video data to obtain time series data; performing frequency spectrum analysis on the time series data to quantize the motion vectors to obtain the time-series vibration data.

3. The method of monitoring the state of an elevator according to claim 2, characterized in that, The frequency spectrum analysis on the time series data to quantize the motion vectors to obtain the time-series vibration data comprises: performing Fourier transform or time-series wavelet transform on the time series data to obtain a video frequency spectrum graph; determining the peak value in the video frequency spectrum graph as the vibration frequency of the elevator at each timestamp; determining the frequency spectrum value or wavelet coefficient corresponding to the vibration frequency as the vibration amplitude of the elevator at the corresponding timestamp; determining the vibration frequency and the vibration amplitude of each timestamp as the vibration data of the corresponding timestamp; replacing the motion vector of the corresponding timestamp in the time series data with the vibration data to obtain the time-series vibration data.

4. The method of monitoring the state of an elevator according to claim 3, characterized in that, After determining the vibration frequency and the vibration amplitude of each timestamp as the vibration data of the corresponding timestamp, the method further comprises: adding an abnormal vibration identifier to the vibration data with a vibration frequency greater than a preset frequency and / or a vibration amplitude greater than a preset amplitude to represent the vibration data as abnormal vibration data; add a normal vibration identifier to the vibration data, where the vibration frequency is not greater than a preset frequency and the vibration amplitude is not greater than a preset amplitude, so as to represent the vibration data as normal vibration data.

5. The method of monitoring the state of an elevator according to claim 1, characterized in that, perform abnormal sound analysis on the audio data to obtain time-series abnormal sound data, including: mapping the audio data to a frequency spectrum arranged according to a mel frequency scale to obtain an audio spectrum graph, where the mel frequency scale is a nonlinear frequency scale, and the audio spectrum graph is used to represent the audio frequency of the audio data at each timestamp; performing feature extraction on the audio spectrum graph to obtain a plurality of audio feature data; performing abnormal sound identification on each piece of audio feature data using an abnormal sound identification model to obtain an abnormal sound type of each piece of audio feature data, where the abnormal sound identification model is obtained by training a convolutional neural network using a plurality of second training data, and each of the plurality of second training data includes historical audio feature data and a corresponding historical abnormal sound type; adding the abnormal sound type to the audio feature data as an abnormal sound identifier to obtain the abnormal sound data corresponding to the timestamp; integrating the abnormal sound data according to the time sequence of the audio data to obtain the time-series abnormal sound data.

6. The method of monitoring the state of an elevator according to claim 1, characterized in that, performing weighted fusion on the time-series vibration data and the time-series abnormal sound data using an attention mechanism to obtain time-series fusion data, including: determining a first weight value and a second weight value corresponding to each timestamp using the attention mechanism, where the first weight value is a weight coefficient corresponding to the vibration data within the timestamp, and the second weight value is a weight coefficient corresponding to the abnormal sound data within the timestamp; performing weighted fusion on the vibration data and the abnormal sound data according to the first weight value and the second weight value corresponding to each timestamp to obtain fusion data corresponding to the timestamp; integrating the fusion data according to the time sequence of the audio data or the video data to obtain the time-series abnormal sound data.

7. The method of monitoring the state of an elevator according to claim 1, characterized in that, After obtaining the video data and audio data collected by the image acquisition device during the operation of the elevator, the method further includes: performing illumination detection on the video data to obtain an illumination index of the video data, and performing noise detection on the audio data to obtain a noise index of the audio data; performing image enhancement processing on the video data to make the illumination index of the video data higher than a preset illumination threshold value, in the case that the illumination index is lower than the preset illumination threshold value; performing noise reduction processing on the audio data to make the noise index of the audio data lower than a preset noise threshold value, in the case that the noise index is higher than the preset noise threshold value.

8. The method of monitoring the state of an elevator according to claim 1, characterized in that, After inputting the time-series fusion data into a state monitoring model to perform state evaluation on the time-series fusion data using the state monitoring model to obtain a time-series health score of the elevator during the operation, the method further includes: comparing each health score in the time-series health score with a preset score threshold to obtain a comparison result; In a case where the comparison result indicates that at least one of the health scores is lower than the preset score threshold, the fusion data corresponding to the low health score in the time-series fusion data is determined as abnormal fusion data, where the low health score is the health score lower than the preset score threshold. An abnormal alarm is issued, and the abnormal fault is sent to a terminal device to prompt a target object to maintain the elevator. The first acquisition unit is configured to acquire video data and audio data collected by an image collection device during operation of an elevator. The analysis unit is configured to perform vibration analysis on the video data to obtain time-series vibration data, and perform abnormal sound analysis on the audio data to obtain time-series abnormal sound data, where the time-series vibration data includes vibration data of each timestamp during the operation of the elevator, and the time-series abnormal sound data includes abnormal sound data of each timestamp during the operation of the elevator.

9. An elevator state monitoring device, characterized by The fusion unit is configured to perform weighted fusion on the time-series vibration data and the time-series abnormal sound data by using an attention mechanism to obtain time-series fusion data. The evaluation unit is configured to input the time-series fusion data into a state monitoring model to perform state evaluation on the time-series fusion data by using the state monitoring model to obtain a time-series health score of the elevator during the operation, where the state monitoring model is obtained by training a plurality of sets of first training data by deep learning, and each set of the plurality of sets of first training data includes historical time-series fusion data and a historical time-series health score corresponding to the historical time-series fusion data, and the time-series health score includes a health score representing a health state of the elevator at each timestamp during the operation. The computer readable storage medium includes a stored program, where the program performs the method for monitoring the state of the elevator according to any one of claims 1 to 8. The computer instructions are executed by the processor to perform the method for monitoring the state of the elevator according to any one of claims 1 to 8. ​ 10. A computer readable storage medium characterized by ​ 11. A computer program product comprising computer instructions, characterized in that, ​

Citation Information

Patent Citations

  • Elevator data acquisition method and device

    CN106395532A

  • Elevator abnormal sound detection method and system

    CN112193959A

  • Elevator running evaluation system

    CN112978531A

  • Elevator noise level monitoring method and system based on audio signals

    CN113581956A

  • Mine audio and video monitoring method based on video analysis

    CN114363551A

Cited By

  • Elevator noise and vibration test system

    CN121048918A

  • Intelligent elevator operation and maintenance management method and system based on multi-mode neural network

    CN121493741A

  • Elevator health degree evaluation system

    CN121672300A