Elevator abnormal behavior detection system based on multi-mode neural network

By applying a multimodal neural network detection system in elevators, combined with multiple data processing and feature fusion technologies, the problem of difficulty in monitoring complex abnormal behaviors in traditional elevator safety monitoring systems is solved, and efficient and safe management of elevator operation is achieved.

CN120024777AInactive Publication Date: 2025-05-23XIAMEN TIANYU INTERNET OF THINGS TECH CO LTD

Patent Information

Application Number
CN202510506698.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-05-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional elevator safety monitoring systems are difficult to effectively monitor complex abnormal behaviors during elevator operation, resulting in potential failures and safety accidents.

Method used

The elevator abnormal behavior detection system based on multimodal neural network is adopted to collect multimodal data through vision sensors, inertial measurement units and audio acquisition devices. Combined with improved wavelet packet transformation, 3D-ResNet architecture and deep reinforcement learning strategies, data processing and feature fusion are carried out to achieve accurate identification of elevator abnormal behavior.

Benefits of technology

It significantly improves the safety of elevator operation, and through comprehensive and accurate abnormal detection and timely and effective alarm mechanisms, it reduces missed and false alarms, improves user experience, and optimizes the elevator operation management process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120024777A_ABST
    Figure CN120024777A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of equipment management and maintenance, in particular to an elevator abnormal behavior detection system based on a multi-modal neural network, which collects visual and inertial measurement and audio data in an elevator car through a data acquisition module, and performs time-frequency domain conversion and feature vectorization on original data by using a data processing module; the multi-modal neural network model is composed of a CNN branch, an RNN branch and a voiceprint recognition network which are parallel, features are integrated through a feature fusion module by adopting a double fusion architecture, and then abnormality is judged through an abnormal behavior classifier based on a hierarchical classification strategy; the alarm and log module gives an alarm in time and records abnormal information, the model optimization module improves the model performance by applying multiple methods, and the system and an elevator control system achieve data interaction through a specific interface. According to the elevator abnormal behavior detection method, multi-modal data and an innovative algorithm are integrated, the accuracy and reliability of elevator abnormal behavior detection are remarkably improved, safe operation of an elevator is effectively guaranteed, and powerful support is provided for intelligent management of the elevator.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of equipment management and maintenance, and in particular to an elevator abnormal behavior detection system based on a multimodal neural network. Background Art

[0002] With the acceleration of urbanization, the number of elevators used, as an indispensable means of vertical transportation in modern buildings, has increased dramatically.

[0003] Traditional elevator safety monitoring mainly relies on simple sensors, such as speed sensors, door limit switches, etc., which can only monitor some operating parameters of the elevator. The monitoring capability is extremely limited for complex abnormal behaviors during elevator operation, such as abnormal behavior of people in the elevator (fighting, fainting, etc.), potential failures of equipment (wear of mechanical parts, abnormal electrical system, etc.), and problems caused by environmental factors (abnormal noise and abnormal smell in the car, etc.). For example, when a fight occurs in the elevator, the traditional monitoring system cannot detect it in time, and the alarm may be triggered only when serious consequences such as elevator damage or casualties occur.

[0004] In the complex elevator operation environment, a single monitoring method is difficult to meet actual needs. On the one hand, the internal space of the elevator is complex and personnel activities are frequent. The traditional rule-based detection method cannot cope with various emergencies. On the other hand, elevator equipment has been in operation for a long time, and problems such as wear and aging of mechanical parts have gradually emerged. Early minor faults are difficult to be discovered in time, which may eventually lead to serious safety accidents. According to relevant safety accident statistics, in the past five years, there have been more than tens of thousands of safety accidents caused by elevator failures, causing a large number of casualties and property losses.

[0005] In addition, with the rapid development of technologies such as the Internet of Things and artificial intelligence, people's expectations for intelligent elevator operation and management are constantly increasing. They hope to be able to grasp the operating status of the elevator in real time and accurately, predict potential faults in advance, and detect and handle abnormal behaviors in a timely manner to ensure the safe operation of the elevator and improve user experience. However, the existing elevator monitoring technology is far from meeting these requirements. Therefore, in response to the above problems, an elevator abnormal behavior detection system based on a multimodal neural network is proposed. Summary of the invention

[0006] The purpose of the present invention is to provide an elevator abnormal behavior detection system based on a multimodal neural network to solve the problems raised in the above background technology.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] An elevator abnormal behavior detection system based on a multimodal neural network, comprising:

[0009] A data acquisition module, including a visual sensor array, an inertial measurement unit (IMU), and an audio acquisition device installed in the elevator car;

[0010] A data processing module is used to perform time-frequency domain conversion and feature vectorization processing on the original data;

[0011] Multimodal neural network model, including parallel convolutional neural network (CNN) branches, sequential neural network (RNN) branches, and voiceprint recognition network;

[0012] The feature fusion module adopts a dual fusion architecture that combines feature-level fusion and decision-level fusion based on the attention mechanism;

[0013] Abnormal behavior classifier, a three-level abnormality discrimination model built based on a hierarchical classification strategy;

[0014] The alarm and log module includes a dynamic threshold alarm unit and an abnormal behavior database.

[0015] As a preferred solution, the data acquisition module includes:

[0016] A combined visual sensor of an infrared thermal imaging camera and a visible light camera, with a spatial resolution of no less than 1920×1080@30fps;

[0017] A 9-axis IMU device consisting of a three-axis accelerometer, a gyroscope, and a magnetometer, with a sampling frequency set to 100 Hz ± 5%;

[0018] The directional microphone array is arranged at a preset angle θ with the elevator control panel to meet the sound field coverage requirement of 45°≤θ≤60°.

[0019] As a preferred solution, the data processing module includes the following improved algorithm:

[0020] The threshold function of the vibration signal denoising unit based on improved wavelet packet transform is defined as:

[0021] When the absolute value of the signal is greater than or equal to the high-order threshold Th, the output value is the original signal minus the product of the threshold and the sign function;

[0022] When the absolute value of the signal is less than the low-order threshold Tl, the output value is zero;

[0023] When the absolute value of the signal is between Tl and Th, the output value is the ratio of the original signal to the S-type smoothing function, where the smoothing coefficient σ controls the attenuation rate of the transition zone, and the intermediate value Tm between Th and Tl is used to calculate the exponential term;

[0024] The video feature extraction unit using the three-dimensional optical flow method has a motion vector calculation that satisfies the product of the partial derivative of the spatiotemporal pixel intensity in the x direction and the x velocity component, plus the product of the partial derivative in the y direction and the y velocity component, and the sum of the partial derivative in time equals zero;

[0025] The audio feature fusion unit uses a hybrid feature of Mel frequency cepstral coefficients MFCC and improved linear predictive coding LPC, where the MFCC feature weight α ranges from 0.6 to 0.8, and the LPC feature weight is 1-α.

[0026] As a preferred solution, the multimodal neural network model includes:

[0027] The visual branch adopts an improved 3D-ResNet architecture, and the output of its spatiotemporal attention module is the sigmoid activation value of the convolution result of the spatial convolution kernel and the temporal feature map, and the sum of the convolution result of the temporal convolution kernel and the spatial feature map;

[0028] The motion data processing branch uses a gated convolution-enhanced Bi-GRU network, whose gated update formula includes the linear transformation result of the previous hidden state and the current input, and the product of the one-dimensional convolution output and the learnable gating coefficient is superimposed and then activated by Sigmoid.

[0029] The voiceprint recognition branch adopts an improved DCGAN model, whose generator loss function consists of a weighted adversarial loss term and an L2 regularization term output by the pre-trained generator.

[0030] As a preferred solution, the feature fusion module includes:

[0031] Feature-level fusion uses an improved channel attention mechanism, whose weight is calculated as the transposed product of the query matrix and the key matrix adjusted by the dimension scaling factor, superimposed with a bias term and normalized by the Softmax function;

[0032] The decision-level fusion adopts the credibility-corrected DS evidence theory, and the basic probability distribution function is defined as the normalized result of the weighted probability distribution value of each evidence source and the category confidence score after being corrected by the credibility adjustment factor;

[0033] The final decision output adopts a gated fusion strategy, which concatenates the feature-level fusion result with the decision-level fusion result, transforms it through a trainable weight matrix, and outputs it through a Sigmoid activation function.

[0034] As a preferred solution, the three-level discrimination strategy of the abnormal behavior classifier includes:

[0035] The first-level physical layer detection uses dynamic threshold comparison, and triggers an alarm when the acceleration modulus exceeds the mean plus three times the standard deviation within the sliding time window;

[0036] The second-level behavior recognition uses an improved dynamic time warping (DTW) distance metric and introduces a time decay coefficient λ to perform exponential weighting on the distance matrix during path search.

[0037] The third-level compound decision adopts a deep reinforcement learning strategy, which updates the Q function by combining the learning rate η and the discount factor γ through the difference between the maximum Q value output by the target network and the current Q value.

[0038] As a preferred solution, the alarm and log module includes:

[0039] The dynamic threshold adjustment unit adopts a sliding window mechanism, and the window length T meets 30 seconds to 120 seconds;

[0040] The anomaly database adopts a spatiotemporal coding storage format, and the record fields include timestamp, three-dimensional spatial coordinates, multimodal feature vector set and feature fingerprint hash value;

[0041] The alarm interface supports RS485 and CAN bus dual protocol parallel transmission.

[0042] As a preferred solution, a model optimization module is also included, and the model optimization module includes:

[0043] The transfer learning method adopts a feature pyramid matching strategy and calculates the square sum of the Frobenius norm of each layer of the feature map of the source domain and the target domain as the transfer loss;

[0044] The federated learning framework uses differential privacy gradient aggregation, adding Gaussian noise after client-side gradient summation;

[0045] The feature dictionary update rule of the incremental learning unit is to retain the new features in the old dictionary whose distance from the class center exceeds the threshold τ, and select the top K features in order of importance.

[0046] As a preferred solution, the training method of the multimodal neural network model includes:

[0047] Adversarial training sample generation uses gradient sign attack to superimpose gradient direction perturbations on the original samples;

[0048] The multi-stage training process sets a joint loss function, which is a linear combination of classification loss, feature reconstruction loss and regularization term;

[0049] The model validation adopts multi-dimensional evaluation indicators, which integrates the weighted scores of F1 value, AUC and ROC curve.

[0050] As a preferred solution, the interface between the system and the elevator control system includes:

[0051] Hardware-isolated CAN bus communication channel, transmission delay is less than 10 milliseconds;

[0052] The triggering condition of the safety relay module is that the first level alarm and the second level or third level alarm are activated at the same time;

[0053] The UPS life of the independent power supply unit shall be no less than twice the average single operation time of the elevator and no less than 300 seconds.

[0054] It can be seen from the technical solution provided by the present invention that the elevator abnormal behavior detection system based on a multimodal neural network provided by the present invention has the following beneficial effects:

[0055] Significantly improve elevator operation safety:

[0056] Comprehensive and accurate anomaly detection: With the help of multimodal data collection, the integrated visual, motion and audio information greatly enriches the perception dimension of the elevator operation status; the multimodal neural network model can deeply explore the complex associations between different modal data and accurately identify various abnormal behaviors; for example, when detecting abnormal behaviors of people in the elevator, it can not only capture the movement posture of the people through visual sensors, but also combine audio information to determine whether there are abnormal shouts, greatly improving the detection accuracy, reducing missed reports and false alarms, and effectively ensuring passenger safety;

[0057] Timely and effective alarm mechanism: The alarm and log module adopts a dynamic threshold alarm unit, which can flexibly adjust the alarm threshold according to the real-time operation of the elevator to ensure that the alarm is issued as soon as an abnormality occurs; the alarm interface supports RS485 / CAN dual-protocol parallel transmission to ensure that the alarm signal is reliably and quickly transmitted to relevant personnel, so that emergency measures can be taken in time to reduce the probability of accidents;

[0058] Optimize elevator operation management process:

[0059] Reasonable maintenance plan: The health assessment module generates node health index and gives maintenance priority ranking through in-depth analysis of node historical fault data and real-time performance indicators. The management department can plan maintenance work in advance based on this result and reasonably allocate resources to elevators with low health index and urgent maintenance, thereby improving maintenance efficiency, extending elevator service life and reducing maintenance costs.

[0060] Improve operation efficiency: The interface between the system and the elevator control system realizes two-way data interaction. The detection system can promptly feed back abnormal information to the elevator control system so that it can make corresponding adjustments, such as preventing the faulty elevator from continuing to operate and reasonably dispatching normal elevators to ensure the efficient operation of the entire elevator system.

[0061] Technological innovation and advancement:

[0062] Innovative algorithm application: Innovative algorithms are used in multiple links such as data processing, feature fusion and abnormal behavior classification. For example, the vibration signal denoising algorithm based on improved wavelet packet transform in the data processing module can more effectively remove noise interference; the feature fusion module adopts a dual fusion architecture combining feature-level fusion and decision-level fusion based on the attention mechanism, which significantly improves the fusion effect; the deep reinforcement learning strategy of the abnormal behavior classifier enhances the ability to handle complex abnormal situations. These innovative algorithms improve system performance and make it stand out among similar technologies.

[0063] Advantages of multi-technology integration: Integrate advanced technologies such as multimodal neural networks and big data analysis to give full play to the advantages of each technology; multimodal neural networks conduct in-depth analysis of multi-source data, and big data analysis provides rich data support for model training and health assessment, so that the system has strong learning and adaptability, can continuously optimize the abnormal behavior detection effect, and adapt to the elevator operation monitoring needs in different scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0064] Figure 1 The present invention is a schematic diagram of the overall structure of an elevator abnormal behavior detection system based on a multimodal neural network. DETAILED DESCRIPTION

[0065] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0066] In order to better understand the above technical solution, the above technical solution will be described in detail below in conjunction with the accompanying drawings and specific implementation methods.

[0067] like Figure 1 As shown, an embodiment of the present invention provides an elevator abnormal behavior detection system based on a multimodal neural network, including a data acquisition module, a data processing module, a multimodal neural network model, a feature fusion module, an abnormal behavior classifier and an alarm and log module.

[0068] In this embodiment, the data acquisition module includes a visual sensor array, an inertial measurement unit IMU and an audio acquisition device installed in the elevator car;

[0069] The data acquisition module includes:

[0070] A combined visual sensor of an infrared thermal imaging camera and a visible light camera, with a spatial resolution of no less than 1920×1080@30fps;

[0071] A 9-axis IMU device consisting of a three-axis accelerometer, a gyroscope, and a magnetometer, with a sampling frequency set to 100 Hz ± 5%;

[0072] The directional microphone array is arranged at a preset angle θ with the elevator control panel to meet the sound field coverage requirement of 45°≤θ≤60°;

[0073] Furthermore, the data acquisition module of the present invention plays a vital role in the elevator abnormal behavior detection system based on multimodal neural network. It is responsible for collecting various types of data during the operation of the elevator and provides a basis for subsequent data processing and analysis. The module mainly includes a visual sensor array, an inertial measurement unit (IMU) and an audio acquisition device. The following is a detailed introduction to each component:

[0074] 1. Vision sensor array:

[0075] The visual sensor array uses a combination of infrared thermal imaging cameras and visible light cameras. This design can fully utilize the advantages of both cameras to capture visual information inside the elevator car from different angles.

[0076] Infrared thermal imaging camera: Using infrared thermal imaging technology, it can detect infrared radiation emitted by objects and convert it into thermal images. In an elevator environment, it can detect the thermal radiation of the human body and clearly capture the activities of people even in dim light, which helps to detect the location, posture and other information of people.

[0077] Visible light camera: can provide clear color images, reflecting the real scene inside the elevator car; it can capture the appearance, movements of people and the status of facilities inside the elevator; it complements the infrared thermal imaging camera to make the visual information more comprehensive and accurate;

[0078] Technical parameters: The spatial resolution of the combined visual sensor is no less than 1920×1080@30fps, which means it can capture high-resolution images at 30 frames per second, ensuring image clarity and smoothness, and providing high-quality data for subsequent video analysis;

[0079] 2. Inertial Measurement Unit (IMU):

[0080] The inertial measurement unit is composed of a three-axis accelerometer, a gyroscope, and a magnetometer. It is a 9-axis IMU device used to measure the motion state and attitude information of the elevator;

[0081] Three-axis accelerometer: It can measure the acceleration of the elevator in the three coordinate axes. By analyzing the acceleration data, it can understand the elevator's acceleration, deceleration, uniform speed and other operating conditions, as well as whether there is abnormal vibration or shaking;

[0082] Gyroscope: used to measure the angular velocity of the elevator. It can detect the rotational motion of the elevator, such as the tilt or twist of the elevator during operation.

[0083] Magnetometer: It can measure the direction of the earth's magnetic field, provide a direction reference for the elevator, and help determine the elevator's attitude and orientation;

[0084] Technical parameters: The sampling frequency of the IMU device is set to 100Hz±5%, that is, data is collected 100 times per second, and the error range is within ±5%; a higher sampling frequency can ensure that the collected motion data is more accurate and detailed, and timely capture the slight changes in the elevator operation process;

[0085] 3. Audio acquisition device:

[0086] The audio collection device uses a directional microphone array, which is mainly used to collect sound information in the elevator car, including the voices of people and the operating sounds of elevator equipment;

[0087] Layout position: The layout position of the directional microphone array forms a preset angle θ with the elevator control panel, meeting the sound field coverage requirement of 45°≤θ≤60°; this layout can effectively capture sounds from all directions in the elevator car, while reducing interference from external noise, ensuring high quality of the collected audio data;

[0088] Through the collection of the above visual, motion and audio data, the data acquisition module provides rich multimodal information for the elevator abnormal behavior detection system, laying a solid foundation for subsequent anomaly detection and analysis.

[0089] In this embodiment, the data processing module is used to perform time-frequency domain conversion and feature vectorization processing on the original data;

[0090] Among them, the data processing module includes the following improved algorithms:

[0091] The threshold function of the vibration signal denoising unit based on improved wavelet packet transform is defined as:

[0092] When the absolute value of the signal is greater than or equal to the high-order threshold Th, the output value is the original signal minus the product of the threshold and the sign function;

[0093] When the absolute value of the signal is less than the low-order threshold Tl, the output value is zero;

[0094] When the absolute value of the signal is between Tl and Th, the output value is the ratio of the original signal to the S-type smoothing function, where the smoothing coefficient σ controls the attenuation rate of the transition zone, and the intermediate value Tm between Th and Tl is used to calculate the exponential term;

[0095] The video feature extraction unit using the three-dimensional optical flow method has a motion vector calculation that satisfies the product of the partial derivative of the spatiotemporal pixel intensity in the x direction and the x velocity component, plus the product of the partial derivative in the y direction and the y velocity component, and the sum of the partial derivative in time equals zero;

[0096] The audio feature fusion unit uses the hybrid features of Mel frequency cepstral coefficients MFCC and improved linear predictive coding LPC, where the MFCC feature weight α ranges from 0.6 to 0.8, and the LPC feature weight is 1-α;

[0097] Furthermore, the data processing module plays an important role in the elevator abnormal behavior detection system based on multimodal neural network. It mainly uses a variety of improved algorithms to perform time-frequency domain conversion and feature vectorization processing on the raw data from the data acquisition module, providing high-quality input data for the subsequent multimodal neural network model, thereby ensuring the accuracy and reliability of the system's detection of elevator abnormal behavior; the following is a detailed and specific description of the module:

[0098] Overall function overview:

[0099] The core task of the data processing module is to convert the raw data collected during the operation of the elevator, including the image data of the visual sensor array, the motion data of the inertial measurement unit (IMU), and the audio data of the audio acquisition device, into feature vectors that can be effectively processed by the multimodal neural network model; through a series of operations such as cleaning, denoising, feature extraction and fusion of the raw data, the module can accurately mine the key information in the data, eliminate noise interference, provide a solid data foundation for abnormal behavior detection, and greatly improve the accuracy of subsequent model analysis and judgment;

[0100] Data source:

[0101] The data of this module mainly comes from the data acquisition module, which specifically covers the following categories:

[0102] Visual sensor array data: collected by a combination of infrared thermal imaging cameras and visible light cameras; this data contains real-time scene images inside the elevator car, which can be used to analyze information such as the activities and postures of people and the status of facilities inside the elevator; its spatial resolution is no less than 1920×1080@30fps, providing high-resolution images at a speed of 30 frames per second, providing a rich and high-quality data foundation for subsequent video analysis;

[0103] IMU motion data: collected by a 9-axis IMU device consisting of a three-axis accelerometer, gyroscope, and magnetometer; these data record information such as the acceleration, angular velocity, and magnetic field direction of the elevator during operation, which can be used to determine the elevator's motion state, whether there is abnormal vibration or tilt, etc.; the sampling frequency is set to 100Hz±5%, which can capture subtle changes in the elevator's operation in a timely and accurate manner;

[0104] Audio collection device data: collected through a directional microphone array; the array is arranged at a preset angle θ (45°≤θ≤60°) with the elevator control panel, which can effectively collect sounds in the elevator car, including human voices, equipment operation sounds, etc.; these audio data help detect abnormal sounds during elevator operation and provide audio-dimensional information support for abnormal behavior detection;

[0105] Multimodal data processing algorithms:

[0106] The data processing module uses a variety of improved algorithms to process data of different modes to extract key features. The specific contents are as follows:

[0107] Vibration signal noise reduction algorithm:

[0108] Vibration signal denoising is achieved based on improved wavelet packet transform, and its threshold function is: ,in, is the input signal value; is a high-order threshold, used to define the boundary between strong signals and noise; is a low-order threshold, used to distinguish obvious noise; ; is the smoothing coefficient, and its value range is , used to control the smoothness of the signal in the transition interval; is the wavelet threshold function control coefficient, and its value range is , determines the degree of signal processing; is a symbolic function, when hour, ,when hour, ,when hour, , used to determine the adjustment direction when processing strong signals;

[0109] During elevator operation, the vibration signal collected by the IMU is susceptible to noise interference. This threshold function uses different processing methods for signals of different intensities to effectively remove noise while retaining the key features of the signal to the greatest extent, providing reliable data for subsequent motion analysis.

[0110] Video feature extraction algorithm:

[0111] The three-dimensional optical flow method is used to extract video features, and its motion vector calculation satisfies: ,in, is the pixel intensity in the spatiotemporal domain, which is about the spatial coordinates and time Function of , , Respectively about , , The partial derivative of is used to describe the rate of change of pixel intensity in the spatial and temporal directions; , They are and The motion velocity component of the direction can be solved by this formula to obtain the motion vector of the object in the video, and then the motion characteristics of the people or objects in the elevator can be extracted to provide a basis for subsequent behavior analysis;

[0112] Audio feature fusion algorithm:

[0113] The audio feature fusion unit uses a hybrid feature of Mel frequency cepstral coefficients (MFCC) and improved linear predictive coding (LPC): ,in, is the feature fusion weight, and its value range is , used to adjust and The proportion of mixed features; It can effectively reflect the spectral envelope characteristics of audio and performs well in tasks such as speech recognition; improved It has unique advantages in predicting audio signals and can better capture the detailed features of audio. By integrating the two, the extracted audio features can more comprehensively and accurately reflect the audio information in the elevator and enhance the recognition of abnormal sounds.

[0114] Feature vectorization and data integration:

[0115] After being processed by the above algorithm, data of different modes are converted into feature vectors. In order for the multimodal neural network model to better process these data, the feature vectors need to be further vectorized and integrated.

[0116] For visual features, the motion features extracted by the three-dimensional optical flow method and the features obtained by other related image processing algorithms are combined and encoded to form a visual feature vector suitable for neural network input; for IMU motion data, the acceleration, angular velocity and other features are normalized and combined into a motion feature vector in a certain order; for audio features, the features after MFCC and LPC fusion are quantized and encoded to generate an audio feature vector;

[0117] Finally, the feature vectors of different modalities are spliced ​​or fused in a specific way to form a comprehensive feature vector containing multimodal information, which serves as the final output of the data processing module and provides comprehensive and high-quality input data for the subsequent multimodal neural network model.

[0118] Application value of the module:

[0119] The data processing module significantly improves the quality and availability of data through fine processing and feature extraction of multimodal raw data; it provides accurate and effective input for the multimodal neural network model, greatly improving the performance of the elevator abnormal behavior detection system; by eliminating noise interference and mining key features, it can more accurately identify abnormal behaviors during elevator operation, providing strong support for ensuring the safe operation of elevators; at the same time, the efficient operation of this module also optimizes the data processing process of the entire system, improves the operating efficiency and stability of the system, and provides key technical support for the intelligent management and maintenance of urban elevators.

[0120] In this embodiment, the multimodal neural network model includes a parallel convolutional neural network CNN branch, a temporal neural network RNN ​​branch and a voiceprint recognition network;

[0121] Among them, the multimodal neural network model includes:

[0122] The visual branch adopts an improved 3D-ResNet architecture, and the output of its spatiotemporal attention module is the sigmoid activation value of the convolution result of the spatial convolution kernel and the temporal feature map, and the sum of the convolution result of the temporal convolution kernel and the spatial feature map;

[0123] The motion data processing branch uses a gated convolution-enhanced Bi-GRU network, whose gated update formula includes the linear transformation result of the previous hidden state and the current input, and the product of the one-dimensional convolution output and the learnable gating coefficient is superimposed and then activated by Sigmoid.

[0124] The voiceprint recognition branch adopts an improved DCGAN model, whose generator loss function is composed of the adversarial loss term and the L2 regularization term weighted by the output of the pre-trained generator;

[0125] The training methods of the multimodal neural network model include:

[0126] Adversarial training sample generation uses gradient sign attack to superimpose gradient direction perturbations on the original samples;

[0127] The multi-stage training process sets a joint loss function, which is a linear combination of classification loss, feature reconstruction loss and regularization term;

[0128] The model validation uses multi-dimensional evaluation indicators, integrating the weighted scores of F1 value, AUC and ROC curve;

[0129] Furthermore, the multimodal neural network model plays a core role in the elevator abnormal behavior detection system based on the multimodal neural network. It can deeply analyze and learn the multimodal features output by the data processing module, and mine the intrinsic correlation between different modal data, so as to accurately identify abnormal behaviors during elevator operation. The following is a detailed description of the model:

[0130] Overview of the overall architecture:

[0131] The multimodal neural network model adopts a parallel branch structure, which mainly includes a visual branch, a motion data processing branch, and a voiceprint recognition branch. Each branch processes data of different modes and extracts the key features of each mode. After that, the features of each branch are fused through the feature fusion module to comprehensively utilize multimodal information and improve the accuracy and reliability of abnormal behavior detection.

[0132] Detailed introduction of each branch:

[0133] Visual branch:

[0134] The visual branch uses an improved 3D-ResNet architecture, which introduces a spatiotemporal attention module based on the traditional 3D-ResNet, which can better capture the spatiotemporal features in video data; the output of its spatiotemporal attention module is: ,in, It is the output of the spatiotemporal attention module, which is used to adjust the weight of the feature map and highlight important spatiotemporal features; is the spatial convolution kernel, used to extract spatial features; It is the temporal convolution kernel, which is used to capture changes in the time dimension; It is a temporal feature, which reflects the temporal dynamic information of the video; It is a spatial feature, reflecting the spatial structure of the video screen; Represents the convolution operation, which is used to perform convolution operations on features to extract higher-level features; is the Sigmoid activation function, which maps the output value to interval, thereby assigning a weight to each feature element;

[0135] Through the spatiotemporal attention module, the model can adaptively focus on important spatiotemporal areas in the video, improve the ability to extract visual information such as people's movements and postures, and provide more accurate visual features for subsequent abnormal behavior analysis;

[0136] Motion data processing branch:

[0137] The motion data processing branch uses a Bi-GRU network enhanced by gated convolution. The Bi-GRU (bidirectional gated recurrent unit) network can process sequence data while taking into account past and future information, and is suitable for processing motion data collected by IMU. On this basis, gated convolution is introduced to enhance its feature extraction capability. The gated update formula is: ,in, It is the gate update value, which is used to control the transmission and update of information; is the Sigmoid activation function, which maps the output value to The interval determines the degree to which information passes; is a learnable weight matrix used to adjust the weight of input information; It is the hidden state of the previous moment, which contains the past information; is the input at the current moment, that is, the motion data at the current moment; It is a learnable gating coefficient used to adjust the influence of the convolution operation; Conv1D Express Performing one-dimensional convolution operations can extract local features from motion data;

[0138] The gated convolution-enhanced Bi-GRU network can better capture the temporal and local features in motion data, accurately analyze the motion state of the elevator, such as acceleration, deceleration, shaking, etc., and provide strong motion feature support for abnormal behavior detection;

[0139] Voiceprint recognition branch:

[0140] The voiceprint recognition branch uses an improved DCGAN (Deep Convolutional Generative Adversarial Network) model; DCGAN consists of a generator and a discriminator, and learns the distribution of data through adversarial training; the improved DCGAN model is optimized on the traditional basis, and its generator loss function is: ,in, is the generator loss function, which is used to measure the performance of the generator; Represents the expectation and statistically averages the data generated by the generator; is the discriminator, used to determine whether the data is real or generated; is a generator used to generate fake data; is random noise, which is used as the input of the generator; It is a pre-trained generator that learns certain patterns through pre-training; is the transfer learning weight factor, the value range can be adjusted according to the actual situation to balance the impact of the two parts of loss; represents the L2 norm, which is used to measure the difference between the data generated by the generator and the data generated by the pre-trained generator;

[0141] The improved DCGAN model can learn the characteristic distribution of sounds in the elevator and generate realistic sound samples. At the same time, it can use the knowledge of the pre-trained generator through transfer learning to improve the ability to recognize abnormal sounds, such as abnormal noise of elevator equipment and abnormal shouting of people.

[0142] Feature fusion and abnormal behavior identification:

[0143] The features extracted by each branch are fused through the feature fusion module, which adopts a dual fusion architecture combining feature-level fusion based on the attention mechanism and decision-level fusion; the feature-level fusion adopts an improved channel attention mechanism to assign weights according to the importance of each channel feature and enhance the expression of important features; the decision-level fusion adopts the DS evidence theory with credibility correction to integrate the decision results of each branch and improve the reliability of the decision;

[0144] The fused features are input into the abnormal behavior classifier. The three-level abnormality discrimination model constructed based on the hierarchical classification strategy judges the running status of the elevator, identifies whether there is abnormal behavior and the type of abnormal behavior, and provides a basis for subsequent alarm and maintenance.

[0145] Application value of the model:

[0146] The multimodal neural network model processes data of different modes in parallel and performs effective feature fusion, making full use of information from multiple aspects such as vision, motion and voiceprint, improving the accuracy and robustness of abnormal behavior detection of elevators. It can timely detect various abnormal situations during the operation of elevators, providing reliable technical support for ensuring the safe operation of elevators, and also providing a strong decision-making basis for the intelligent management and maintenance of elevators.

[0147] Furthermore, the training method of the multimodal neural network model is crucial in the elevator abnormal behavior detection system based on the multimodal neural network, which directly affects the model's recognition ability and accuracy of elevator abnormal behavior; the following will elaborate on the training method of the model:

[0148] Physical training ideas:

[0149] The training of this multimodal neural network model aims to enable the model to learn the complex relationship between multimodal data in order to accurately identify abnormal behaviors in elevator operation. The training process comprehensively uses adversarial training sample generation, multi-stage training process and multi-dimensional evaluation indicators to improve the performance and generalization ability of the model.

[0150] Adversarial training sample generation:

[0151] Adversarial training samples are generated using gradient sign attacks, with the goal of enhancing the robustness of the model so that it can resist adversarial attacks and improve recognition capabilities in complex environments. The specific formula is: ,in, Adversarial samples are samples generated by adding perturbations to the original samples. is the original sample, i.e., the multimodal feature data obtained from the data acquisition module and processed by the data processing module; is the perturbation intensity coefficient, which controls the size of the added perturbation and can be adjusted according to the actual situation; sign is the sign function, which determines the direction of the perturbation according to the positive or negative gradient; is the loss function about The gradient of the loss function reflects the gradient of the input sample Changing trends; are model parameters, representing the structure and weight of the model; is a label indicating the true category to which the sample corresponds, such as normal operation or abnormal behavior type;

[0152] By generating adversarial samples and adding them to the training set, the model can learn the robust features of the data during the training process, thereby improving the ability to identify various abnormal situations;

[0153] Multi-stage training process:

[0154] The multi-stage training process sets a joint loss function to balance the goals of different training stages and improve the overall performance of the model; the formula of the joint loss function is as follows: ,in, is the total loss function, which comprehensively considers the losses of classification, reconstruction and regularization; is the classification loss weight, and its value range is , which reflects the importance of the classification task in the total loss; It is the classification loss, which measures the difference between the category predicted by the model and the true label. Common classification loss functions include cross entropy loss, etc. is the reconstruction loss weight, which is used to adjust the proportion of reconstruction loss in the total loss; It is the feature reconstruction loss, which reflects the model's ability to reconstruct the input features and can enable the model to learn the essential characteristics of the data; 1 is the regularization loss weight, and its value range is , used to control the influence of the regularization term; Regularization terms, such as L1 or L2 regularization, can prevent the model from overfitting and improve the generalization ability of the model;

[0155] In the multi-stage training process, the weights of each loss item can be adjusted according to the actual situation at different stages to gradually optimize the performance of the model; for example, in the initial stage, more emphasis can be placed on classification loss to allow the model to quickly learn the classification features of the data; in the subsequent stages, the weights of reconstruction loss and regularization loss can be appropriately increased to improve the model's feature extraction and generalization capabilities;

[0156] Model Validation:

[0157] Model validation uses multi-dimensional evaluation indicators to comprehensively evaluate the performance of the model; the evaluation indicator formula is: ,in, For model validation score, comprehensive consideration is given to Fraction, and Three indicators; The F1 score is the harmonic mean of precision and recall, which is used to measure the overall performance of the model in classification tasks; is the area under the curve, The area under the curve reflects the classification and ranking capabilities of the model; The receiver operating characteristic curve describes the relationship between the true positive rate and false positive rate of the model at different thresholds;

[0158] Through multi-dimensional evaluation indicators, the performance of the model can be evaluated from different angles to avoid the limitations of a single indicator. During the training process, the model parameters and training strategies can be adjusted according to the verification score to improve the final performance of the model.

[0159] Summary of the training process:

[0160] The training process of the multimodal neural network model combines adversarial training sample generation, multi-stage training process and multi-dimensional model verification; the robustness of the model is enhanced by generating adversarial samples, and the model performance is optimized in multi-stage training using a joint loss function. Finally, the model is comprehensively evaluated through multi-dimensional evaluation indicators to ensure that the model can accurately identify abnormal behaviors in elevator operation and provide reliable protection for the safe operation of elevators.

[0161] In this embodiment, the feature fusion module adopts a dual fusion architecture combining feature-level fusion and decision-level fusion based on the attention mechanism;

[0162] Among them, the feature fusion module includes:

[0163] Feature-level fusion uses an improved channel attention mechanism, whose weight is calculated as the transposed product of the query matrix and the key matrix adjusted by the dimension scaling factor, superimposed with a bias term and normalized by the Softmax function;

[0164] The decision-level fusion adopts the credibility-corrected DS evidence theory, and the basic probability distribution function is defined as the normalized result of the weighted probability distribution value of each evidence source and the category confidence score after being corrected by the credibility adjustment factor;

[0165] The final decision output adopts a gated fusion strategy, which concatenates the feature-level fusion result with the decision-level fusion result, transforms it through a trainable weight matrix, and outputs it through a Sigmoid activation function.

[0166] Furthermore, the feature fusion module plays a key role in the elevator abnormal behavior detection system based on multimodal neural network. Its main responsibility is to effectively fuse the features extracted by different branches in the multimodal neural network model, fully explore the complementary information between the modal data, and improve the system's detection accuracy of elevator abnormal behavior. The following is a detailed description of the module:

[0167] Overall function overview:

[0168] The feature fusion module adopts a dual fusion architecture that combines feature-level fusion and decision-level fusion based on the attention mechanism. Feature-level fusion aims to integrate the original features of each modality to generate fusion features containing rich multimodal information. Decision-level fusion comprehensively considers the decision results made by each modal branch and finally outputs a comprehensive and reliable detection result. Through this dual fusion architecture, the module can maximize the advantages of multimodal data and significantly improve the accuracy and reliability of elevator abnormal behavior detection.

[0169] Feature-level fusion:

[0170] Feature-level fusion adopts an improved channel attention mechanism, which adaptively highlights the feature channels that are more critical to abnormal behavior detection by calculating the attention weights on the channel dimension; its weight calculation is: ,in, is the channel attention weight, which is used to adjust the importance of each channel feature; is the query matrix, which extracts specific information from the input features for query; is the key matrix, which is used to provide key values ​​related to the query so as to find the corresponding information in the features; It represents the transpose multiplication of the query matrix and the key matrix. Through this operation, the correlation information between different channels can be obtained; It is the dimension scaling factor, which is used to normalize the calculation results to avoid the stability of weight calculation affected by the dimension size; It is a bias term that increases the learning ability of the model so that the model can better fit the data; is the Softmax function, mapping the above calculation results to Interval, get the normalized weight value of each channel, the larger the weight value, the more important the channel feature is;

[0171] In this way, feature-level fusion can re-weight the features of each channel according to the characteristics of different modal data, generate more representative fusion features, and provide better feature input for subsequent decision-making processes;

[0172] Decision-level fusion:

[0173] The decision-level fusion adopts the credibility-corrected DS evidence theory; DS evidence theory is an uncertain reasoning method that can integrate information from multiple evidence sources to make a comprehensive decision; in this module, the credibility is corrected to meet the actual needs of elevator abnormal behavior detection; the basic probability distribution function is defined as: ,in, For Category The basic probability distribution of , which indicates the probability that the category is considered correct in the fusion decision; is the number of evidences, that is, the number of decision results output by different branches (such as the visual branch, motion data processing branch, and voiceprint recognition branch) in the multimodal neural network model; For the The weight of each piece of evidence reflects the relative importance of each branch decision result, and its value can be determined through training or experience according to the actual situation; For the Evidence pair category The basic probability distribution of each branch is to judge the category separately is the probability of being correct; is the credibility adjustment factor, and its value range is , used to adjust the class The degree to which other confidence scores influence the final decision; For Category The confidence score of the category is calculated based on the decision results of each branch and related information. Credibility of judgments; To identify the framework, all possible categories were included; for When calculating the denominator, it is necessary to sum the basic probability distributions of all subsets in the recognition framework to ensure The value of The sum of the basic probability distribution of all categories is 1;

[0174] Through the credibility-corrected DS evidence theory, decision-level fusion can comprehensively consider the decision results and credibility of each modal branch, avoid the impact of a single branch's decision-making error on the final result, and thus improve the accuracy and reliability of the decision;

[0175] Final decision output:

[0176] In order to further optimize the fusion effect, the feature fusion module adopts the gated fusion strategy to output the final decision; its formula is: ,in, The final decision output is the final judgment result given by the feature fusion module on whether the elevator has abnormal behavior; is the Sigmoid activation function, which maps the output value to Interval, which is convenient for interpretation and judgment of the results. For example, if 0.5 is used as the threshold, it is judged as abnormal behavior if it is greater than 0.5, and it is judged as normal operation if it is less than 0.5; It is a trainable weight matrix, which learns the optimal weight relationship between the feature-level fusion results and the decision-level fusion results through training to achieve the optimal combination of the two; Indicates the result of feature-level fusion and decision-level fusion results Splice to form a vector containing two fusion information as Input;

[0177] Through the gated fusion strategy, the advantages of feature-level fusion and decision-level fusion are fully utilized, making the final decision result more accurate and stable, providing a reliable output for the elevator abnormal behavior detection system;

[0178] Application value of the module:

[0179] The feature fusion module effectively integrates the information of multimodal data through a unique dual fusion architecture, significantly improving the performance of the elevator abnormal behavior detection system; it can more accurately identify abnormal behavior during elevator operation, reduce misjudgments and missed judgments, and provide key support for ensuring the safe operation of elevators; at the same time, the module's efficient fusion strategy also provides ideas and methods that can be used as reference for other multimodal data analysis applications.

[0180] In this embodiment, the abnormal behavior classifier is based on a three-level abnormal discrimination model constructed based on a hierarchical classification strategy. The three-level discrimination strategy of the abnormal behavior classifier includes:

[0181] The first-level physical layer detection uses dynamic threshold comparison, and triggers an alarm when the acceleration modulus exceeds the mean plus three times the standard deviation within the sliding time window;

[0182] The second-level behavior recognition uses an improved dynamic time warping (DTW) distance metric and introduces a time decay coefficient λ to perform exponential weighting on the distance matrix during path search.

[0183] The third-level compound decision adopts a deep reinforcement learning strategy, which updates the Q function by combining the learning rate η and the discount factor γ with the difference between the maximum Q value output by the target network and the current Q value;

[0184] Furthermore, the abnormal behavior classifier plays a key role in the elevator abnormal behavior detection system based on multimodal neural network. It accurately judges the elevator operation status based on the data processed by the multimodal neural network model and integrated by the feature fusion module, and identifies whether there is abnormal behavior and the type of abnormal behavior. The following is a detailed description of the abnormal behavior classifier:

[0185] Overall function overview:

[0186] The abnormal behavior classifier is built based on a hierarchical classification strategy. Through a three-level discrimination model, it gradually and deeply analyzes the elevator operation data, captures abnormal information from different levels, and finally outputs accurate abnormal behavior judgment results, providing a solid basis for subsequent alarms and maintenance measures. This hierarchical structure can improve the accuracy and reliability of classification and effectively deal with complex and changeable situations in elevator operation.

[0187] Three-level discrimination strategy:

[0188] First level physical layer detection:

[0189] The first-level physical layer detection mainly monitors the basic physical parameters of elevator operation, using a dynamic threshold comparison method; its judgment formula is: ,in, It is the alarm signal detected by the first level physical layer. When , it indicates that an abnormality is detected and an alarm is triggered; when When the elevator is in normal physical operation state; is a time variable, representing the detection moment; It is a sliding time window. By observing the physical parameters within a certain time range, it can avoid misjudgment due to instantaneous fluctuations. The window duration can be set according to the actual elevator operation characteristics; for Physical signals at a given moment, such as a vector of physical quantities such as acceleration and angular velocity obtained from the IMU; express The L2 norm of is used to measure the size of the physical signal vector; For physical signals in a sliding time window The mean value within reflects the average level of the physical parameter when the elevator is operating normally; is the standard deviation, which reflects the degree of fluctuation of the physical parameter around the mean; when in the sliding time window There is a moment in time Make the L2 norm of the physical signal vector greater than the mean Plus three times the standard deviation When it is judged as abnormal, an alarm signal is triggered ;

[0190] This dynamic threshold comparison can quickly capture significant anomalies in physical parameters during elevator operation, providing preliminary warnings for timely discovery of potential problems.

[0191] Second level behavior recognition:

[0192] The second-level behavior recognition mainly focuses on the analysis of the behavior patterns of people and equipment in the elevator, using an improved DTW (dynamic time warping) distance measurement method; its calculation formula is: ,in, To improve the DTW distance, it is used to measure the similarity between two behavior sequences; For all possible alignment path sets, during the calculation process, it is necessary to find an optimal alignment path so that the corresponding elements of the two behavior sequences on the time axis match as much as possible; are the elements in the alignment path, representing the index positions in the two behavior sequences respectively; and are the elements at the corresponding index positions in the two behavior sequences, which can be the action feature vectors obtained through visual analysis or the sound feature vectors obtained from audio analysis; Used to calculate the distance between two corresponding elements and measure their degree of difference; is the time attenuation coefficient, and its value range can be adjusted according to the actual situation. It reflects that as the time interval increases, the importance of the corresponding relationship between the two elements gradually decreases; As a weight term, the distances of elements at different time intervals are weighted, so that when calculating the total distance, more attention is paid to the matching of elements that are close in time;

[0193] By calculating the improved DTW distance, the currently collected behavior sequence is compared with the pre-set normal behavior pattern library. If the distance exceeds a certain threshold, it is determined that abnormal behavior exists, thereby achieving accurate identification of abnormal behavior in the elevator;

[0194] Third level compound decision:

[0195] The third-level compound decision adopts a deep reinforcement learning strategy to optimize the decision-making process through continuous trial and error and learning to improve the ability to judge complex abnormal situations; its update formula is: ,in, For the current state Take action of Value, representing the expected long-term cumulative reward for performing this action in this state; is the learning rate, and its value range is generally It controls each learning The step size of the value update. If the learning rate is too large, the algorithm may become unstable, while if the learning rate is too small, the learning process will become slow. Instant reward: When the elevator status changes, an instant reward value is given according to the change. For example, a negative reward is given when abnormal behavior is detected, and a positive reward is given when it is running normally, so as to guide the decision-making towards a better direction. is the discount factor, and its value range is usually It determines the importance of future rewards in current decisions. The closer it is to 1, the more attention is paid to future rewards and long-term optimization effects; The closer it is to 0, the more attention is paid to immediate rewards; For the target network in state Take action of The target network has the same structure as the current network but with relatively slow parameter updates, which is used to stabilize the learning process. Indicates in status Next, for all possible actions of The value takes the maximum value, that is, from the perspective of the target network, choose the The best action value;

[0196] Through deep reinforcement learning strategies, the abnormal behavior classifier can dynamically adjust the decision-making strategy according to the changing elevator operation status, and improve the ability to identify and handle complex abnormal behaviors;

[0197] Application value of the module:

[0198] The abnormal behavior classifier uses a hierarchical three-level discrimination strategy to comprehensively and deeply analyze the elevator operation data, greatly improving the accuracy and reliability of elevator abnormal behavior detection; it can timely and accurately detect various abnormal situations in the elevator operation process, provide strong guarantees for the safe operation of the elevator, reduce the safety risks and economic losses caused by elevator failures, and also provide clear guidance for elevator maintenance and management, improving the efficiency and quality of maintenance work.

[0199] In this embodiment, the alarm and log module includes a dynamic threshold alarm unit and an abnormal behavior database, and the alarm and log module includes:

[0200] The dynamic threshold adjustment unit adopts a sliding window mechanism, and the window length T meets 30 seconds to 120 seconds;

[0201] The anomaly database adopts a spatiotemporal coding storage format, and the record fields include timestamp, three-dimensional spatial coordinates, multimodal feature vector set and feature fingerprint hash value;

[0202] The alarm interface supports RS485 and CAN bus dual protocol parallel transmission;

[0203] Furthermore, the alarm and log module plays a vital role in the elevator abnormal behavior detection system based on multimodal neural network. It is responsible for timely feedback on detected abnormal situations and recording relevant information in detail, providing key support for subsequent troubleshooting and system optimization. The following is a detailed description of the module:

[0204] Overall function overview:

[0205] The alarm and log module has two core functions: first, when the system detects abnormal behavior of the elevator, it quickly sends out an alarm signal to notify relevant personnel to handle it in time; second, it records the time, relevant data and detection results of the abnormal behavior in detail to form a complete log for subsequent analysis and tracing. Through these two functions, the module can effectively ensure the safe operation of the elevator and provide strong data support for elevator maintenance and management;

[0206] Alarm function:

[0207] Dynamic Threshold Alarm Unit:

[0208] The alarm and log module adopts a dynamic threshold alarm mechanism to adapt to the complex and changeable situations during elevator operation; the dynamic threshold adjustment unit adopts a sliding window mechanism, and the window length T satisfies 30s≤T≤120s; within this time window, the system monitors the elevator's operating parameters and abnormal behavior detection results in real time; for the first-level physical layer detection and subsequent levels of detection results, when the relevant indicators exceed the threshold calculated according to the dynamic threshold algorithm, the system triggers an alarm;

[0209] For example, in the first-level physical layer detection, based on the average value of the physical signal within the sliding time window T and standard deviation Determine dynamic thresholds (such as + ), once the relevant measure of the physical signal is detected (such as ) exceeds the threshold, an alarm signal is triggered; this method of dynamically adjusting the threshold can better cope with various normal fluctuations and abnormal changes in the elevator operation process, and reduce false alarms and missed alarms;

[0210] Alarm interface:

[0211] The alarm interface supports RS485 / CAN dual-protocol parallel transmission, ensuring that the alarm signal can be stably and quickly transmitted to related equipment and management systems; RS485 and CAN protocols are widely used in the field of industrial control, with the advantages of high reliability, long transmission distance, and strong anti-interference ability; dual-protocol parallel transmission further improves the reliability of alarm transmission. Even if one of the protocols fails, the other protocol can still ensure the transmission of alarm information; through the alarm interface, the alarm signal can be sent to the monitoring terminal of the elevator manager, the property management system and the relevant emergency rescue department in a timely manner, so that they can quickly take measures to deal with abnormal elevator conditions;

[0212] Log function:

[0213] Abnormal database:

[0214] The abnormality database adopts the time-space coding storage format to record the detailed information of abnormal behavior of elevators; the record fields include: ,in, Accurately record the time when abnormal behavior occurs to the specific moment, providing a basis for subsequent analysis of the time pattern of abnormal occurrence; The spatial coordinates of the elevator help determine the specific location of the elevator when an abnormality occurs, making it easier for maintenance personnel to locate the elevator quickly. It is a set of variable values ​​related to abnormal behavior. These variables can be physical parameters of elevator operation (such as acceleration, speed, etc.), output results of abnormal behavior classifier (such as identification of abnormal type), etc., which comprehensively records data related to abnormality. It is a multimodal feature vector, which is collected by the data acquisition module and processed by the data processing module and the multimodal neural network model, and contains visual, motion, audio and other information. is the characteristic fingerprint function value of the multimodal feature vector, which is obtained by performing hash calculation on the multimodal feature vector and is used to uniquely identify the group of feature vectors, facilitating rapid retrieval and comparison in the database;

[0215] Through this spatiotemporal coding storage format, the anomaly database can efficiently store and manage a large number of anomaly records, providing rich data resources for subsequent data analysis and troubleshooting;

[0216] Application value of the module:

[0217] The alarm and log module can quickly notify relevant personnel when the elevator exhibits abnormal behavior through timely and accurate alarm functions, thereby minimizing the probability of safety accidents and ensuring the safety of passengers' lives. At the same time, detailed log records provide comprehensive fault information for elevator maintenance personnel, helping them to quickly diagnose the cause of the fault, develop effective maintenance plans, improve elevator maintenance efficiency, and extend the service life of the elevator. In addition, in-depth analysis of log data can also provide data support for the optimization of the elevator abnormal behavior detection system, continuously improve the system's detection performance and reliability, and provide strong guarantees for the safe operation and intelligent management of urban elevators.

[0218] In this embodiment, the system further includes a model optimization module, which includes:

[0219] The transfer learning method adopts a feature pyramid matching strategy and calculates the square sum of the Frobenius norm of each layer of the feature map of the source domain and the target domain as the transfer loss;

[0220] The federated learning framework uses differential privacy gradient aggregation, adding Gaussian noise after client-side gradient summation;

[0221] The feature dictionary update rule of the incremental learning unit is to retain the new features in the old dictionary whose distance from the class center exceeds the threshold τ, and select the top K features in order of importance;

[0222] Furthermore, the model optimization module plays a key role in the elevator abnormal behavior detection system based on multimodal neural network. It can continuously improve the performance of the multimodal neural network model and enhance the accuracy and robustness of the system in detecting elevator abnormal behavior. The following is a detailed description of the module:

[0223] Overall function overview:

[0224] The core task of the model optimization module is to continuously adjust the parameters of the multimodal neural network model through a variety of optimization strategies and algorithms, so that it can better adapt to the complex situations in the elevator operation process, improve the ability to identify abnormal behaviors, and reduce the false alarm rate and missed alarm rate. This module comprehensively considers the model's training effect, generalization ability, and the efficiency of computing resources to achieve the optimization of model performance.

[0225] Adaptive learning rate adjustment strategy:

[0226] The model optimization module adopts an adaptive learning rate adjustment strategy to balance the convergence speed and stability during model training;

[0227] Through this adaptive learning rate adjustment strategy, the model can dynamically adjust the learning rate according to the changes in the loss function during training, thereby improving training efficiency and model performance;

[0228] Model pruning and quantization:

[0229] In order to improve the computational efficiency and resource utilization of the model, the model optimization module uses model pruning and quantization technology;

[0230] Model pruning:

[0231] Model pruning analyzes the importance of each parameter in the model and removes parameters that have little impact on model performance. Specifically, an amplitude-based pruning method is used. For a weight parameter W in a neural network, if its absolute value is less than a preset threshold τ, the parameter is set to 0.

[0232] Through model pruning, the number of model parameters is reduced, the computational complexity of the model is reduced, and the problem of overfitting is avoided to a certain extent;

[0233] Model Quantization:

[0234] Model quantization converts floating-point parameters in the model into low-precision integer representations to reduce the storage and computation requirements of the model;

[0235] During the inference process, quantized integer parameters are used for calculation, which greatly improves the calculation speed and resource utilization of the model;

[0236] Transfer learning and knowledge distillation:

[0237] In order to make full use of existing data and model knowledge, the model optimization module uses transfer learning and knowledge distillation technology;

[0238] Transfer Learning:

[0239] Transfer learning transfers model parameters pre-trained in other related fields or tasks to the elevator abnormal behavior detection model. For example, in the visual branch, the parameters of the 3D-ResNet model pre-trained on a large-scale image dataset can be used as the initial parameters, and then fine-tuned on the elevator vision data. This can speed up the convergence of the model and improve the performance of the model on a small dataset.

[0240] Knowledge Distillation:

[0241] Knowledge distillation is to transfer the knowledge of a complex teacher model to a simple student model. The teacher model is usually a well-trained model with good performance, while the student model is a model with relatively simple structure and high computational efficiency. By minimizing the difference between the output of the student model and the output of the teacher model, the student model learns the knowledge of the teacher model.

[0242] Through knowledge distillation, a lighter and more computationally efficient student model is obtained without significantly reducing model performance;

[0243] Application value of the module:

[0244] The model optimization module significantly improves the performance and computational efficiency of the multimodal neural network model through various technologies such as adaptive learning rate adjustment, model pruning and quantization, transfer learning and knowledge distillation. It enables the model to converge faster, reduces training time, and reduces the storage and computing requirements of the model, making it suitable for deployment on devices with limited resources. In addition, through transfer learning and knowledge distillation, it makes full use of existing data and model knowledge, improves the performance and generalization ability of the model on small data sets, and provides strong support for the practical application of elevator abnormal behavior detection systems.

[0245] In this embodiment, the interface between the system and the elevator control system includes:

[0246] Hardware-isolated CAN bus communication channel, transmission delay is less than 10 milliseconds;

[0247] The triggering condition of the safety relay module is that the first level alarm and the second level or third level alarm are activated at the same time;

[0248] The UPS life of the independent power supply unit shall be no less than twice the average single operation time of the elevator and no less than 300 seconds;

[0249] Furthermore, the interface between the system and the elevator control system plays a vital role in the elevator abnormal behavior detection system based on multimodal neural networks. It is the key bridge to realize data interaction and collaborative work between the detection system and the elevator control system. Through this interface, the detection system can timely feed back the abnormal behavior detection results to the elevator control system, so that the elevator control system can take corresponding measures to ensure the safe operation of the elevator. At the same time, the elevator control system can also provide the detection system with necessary operating parameters to help the detection system detect abnormal behavior more accurately. The following is a detailed description of the interface:

[0250] Overall function overview:

[0251] The main function of this interface is to realize the two-way data transmission and interaction between the detection system and the elevator control system. On the one hand, the detection system transmits the detected abnormal behavior information of the elevator and the related analysis results to the elevator control system; on the other hand, the elevator control system feeds back the real-time operation status, control instructions and other information of the elevator to the detection system. Through this two-way data interaction, the two systems can work closely together to improve the safety and reliability of elevator operation.

[0252] Data transmission protocol:

[0253] The interface uses CANopen protocol for data transmission; CANopen protocol is a high-level communication protocol based on CAN bus, which has the advantages of high reliability, strong real-time performance, and fast communication rate, and is very suitable for data communication in the field of industrial control; in this system, CANopen protocol can ensure accurate and timely transmission of data between the detection system and the elevator control system;

[0254] The CANopen protocol defines a series of communication objects, including service data objects (SDO) and process data objects (PDO). SDO is mainly used to transmit parameter configuration information and non-periodic data, such as abnormal behavior classification information and abnormal occurrence time sent by the detection system to the elevator control system. PDO is used to transmit periodic real-time data, such as the elevator running speed, position and other parameters sent by the elevator control system to the detection system.

[0255] Data interaction content:

[0256] Data transmitted from the detection system to the elevator control system:

[0257] Abnormal behavior information: When the detection system detects abnormal behavior in the elevator, it will send information such as the type and severity of the abnormal behavior to the elevator control system. For example, if violent behavior is detected in the elevator, the detection system will identify the abnormal behavior type as "violent behavior" and assess the severity based on the intensity of the behavior, and send it to the elevator control system in numerical form (such as 1-5, 5 being the most severe);

[0258] Abnormality time and location: The detection system will record the specific time when the abnormal behavior occurs and the location information of the elevator, and transmit this information to the elevator control system; the time information is accurate to the second, and the location information includes the floor where the elevator is located, the specific coordinates of the car, etc., so that the elevator control system can accurately grasp the abnormality;

[0259] Abnormal behavior analysis results: The results obtained by the detection system after analyzing abnormal behavior, such as the possible causes of abnormal behavior and the impact assessment on elevator operation safety, will also be sent to the elevator control system; these analysis results can help the elevator control system formulate more reasonable response strategies;

[0260] Data transmitted from the elevator control system to the detection system:

[0261] Elevator operation status parameters: including the elevator's operating speed, acceleration, current floor, door status (open or closed), etc. These parameters are important bases for the detection system to detect abnormal behavior. The detection system can use these parameters to determine whether the elevator is operating normally and whether there is a potential risk of abnormal behavior.

[0262] Control command information: The control command currently executed by the elevator control system, such as the direction of the elevator (up or down), whether to stop at a certain floor, etc. The detection system can combine these control command information to better understand the operation intention of the elevator, thereby more accurately detecting abnormal behavior;

[0263] Security and reliability design of the interface:

[0264] In order to ensure the security and reliability of interface data transmission, the following measures are taken:

[0265] Data Encryption:

[0266] During the data transmission process, the transmitted data is encrypted; a symmetric encryption algorithm, such as AES (Advanced Encryption Standard), is used to encrypt the data to prevent the data from being stolen or tampered with during the transmission process; the encryption key is pre-negotiated and determined by the detection system and the elevator control system, and is updated regularly to improve the security of encryption;

[0267] Error detection and correction:

[0268] On the basis of CANopen protocol, an error detection and correction mechanism is added; the cyclic redundancy check (CRC) technology is used to check the transmitted data; when sending data, the sender will calculate the CRC value of the data and attach it to the data and send it together; after receiving the data, the receiver will recalculate the CRC value of the data and compare it with the received CRC value; if the two are inconsistent, it means that an error has occurred in the data transmission process, and the receiver will ask the sender to resend the data;

[0269] Redundant design:

[0270] In order to improve the reliability of the interface, a redundant design is adopted; two independent CAN buses are set up for data transmission. When one bus fails, the other bus can continue to work to ensure the continuity of data transmission; at the same time, spare communication modules are set up in the detection system and the elevator control system respectively. When the main communication module fails, the spare communication module can automatically switch and continue to work;

[0271] Application value of the module:

[0272] The interface between the system and the elevator control system realizes close collaboration between the detection system and the elevator control system through efficient, safe and reliable data interaction. On the one hand, the detection system can promptly feedback abnormal behavior information to the elevator control system, so that the elevator control system can quickly take measures, such as emergency braking, notifying maintenance personnel, etc., to ensure the safety of people in the elevator. On the other hand, the operating parameters provided by the elevator control system to the detection system help the detection system to detect abnormal behavior more accurately and improve the accuracy and reliability of detection. This collaborative working mode can effectively reduce safety risks during elevator operation and improve the operating efficiency and management level of elevators.

[0273] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An elevator abnormal behavior detection system based on a multimodal neural network, characterized in that: include: A data acquisition module, including a visual sensor array, an inertial measurement unit (IMU), and an audio acquisition device installed in the elevator car; A data processing module is used to perform time-frequency domain conversion and feature vectorization processing on the original data; Multimodal neural network model, including parallel convolutional neural network (CNN) branches, sequential neural network (RNN) branches, and voiceprint recognition network; The feature fusion module adopts a dual fusion architecture that combines feature-level fusion and decision-level fusion based on the attention mechanism; Abnormal behavior classifier, a three-level abnormality discrimination model built based on a hierarchical classification strategy; The alarm and log module includes a dynamic threshold alarm unit and an abnormal behavior database.

2. The elevator abnormal behavior detection system based on multimodal neural network according to claim 1, characterized in that: The data acquisition module comprises: A combined visual sensor of an infrared thermal imaging camera and a visible light camera, with a spatial resolution of no less than 1920×1080@30fps; A 9-axis IMU device consisting of a three-axis accelerometer, a gyroscope, and a magnetometer, with a sampling frequency set to 100 Hz ± 5%; The directional microphone array is arranged at a preset angle θ with the elevator control panel to meet the sound field coverage requirement of 45°≤θ≤60°.

3. The elevator abnormal behavior detection system based on multimodal neural network according to claim 1 is characterized in that: The data processing module includes the following improved algorithms: The threshold function of the vibration signal denoising unit based on improved wavelet packet transform is defined as: When the absolute value of the signal is greater than or equal to the high-order threshold Th, the output value is the original signal minus the product of the threshold and the sign function; When the absolute value of the signal is less than the low-order threshold Tl, the output value is zero; When the absolute value of the signal is between Tl and Th, the output value is the ratio of the original signal to the S-type smoothing function, where the smoothing coefficient σ controls the attenuation rate of the transition zone, and the intermediate value Tm between Th and Tl is used to calculate the exponential term; The video feature extraction unit using the three-dimensional optical flow method has a motion vector calculation that satisfies the product of the partial derivative of the spatiotemporal pixel intensity in the x direction and the x velocity component, plus the product of the partial derivative in the y direction and the y velocity component, and the sum of the partial derivative in time equals zero; The audio feature fusion unit uses a hybrid feature of Mel frequency cepstral coefficients MFCC and improved linear predictive coding LPC, where the MFCC feature weight α ranges from 0.6 to 0.8, and the LPC feature weight is 1-α.

4. The elevator abnormal behavior detection system based on multimodal neural network according to claim 1, characterized in that: The multimodal neural network model comprises: The visual branch adopts an improved 3D-ResNet architecture, and the output of its spatiotemporal attention module is the sigmoid activation value of the convolution result of the spatial convolution kernel and the temporal feature map, and the sum of the convolution result of the temporal convolution kernel and the spatial feature map; The motion data processing branch uses a gated convolution-enhanced Bi-GRU network, whose gated update formula includes the linear transformation result of the previous hidden state and the current input, and the product of the one-dimensional convolution output and the learnable gating coefficient is superimposed and then activated by Sigmoid. The voiceprint recognition branch adopts an improved DCGAN model, whose generator loss function consists of a weighted adversarial loss term and an L2 regularization term output by the pre-trained generator.

5. The elevator abnormal behavior detection system based on multimodal neural network according to claim 1 is characterized in that: The feature fusion module comprises: Feature-level fusion uses an improved channel attention mechanism, whose weight is calculated as the transposed product of the query matrix and the key matrix adjusted by the dimension scaling factor, superimposed with a bias term and normalized by the Softmax function; The decision-level fusion adopts the credibility-corrected DS evidence theory, and the basic probability distribution function is defined as the normalized result of the weighted probability distribution value of each evidence source and the category confidence score after being corrected by the credibility adjustment factor; The final decision output adopts a gated fusion strategy, which concatenates the feature-level fusion result with the decision-level fusion result, transforms it through a trainable weight matrix, and outputs it through a Sigmoid activation function.

6. The elevator abnormal behavior detection system based on multimodal neural network according to claim 1, characterized in that: The three-level discrimination strategy of the abnormal behavior classifier includes: The first-level physical layer detection uses dynamic threshold comparison, and triggers an alarm when the acceleration modulus exceeds the mean plus three times the standard deviation within the sliding time window; The second-level behavior recognition uses an improved dynamic time warping (DTW) distance metric and introduces a time decay coefficient λ to perform exponential weighting on the distance matrix during path search. The third-level compound decision adopts a deep reinforcement learning strategy, which updates the Q function by combining the learning rate η and the discount factor γ through the difference between the maximum Q value output by the target network and the current Q value.

7. The elevator abnormal behavior detection system based on multimodal neural network according to claim 1, characterized in that: The alarm and log module includes: The dynamic threshold adjustment unit adopts a sliding window mechanism, and the window length T meets 30 seconds to 120 seconds; The anomaly database adopts a spatiotemporal coding storage format, and the record fields include timestamp, three-dimensional spatial coordinates, multimodal feature vector set and feature fingerprint hash value; The alarm interface supports RS485 and CAN bus dual protocol parallel transmission.

8. The elevator abnormal behavior detection system based on multimodal neural network according to claim 1, characterized in that: It also includes a model optimization module, which includes: The transfer learning method adopts a feature pyramid matching strategy and calculates the square sum of the Frobenius norm of each layer of the feature map of the source domain and the target domain as the transfer loss; The federated learning framework uses differential privacy gradient aggregation, adding Gaussian noise after client-side gradient summation; The feature dictionary update rule of the incremental learning unit is to retain the new features in the old dictionary whose distance from the class center exceeds the threshold τ, and select the top K features in order of importance.

9. The elevator abnormal behavior detection system based on multimodal neural network according to claim 1, characterized in that: The training method of the multimodal neural network model includes: Adversarial training sample generation uses gradient sign attack to superimpose gradient direction perturbations on the original samples; The multi-stage training process sets a joint loss function, which is a linear combination of classification loss, feature reconstruction loss and regularization term; The model validation adopts multi-dimensional evaluation indicators, which integrates the weighted scores of F1 value, AUC and ROC curve.

10. The elevator abnormal behavior detection system based on multimodal neural network according to claim 1, characterized in that: The interface between the system and the elevator control system includes: Hardware-isolated CAN bus communication channel, transmission delay is less than 10 milliseconds; The triggering condition of the safety relay module is that the first level alarm and the second level or third level alarm are activated at the same time; The UPS life of the independent power supply unit shall be no less than twice the average single operation time of the elevator and no less than 300 seconds.

Citation Information

Patent Citations

  • Multi-modal feature fusion emotion recognition method based on gating cross-attention mechanism

    CN117370828A

  • Property safety prevention and control method and system based on artificial intelligence

    CN118694899A

  • Elevator safety risk early warning method based on multi-modal data

    CN118877676A

  • Product visual image accurate identification and processing integrated platform

    CN118982543A

  • In-station anomaly analysis method and system based on multi-modal retrieval enhancement generation and readable medium

    CN119360119A

Cited By

  • Video monitoring early warning method and system based on multi-mode behavior mode

    CN120219903A

  • Drain valve action monitoring method and device, medium and program product

    CN120488096A

  • Power plant specific area personnel behavior identification and early warning system

    CN120632785A

  • Building equipment operation state monitoring system and method

    CN121050339A

  • Smart city elevator operation method, Internet of Things large model system and medium

    CN121063350A