An elevator video retrieval and recognition method and system based on an improved neural network

By combining an improved twin liquid hybrid neural network model with elevator physical constraints, the problems of low accuracy and weak anti-interference ability in elevator video retrieval and recognition are solved, realizing efficient and real-time elevator fault monitoring and early warning, and improving the safety and efficiency of elevator operation and maintenance.

CN122346565BActive Publication Date: 2026-08-25SICHUAN SPECIAL EQUIP INSPECTION & RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202610821507.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-09
Publication Date
2026-08-25
Estimated Expiration
2046-06-09

AI Technical Summary

Technical Problem

Existing video retrieval and recognition technologies have problems in elevator applications, such as low accuracy, weak anti-interference ability, difficulty in balancing accuracy and temporal correlation, and low retrieval efficiency of massive video data. They cannot meet the high-precision and high-efficiency requirements of elevator operation and maintenance, and are prone to delaying fault handling and causing safety hazards.

Method used

An improved twin liquid hybrid neural network model is adopted, combined with the physical constraints of the elevator. Through video acquisition, preprocessing, feature screening, feature extraction and similarity matching, a dynamic feature screening mechanism is constructed to eliminate environmental interference, improve the model's anti-interference ability, take into account the local details of the equipment and the temporal correlation, and build a dynamically updatable feature sample library to optimize the efficiency of feature extraction and retrieval.

Benefits of technology

It improves the accuracy and efficiency of elevator video retrieval and recognition, reduces the false judgment rate, realizes real-time monitoring and rapid early warning, meets the long-term dynamic needs of elevator operation and maintenance, and enhances the scalability and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122346565B_ABST
    Figure CN122346565B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of fault monitoring of feature equipment, and relates to an elevator video retrieval and identification method and system based on an improved neural network, and the technical points are as follows: real-time original video data is collected through a monitoring device, preprocessed, and video frame data is obtained; in combination with physical constraint conditions, feature screening is performed on the video frame data, the screened effective feature area is input into a twin liquid mixed neural network, feature extraction, similarity matching and time sequence optimization are completed, and the final retrieval and identification result is obtained. The application designs a device-specific physical constraint feature screening mechanism for the feature equipment such as the elevator, and eliminates interference features; a mixed neural network model is used to realize accurate extraction of device features, similarity matching and time sequence correlation capture; a complete video retrieval and identification system is constructed to form a complete process technical scheme from data collection to result application, and high-precision and high-efficiency retrieval and identification of elevator videos are realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of fault monitoring technology for feature devices, and specifically discloses an elevator video retrieval and recognition method and system based on an improved neural network. Background Technology

[0002] Elevators are critical core equipment in modern urban infrastructure and industrial production. Their safe and stable operation is crucial to the safety of people's lives and property, urban operations, and the continuity of industrial production. Elevators are widely used in various types of buildings and are the core carrier of vertical transportation. my country has over 9 million elevators, with an annual growth rate exceeding 10%, creating an urgent need for elevator fault diagnosis and condition monitoring. With the development of video surveillance and artificial intelligence, video-based retrieval and recognition technology has become a core means of elevator operation and maintenance monitoring. This technology analyzes surveillance video to locate target equipment, extract operational characteristics, identify abnormal states, and accurately retrieve target segments and key information, providing efficient data support for equipment operation and maintenance, reducing the workload of manual inspections, and improving operation and maintenance efficiency and monitoring accuracy.

[0003] However, existing video retrieval and recognition technologies have shortcomings in applications involving specific characteristic equipment such as elevators, making it difficult to meet the high-precision and high-efficiency requirements of operation and maintenance monitoring. Specific defects are as follows: First, they lack dynamic adaptability specifically for elevators, resulting in low retrieval and recognition accuracy; second, they do not incorporate constraint mechanisms designed based on the physical characteristics of elevators, leading to weak anti-interference capabilities; third, traditional neural network models have limitations, making it difficult to balance accuracy and temporal correlation; fourth, the retrieval efficiency of massive video data is low, making real-time monitoring impossible, easily delaying fault handling, and causing safety hazards.

[0004] Therefore, the present invention aims to provide an elevator video retrieval and recognition method and system based on an improved neural network to solve the above problems. Summary of the Invention

[0005] The purpose of this invention is to provide an elevator video retrieval and recognition method and system based on an improved neural network. By combining the physical constraints of the equipment and adopting a novel neural network model, a high-precision, high-efficiency, and highly anti-interference video retrieval and recognition technical solution is achieved, overcoming the shortcomings of the existing technology and providing strong support for the safe operation and maintenance of elevators. The specific plan is as follows: An elevator video retrieval and recognition method based on an improved neural network, the method comprising the following steps: S1. By deploying monitoring equipment in elevator scenarios, real-time raw video data is collected to build a video acquisition network; S2. Preprocess the acquired raw video data to obtain video frame data; S3. Combine physical constraints to perform feature filtering on video frame data, obtain effective feature regions, and normalize the effective feature regions. S4. Input the filtered effective feature regions into the twin liquid hybrid neural network to complete feature extraction, similarity matching and temporal optimization, and obtain the final retrieval and recognition results; S5. Organize and output the final retrieval and identification results, and feed the retrieval results back to the feature sample library in the system server to dynamically update the elevator fault status features in the feature sample library.

[0006] Furthermore, in step S2, video frame data is obtained by performing denoising, deblurring, frame extraction, and grayscale conversion on the original video data, as follows: S201. Noise Reduction Processing: Gaussian filtering algorithm is used to denoise the original video data. The filter kernel size is set to 3×3. The Gaussian filtering formula is as follows:

[0007] in, The standard deviation is Gaussian, with a value of 1.0. The coordinates of the filter kernel; S202 Deblurring: Adaptive histogram equalization algorithm is used to remove video frame blurring in the original video data; S203, Frame Extraction: Using an interval frame extraction strategy, one video frame is extracted from the elevator scene in the original video data at intervals of one frame, resulting in a standardized video frame sequence. The frame size is uniformly adjusted to 224×224 pixels. S204. Grayscale Conversion: Convert the original color video frames into grayscale frames. The grayscale conversion formula is as follows:

[0008] in, , , These are the pixel values ​​of the red, green, and blue channels of a color video frame, respectively. These are the pixel values ​​of the grayscale frame.

[0009] Furthermore, in step S3, the specific steps for filtering out effective feature regions based on video frame data are as follows: S301. Use an object detection algorithm to filter out candidate feature regions from video frame data; S302. The physical constraints include motion state constraints, size ratio constraints and state association constraints. The candidate feature regions are verified. Candidate feature regions that satisfy all three physical constraints are valid feature regions. Candidate feature regions that do not satisfy the physical constraints are eliminated. S303. The effective feature region is normalized to a uniform size of 224×224 pixels. The normalization formula is as follows:

[0010] in, These are the original pixel values. , These are the minimum and maximum pixel values, respectively. These are the normalized pixel values.

[0011] The motion state constraint verifies the motion trajectory of the candidate feature region using a trajectory offset calculation formula, which is as follows:

[0012] in, This represents the trajectory offset of the candidate feature region. For candidate feature regions in the th The center coordinates in a frame of video; For candidate feature regions in the th The center coordinates in a frame of video; The preset standard horizontal displacement is used when the car is constrained. When the gate body is constrained This represents the standard horizontal displacement for the opening and closing of the door. The preset standard vertical displacement is used when the car is constrained. For the standard vertical displacement of the car during lifting and lowering, when the door is constrained... ; Validation rules: When ,in A preset offset threshold of 5 pixels is used to determine whether the motion trajectory of the candidate feature region conforms to the constraint; otherwise, it is determined to be an interfering feature and is removed. The size ratio constraint is used to quantitatively verify the aspect ratio of the elevator in the candidate feature region through a size ratio verification formula, which is as follows:

[0013] in, The aspect ratio of the candidate feature region; The width of the candidate feature region; The height of the candidate feature region; Verification rules: Based on the elevator's intended use, a preset aspect ratio threshold range is established. ,when If the size ratio of the candidate feature region meets the constraints, it is determined that the feature region is a distractor and is therefore removed. The state association constraint uses a state association verification function to quantify the state association of candidate feature regions. The state association verification function is as follows:

[0014] in, The car's operating status (stationary / running); The state of the door (open / closed / open / closed); For the verification results, 1 indicates that the state association meets the constraints, and 0 indicates that it does not meet the constraints and is judged as an interference feature and is removed.

[0015] Furthermore, in S4, the twin-liquid hybrid neural network includes a twin network branch and a liquid neural network branch. The twin network branch performs feature extraction and similarity matching on the effective feature regions, and the liquid neural network branch captures the temporal correlation based on the feature extraction results, and finally outputs the retrieval and recognition results.

[0016] The twin network branch includes a convolutional neural network branch for the features to be retrieved and a convolutional neural network branch for the sample features; The normalized effective feature region is input into the convolutional neural network branch of the feature to be retrieved. Six convolutional layers and max pooling layers are used to extract local image detail features in the effective feature region. The extracted local image detail features are then mapped to a 128-dimensional feature vector in the fully connected layer. Elevator fault status features from the feature sample library are input into the sample feature convolutional neural network branch. Six convolutional layers and max pooling layers are used to extract local detail features of the image. The extracted local detail features are then mapped to a 128-dimensional feature vector in the fully connected layer. The cosine similarity of the feature vectors extracted by the convolutional neural network branch of the feature to be retrieved and the convolutional neural network branch of the sample features is compared. The formula for calculating cosine similarity is as follows:

[0017] in, The feature vector to be retrieved With sample feature vector The cosine similarity takes values ​​in the range [0,1]. The feature vector to be retrieved The One component; For sample feature vectors The One component; Judgment rule: When If the match is successful, the initial search result is obtained; otherwise, the match is considered unsuccessful and the search is marked as invalid. A contrastive loss function is used to optimize the Siamese network parameters. The formula for the contrastive loss function is as follows:

[0018] in: To compare the loss values; The number of training sample pairs; For sample pairs of labels, This indicates a positive sample pair, meaning that two samples belong to the same category; This represents a negative sample pair, meaning two samples belong to different categories; For the first Euclidean distance between the feature vectors of each sample pair; This is the marginal value, preset to 1.0.

[0019] The feature extraction formula for the convolutional layer is as follows:

[0020] in, The output features of the convolutional layer; For the first The weights of each convolutional kernel; For the first input feature map A local area; This refers to the bias term of the convolutional layer; It is the ReLU activation function. ; The number of convolution kernels; The feature extraction formula for the max pooling layer is as follows:

[0021] in, The output features of the pooling layer; The coordinates in the input feature map are Pixel values; This defines the region of the pooled window. The feature mapping formula for a fully connected layer is as follows:

[0022] in, This is the 128-dimensional feature vector output by the fully connected layer; is the weight matrix of the fully connected layer, with a dimension of 128×N, where N is the dimension of the output features of the pooling layer; The output features of the pooling layer; This is the bias vector of the fully connected layer, with a dimension of 128×1.

[0023] Furthermore, the liquid neural network branch includes an input layer, a dynamic hidden layer, and an output layer. The 128-dimensional feature vectors corresponding to the preliminary search results in the Siamese network branch are organized into a video frame feature vector sequence with 10 frames as a group. This sequence is then input into the input layer of the liquid neural network branch. Through the dynamic hidden layer, the weights of the neurons in the dynamic hidden layer of the video frame feature vector sequence are updated and the temporal loss function is optimized to capture the temporal correlation in the sequence, verify the temporal logic of the preliminary search results, eliminate mismatch results, and output the temporal logic verification results through the output layer. The formula for updating the weights of neurons in the dynamic hidden layer is as follows:

[0024] in, For the first Connection weights of neurons in real time; For the first Connection weights of neurons in real time; The learning rate is set to 0.0005 by default. The time-series loss function is used to optimize dynamic weights; Represents partial differentials; The formula for the time-series loss function is as follows:

[0025] in, The length of the time series; For the first The actual timing label of the frame, 1 indicates that the timing logic is reasonable, and 0 indicates that the timing logic is unreasonable; For the first The timing logic prediction result output by the frame Sigmoid activation function; The formula for the output layer Sigmoid activation function is as follows:

[0026] in, The result is the timing logic prediction result, with a value range of [0,1]. When, the timing logic prediction result is reasonable; when If the result is not found to be reasonable, the time-series logic prediction result will be discarded. This is the output value of the dynamic hidden layer.

[0027] Furthermore, in S5, the retrieval and identification results of the twin liquid hybrid neural network are visualized and output, and automatic anomaly warnings are given based on the fault status in the retrieval and identification results. At the same time, the feature sample library is dynamically updated based on the fault status in the retrieval and identification results.

[0028] The present invention also provides an elevator video retrieval and recognition system based on an improved neural network, the system comprising a video acquisition module, a preprocessing module, a physical constraint screening module, a hybrid neural network module, a feature sample library module, and an output feedback module; The video acquisition module collects real-time raw video data through monitoring equipment deployed in the elevator scene, and transmits it to the preprocessing module for data preprocessing in real time. The preprocessing module performs noise reduction, deblurring, frame extraction, and grayscale conversion on the acquired raw video data to obtain video frame data. The physical constraint filtering module combines physical constraint conditions to perform feature filtering on video frame data, obtain effective feature regions, and normalize the effective feature regions. The hybrid neural network module inputs the filtered effective feature regions into the twin liquid hybrid neural network to complete feature extraction, similarity matching and temporal optimization, and obtain the final retrieval and recognition results; The feature sample library module stores elevator fault status features, providing a data source for the hybrid neural network module to perform similarity comparison, and is dynamically updated based on the retrieval and recognition results of the hybrid neural network module; The output feedback module visualizes the retrieval and recognition results of the hybrid neural network module and provides automatic anomaly warnings based on the retrieval and recognition results.

[0029] Compared with existing technologies, the beneficial effects of this solution are: 1. This invention establishes specific physical constraints on elevators, constructs a feature filtering mechanism for dynamic characteristics, eliminates environmental interference features, improves the model's anti-interference ability, reduces the false positive rate of retrieval, and increases the recall rate. 2. This invention employs a hybrid neural network model that combines Siamese networks and liquid neural networks, taking into account both the extraction of local detailed features of the device and temporal correlations, thereby overcoming the limitations of traditional single neural network models and improving the accuracy of retrieval and recognition. 3. This invention constructs a dedicated physical constraint condition to optimize feature extraction, integrates a neural network model, improves the retrieval efficiency of massive video data, reduces retrieval response time, and meets the actual needs of elevator real-time monitoring, rapid early warning, and fault tracing.

[0030] 4. By constructing a dynamically updatable feature sample library, this invention effectively improves the scalability and long-term adaptability of the model, thereby better meeting the long-term and dynamic needs of elevator operation and maintenance monitoring. Attached Figure Description

[0031] Figure 1 This is a flowchart of the elevator video retrieval and recognition method based on an improved neural network according to the present invention.

[0032] Figure 2 This is a structural diagram of the elevator video retrieval and recognition system based on an improved neural network according to the present invention.

[0033] Figure 3 This is a technical roadmap for the elevator video retrieval and recognition system based on an improved neural network, as described in this invention. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0035] like Figures 1-3 As shown, this invention designs a device-specific physical constraint feature screening mechanism for characteristic devices such as elevators to eliminate interfering features; it adopts a hybrid neural network model combining Siamese networks and liquid neural networks to achieve accurate extraction of device features, similarity matching, and temporal correlation capture; and it constructs a complete video retrieval and recognition system, covering modules such as video acquisition, preprocessing, device type determination, physical constraint screening, feature extraction and matching, and result output and feedback, forming a complete technical solution from data acquisition to result application, realizing high-precision and high-efficiency retrieval and recognition of elevator videos.

[0036] The specific plan is as follows: An elevator video retrieval and recognition method based on an improved neural network, the method comprising the following steps: S1. By deploying monitoring equipment in elevator scenarios, real-time raw video data is collected to build a video acquisition network; S2. Preprocess the acquired raw video data to obtain video frame data; S3. Combine physical constraints to perform feature filtering on video frame data, obtain effective feature regions, and normalize the effective feature regions. S4. Input the filtered effective feature regions into the twin liquid hybrid neural network to complete feature extraction, similarity matching and temporal optimization, and obtain the final retrieval and recognition results; S5. Organize and output the final retrieval and identification results, and feed the retrieval results back to the feature sample library in the system server to dynamically update the elevator fault status features in the feature sample library.

[0037] Furthermore, a high-definition infrared camera deployed on the top of the elevator car captures video of the car's interior and the operation of the car doors; a high-definition infrared camera deployed above the elevator hall door captures video of the elevator doors opening and closing and the car stopping; the camera supports low-light environment shooting (infrared night vision function), adapting to scenes where the light inside the elevator changes frequently, with the frame rate set to 25 frames / second and the video resolution to 1920×1080, ensuring that the dynamic characteristics of the elevator can be clearly captured.

[0038] The collected video data is transmitted in real time via wired network (fiber optic) or wireless network (5G) to ensure the stability and real-time performance of data transmission and avoid data loss or delay.

[0039] Furthermore, in S2, video frame data is obtained by performing denoising, deblurring, frame extraction, and grayscale conversion on the original video data, as follows: S201. Denoising Processing: A Gaussian filtering algorithm is used to denoise the original video data. Gaussian filtering can effectively eliminate Gaussian noise in the video (such as noise caused by changes in lighting and noise from the camera itself) while preserving effective features in the video frames. The filter kernel size is set to 3×3 to balance denoising effectiveness and feature preservation. The Gaussian filtering formula is as follows:

[0040] in, The standard deviation is Gaussian, with a value of 1.0. The coordinates of the filter kernel; S202 Deblurring: Adaptive histogram equalization algorithm is used to solve the problem of video frame blurring caused by changes in light inside the elevator and dust obstruction in the pipeline scene, enhance the contrast and clarity of video frames, and ensure that equipment features can be clearly identified. S203, Frame Extraction: Considering the huge amount of video data, in order to improve the efficiency of subsequent processing, an interval frame extraction strategy is adopted. In the elevator scene, one video frame is extracted every 1 frame interval, and in the pipeline scene, one video frame is extracted every 2 frames interval, so as to obtain a standardized video frame sequence. The frame size is uniformly adjusted to 224×224 pixels. S204. Grayscale Conversion: Converting color video frames to grayscale frames reduces data dimensionality and the computational load for subsequent feature extraction, while preserving the core form and feature information of the device. The grayscale conversion formula is as follows:

[0041] in, , , These are the pixel values ​​of the red, green, and blue channels of a color video frame, respectively. These are the pixel values ​​of the grayscale frame.

[0042] Furthermore, S3, the specific steps for filtering out effective feature regions based on video frame data are as follows: S301. The YOLOv8 target detection algorithm is used to filter out candidate feature regions (such as elevator cars, doors, surrounding interference objects, etc.) from video frame data, and information such as the coordinates, size, and motion trajectory of the candidate regions are obtained. S302. Physical constraints include motion state constraints, size ratio constraints, and state association constraints. Candidate feature regions are verified as follows: the motion trajectory is verified to meet the vertical (car) or horizontal (door) constraints; the size ratio is verified to meet the standard range of elevator car and door; and the motion state and door state are verified to meet the association constraints. Candidate feature regions that meet all three physical constraints are valid feature regions. Candidate feature regions that do not meet the physical constraints are eliminated. S303. The effective feature regions are normalized to a uniform size of 224×224 pixels to facilitate subsequent feature extraction by the neural network model. The normalization formula is as follows:

[0043] in, These are the original pixel values. , These are the minimum and maximum pixel values, respectively. These are the normalized pixel values.

[0044] The core operating components of an elevator (car and doors) have fixed motion trajectories. The car moves only in the vertical direction (up and down), and its motion trajectory is consistent with the elevator shaft, with no horizontal deviation. The doors open and close only in the horizontal direction (left and right), and their motion trajectory is parallel to the front surface of the elevator car, with no vertical displacement.

[0045] Motion state constraints are used to verify the motion trajectory of candidate feature regions through a trajectory offset calculation formula, which is as follows:

[0046] in, This represents the trajectory offset of the candidate feature region. For candidate feature regions in the th The center coordinates in a frame of video; For candidate feature regions in the th The center coordinates in a frame of video; The preset standard horizontal displacement is used when the car is constrained. (No horizontal displacement), when the gate is constrained This represents the standard horizontal displacement for the opening and closing of the door. The preset standard vertical displacement is used when the car is constrained. For the standard vertical displacement of the car during lifting and lowering, when the door is constrained... (No vertical displacement); Validation rules: When ,in A preset offset threshold of 5 pixels is used to determine whether the motion trajectory of the candidate feature region conforms to the constraint; otherwise, it is determined to be an interfering feature and is removed. Elevator cars and doors have industry-standard dimensional proportions. Although the dimensional proportions of different types of elevators (residential elevators and commercial elevators) vary, they all fall within a fixed range. Specifically, the length-to-width ratio of residential elevator cars ranges from 3:1 to 4:1, and the length-to-width ratio of doors ranges from 1:2 to 1:3; while the length-to-width ratio of commercial elevator cars ranges from 2.5:1 to 3.5:1, and the length-to-width ratio of doors ranges from 1:1.8 to 1:2.8.

[0047] The aspect ratio constraint is used to quantitatively verify the aspect ratio of the elevator in the candidate feature region through a aspect ratio verification formula, which is as follows:

[0048] in, The aspect ratio of the candidate feature region; The width of the candidate feature region; The height of the candidate feature region; Verification rules: Based on the elevator's usage type (residential / commercial), preset aspect ratio threshold ranges are established. ,when If the size ratio of the candidate feature region meets the constraints, it is determined that the feature region is a distractor and is therefore removed. There is a fixed correlation between the elevator's operating state and the door's state. This correlation has a clear physical logic and is irreversible. The specific correlation rules are as follows: when the elevator car is in a stopped state (stationary and level with the floor), the door must be in an open state; when the elevator car is in a running state (vertical movement), the door must be in a closed state; when the door is in the process of opening or closing, the car must be in a stationary state (stopped state).

[0049] State association constraints use a state association verification function to quantify the state association of candidate feature regions. The state association verification function is as follows:

[0050] in, The car's operating status (stationary / running); The state of the door (open / closed / open / closed); For the verification results, 1 indicates that the state association meets the constraints, and 0 indicates that it does not meet the constraints and is judged as an interference feature and is removed.

[0051] Furthermore, in S4, the Siamese-Liquid Hybrid Neural Network comprises a Siamese network branch and a liquid neural network branch. The Siamese network branch extracts features and performs similarity matching on effective feature regions, while the liquid neural network branch captures the temporal correlation of the feature extraction results, ultimately outputting the retrieval and recognition results. During the feature extraction and matching process, the model's loss value is calculated in real time (the Siamese network uses a contrastive loss function, and the liquid neural network uses a temporal loss function). When the loss value is higher than 0.01, the weights of the convolutional layers of the Siamese network and the neuron connection weights of the liquid neural network are dynamically adjusted to ensure that the model is always in an optimal state, thereby improving the retrieval and recognition accuracy.

[0052] The Siamese network branch consists of two structurally symmetric, parameter-sharing convolutional neural network (CNN) branches: the CNN branch for the features to be retrieved and the CNN branch for the sample features; The normalized effective feature regions are input into the convolutional neural network branch for the features to be retrieved. Six convolutional layers are used to extract local detail features of the image. The kernel size is mainly 3×3, focusing on enhancing the extraction of dynamic features (such as motion trajectory and speed changes). Max pooling is used to reduce the feature dimensionality and retain key features. Fully connected layers are used to map the features extracted by the convolutional and pooling layers into 128-dimensional feature vectors for subsequent similarity calculation. Elevator fault status features from the feature sample library are input into the sample feature convolutional neural network branch. Six convolutional layers and max pooling layers are used to extract local detail features of the image. The extracted local detail features are then mapped to a 128-dimensional feature vector in the fully connected layer. The cosine similarity of the feature vectors extracted by the convolutional neural network branch of the feature to be retrieved and the convolutional neural network branch of the sample features is compared. The formula for calculating cosine similarity is as follows:

[0053] in, The feature vector to be retrieved With sample feature vector The cosine similarity takes values ​​in the range [0,1]. The feature vector to be retrieved The One component; For sample feature vectors The One component; Judgment rule: When When the initial match is successful, the preliminary search results are obtained (including target location, status, and corresponding video frame). If the match fails, the search is marked as invalid. A contrastive loss function is used to optimize the Siamese network parameters. The formula for the contrastive loss function is as follows:

[0054] in: To compare the loss values; The number of training sample pairs; For sample pairs of labels, This indicates a positive sample pair, meaning that two samples belong to the same category; This represents a negative sample pair, meaning two samples belong to different categories; For the first Euclidean distance between the feature vectors of each sample pair; This is a marginal value used to control the loss contribution of negative sample pairs, and is preset to 1.0.

[0055] By sharing parameters, the feature extraction standards of the two branches are kept consistent, improving the accuracy of feature matching. Through customized convolutional layer design, the feature differences of elevators are adapted, and the extraction of dynamic and static features is strengthened respectively, solving the problem of poor adaptability of general models. The cosine similarity calculation method is simple and efficient, which can quickly achieve feature matching and improve retrieval efficiency.

[0056] The feature extraction formula for convolutional layers is as follows:

[0057] in, The output features of the convolutional layer; For the first The weights of each convolutional kernel; For the first input feature map A local area; This refers to the bias term of the convolutional layer; It is the ReLU activation function. ; The number of convolution kernels; The feature extraction formula for the max pooling layer is as follows:

[0058] in, The output features of the pooling layer; The coordinates in the input feature map are Pixel values; This defines the region of the pooled window. The feature mapping formula for a fully connected layer is as follows:

[0059] in, This is the 128-dimensional feature vector output by the fully connected layer; is the weight matrix of the fully connected layer, with a dimension of 128×N, where N is the dimension of the output features of the pooling layer; The output features of the pooling layer; This is the bias vector of the fully connected layer, with a dimension of 128×1.

[0060] Furthermore, the liquid neural network branch includes an input layer, a dynamic hidden layer, and an output layer. It organizes the 128-dimensional feature vectors corresponding to the initial retrieval results from the Siamese network branch into a video frame feature vector sequence, grouped into sets of 10 frames, and inputs this sequence into the input layer of the liquid neural network branch. The dynamic hidden layer contains three dynamic neuron layers, where the connection weights of each neuron can adaptively adjust according to the temporal changes in the input sequence. This eliminates the need for manually pre-setting a fixed topology and allows for rapid capture of temporal patterns within the sequence. Through the dynamic hidden layer, the weights of the neurons in the dynamic hidden layer are updated, and the temporal loss function is optimized to capture temporal correlations within the sequence. This verifies the temporal logic of the initial retrieval results, eliminates mismatches, and outputs the temporal logic verification results through the output layer. For example, if the initial retrieval results show "the elevator car moves while the door opens," the liquid neural network will determine that the temporal logic of this result is incorrect. ), and remove it.

[0061] The formula for updating the weights of neurons in the dynamic hidden layer is as follows:

[0062] in, For the first Connection weights of neurons in real time; For the first Connection weights of neurons in real time; The learning rate is set to 0.0005 by default. The time-series loss function is used to optimize dynamic weights; Represents partial differentials; The formula for the time-series loss function is as follows:

[0063] in, The length of the time series; For the first The actual timing label of the frame, 1 indicates that the timing logic is reasonable, and 0 indicates that the timing logic is unreasonable; For the first The timing logic prediction result output by the frame Sigmoid activation function; The formula for the output layer Sigmoid activation function is as follows:

[0064] in, The result is the timing logic prediction result, with a value range of [0,1]. When, the timing logic prediction result is reasonable; when If the result is not found to be reasonable, the time-series logic prediction result will be discarded. This is the output value of the dynamic hidden layer.

[0065] The dynamic topology can adaptively adapt to the temporal characteristics of elevators, eliminating the need for a separate temporal model designed for elevators; it has a fast convergence speed, improving convergence efficiency by more than 30% compared to traditional RNN models, enabling rapid processing of massive video sequence data; and it has a strong ability to capture temporal correlations, effectively eliminating mismatched results with inconsistent temporal logic, thus improving the accuracy of retrieval and recognition.

[0066] Hybrid model collaborative workflow: After preprocessing and physical constraint filtering, the video to be retrieved obtains effective feature regions; the effective feature regions are input into a Siamese network, which extracts 128-dimensional feature vectors through convolutional layers, pooling layers, and fully connected layers, and matches them with the sample feature library using the cosine similarity formula to obtain preliminary retrieval results; the feature vector sequence corresponding to the preliminary retrieval results is input into a liquid neural network, which performs temporal logic verification through dynamic weight updates and temporal loss function optimization, eliminates false matching results, and obtains the final retrieval and recognition results.

[0067] Furthermore, in S5, the retrieval and recognition results of the twin liquid hybrid neural network are displayed in a combination of charts, text, and video clips. Specifically, these include: target location coordinates (accurate to meters), equipment operating status (elevator: stopped / operating / faulty), target appearance time (accurate to seconds), and corresponding video clips (extracting key frames and continuous segments). At the same time, the results can be exported (Excel and PDF formats) for easy viewing and archiving by maintenance personnel.

[0068] If abnormal equipment status is detected in the search and identification results (such as elevator door jamming or abnormal car stopping), an abnormality warning will be automatically triggered, and maintenance personnel will be notified in a timely manner through sound, pop-up window, SMS and other means to avoid safety hazards.

[0069] New features discovered during the retrieval and identification process (such as elevator malfunction features not included in the database) are automatically fed back to the feature sample library to supplement new sample data. The samples in the sample library are then retrained to optimize model parameters and improve the accuracy and adaptability of subsequent retrieval and identification.

[0070] The present invention also provides an elevator video retrieval and recognition system based on an improved neural network. The system includes a video acquisition module, a preprocessing module, a physical constraint screening module, a hybrid neural network module, a feature sample library module, and an output feedback module. The video acquisition module collects real-time raw video data through monitoring equipment deployed in the elevator scene and transmits it to the preprocessing module for data preprocessing in real time. The preprocessing module performs noise reduction, deblurring, frame extraction, and grayscale conversion on the acquired raw video data to obtain video frame data. The physical constraint filtering module combines physical constraints to perform feature filtering on video frame data, removing interfering features and retaining effective feature regions; the filtered effective feature regions are then transmitted to the hybrid neural network module. The hybrid neural network module inputs the filtered effective feature regions into the twin liquid hybrid neural network to complete feature extraction, similarity matching and temporal optimization, and obtain the final retrieval and recognition results; The feature sample library module stores elevator fault status features, providing a data source for the hybrid neural network module to perform similarity comparison, and is dynamically updated based on the retrieval and recognition results of the hybrid neural network module; The output feedback module visualizes the retrieval and recognition results of the hybrid neural network module and provides automatic anomaly warnings based on the retrieval and recognition results.

[0071] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for elevator video retrieval and recognition based on an improved neural network, characterized in that, The method includes the following steps: S1. By deploying monitoring equipment in elevator scenarios, real-time raw video data is collected to build a video acquisition network; S2. Preprocess the acquired raw video data to obtain video frame data; S3. Combine physical constraints to perform feature filtering on video frame data, obtain effective feature regions, and normalize the effective feature regions. S4. Input the filtered effective feature regions into the twin liquid hybrid neural network to complete feature extraction, similarity matching and temporal optimization, and obtain the final retrieval and recognition results; S5. Organize and output the final retrieval and identification results, and feed the retrieval results back to the feature sample library in the system server to dynamically update the elevator fault status features in the feature sample library. The specific steps for S3, which involves filtering out effective feature regions based on video frame data, are as follows: S301. Use an object detection algorithm to filter out candidate feature regions from video frame data; S302. The physical constraints include motion state constraints, size ratio constraints and state association constraints. The candidate feature regions are verified. Candidate feature regions that satisfy all three physical constraints are valid feature regions. Candidate feature regions that do not satisfy the physical constraints are eliminated. S303. The effective feature region is normalized to a uniform size of 224×224 pixels. The normalization formula is as follows: in, These are the original pixel values. , These are the minimum and maximum pixel values, respectively. These are the normalized pixel values; In S4, the twin liquid hybrid neural network includes a twin network branch and a liquid neural network branch. The twin network branch performs feature extraction and similarity matching on the effective feature regions, and the liquid neural network branch captures the temporal correlation based on the feature extraction results, and finally outputs the retrieval and recognition results. The twin network branch includes a convolutional neural network branch for the features to be retrieved and a convolutional neural network branch for the sample features; The normalized effective feature region is input into the convolutional neural network branch of the feature to be retrieved. Six convolutional layers and max pooling layers are used to extract local image detail features in the effective feature region. The extracted local image detail features are then mapped to a 128-dimensional feature vector in the fully connected layer. Elevator fault status features from the feature sample library are input into the sample feature convolutional neural network branch. Six convolutional layers and max pooling layers are used to extract local detail features of the image. The extracted local detail features are then mapped to a 128-dimensional feature vector in the fully connected layer. The cosine similarity of the feature vectors extracted by the convolutional neural network branch of the feature to be retrieved and the convolutional neural network branch of the sample features is compared. The formula for calculating cosine similarity is as follows: in, The feature vector to be retrieved With sample feature vector The cosine similarity takes values ​​in the range [0,1]. The feature vector to be retrieved The One component; For sample feature vectors The One component; Judgment rule: When If the match is successful, the initial search result is obtained; otherwise, the match is considered unsuccessful and the search is marked as invalid. A contrastive loss function is used to optimize the Siamese network parameters. The formula for the contrastive loss function is as follows: in: To compare the loss values; The number of training sample pairs; For sample pairs of labels, This indicates a positive sample pair, meaning that two samples belong to the same category; This represents a negative sample pair, meaning two samples belong to different categories; For the first Euclidean distance between the feature vectors of each sample pair; This is the marginal value, preset to 1.0; The feature extraction formula for the convolutional layer is as follows: in, The output features of the convolutional layer; For the first The weights of each convolutional kernel; For the first input feature map A local area; This refers to the bias term of the convolutional layer; It is the ReLU activation function. ; The number of convolution kernels; The feature extraction formula for the max pooling layer is as follows: in, The output features of the pooling layer; The coordinates in the input feature map are Pixel values; This defines the region of the pooled window. The feature mapping formula for a fully connected layer is as follows: in, This is the 128-dimensional feature vector output by the fully connected layer; is the weight matrix of the fully connected layer, with a dimension of 128×N, where N is the dimension of the output features of the pooling layer; The output features of the pooling layer; This is the bias vector of the fully connected layer, with a dimension of 128×1; The liquid neural network branch includes an input layer, a dynamic hidden layer, and an output layer. The 128-dimensional feature vectors corresponding to the preliminary search results in the Siamese network branch are organized into a video frame feature vector sequence with 10 frames as a group. This sequence is then input into the input layer of the liquid neural network branch. Through the dynamic hidden layer, the weights of the neurons in the dynamic hidden layer of the video frame feature vector sequence are updated and the temporal loss function is optimized to capture the temporal correlation in the sequence, verify the temporal logic of the preliminary search results, eliminate mismatch results, and output the temporal logic verification results through the output layer. The formula for updating the weights of neurons in the dynamic hidden layer is as follows: in, For the first Connection weights of neurons in real time; For the first Connection weights of neurons in real time; The learning rate is set to 0.0005 by default. The time-series loss function is used to optimize dynamic weights; Represents partial differentials; The formula for the time-series loss function is as follows: in, The length of the time series; For the first The actual timing label of the frame, 1 indicates that the timing logic is reasonable, and 0 indicates that the timing logic is unreasonable; For the first The timing logic prediction result output by the frame Sigmoid activation function; The formula for the output layer Sigmoid activation function is as follows: in, The result is the timing logic prediction result, with a value range of [0,1]. When, the timing logic prediction result is reasonable; when If the result is not found to be reasonable, the time-series logic prediction result will be discarded. This is the output value of the dynamic hidden layer.

2. The elevator video retrieval and recognition method based on an improved neural network according to claim 1, characterized in that, The motion state constraint verifies the motion trajectory of the candidate feature region using a trajectory offset calculation formula, which is as follows: in, This represents the trajectory offset of the candidate feature region. For candidate feature regions in the th The center coordinates in a frame of video; For candidate feature regions in the th The center coordinates in a frame of video; The preset standard horizontal displacement is used when the car is constrained. When the gate body is constrained This represents the standard horizontal displacement for the opening and closing of the door. The preset standard vertical displacement is used when the car is constrained. For the standard vertical displacement of the car during lifting and lowering, when the door is constrained... ; Validation rules: When ,in A preset offset threshold of 5 pixels is used to determine whether the motion trajectory of the candidate feature region conforms to the constraint; otherwise, it is determined to be an interfering feature and is removed. The size ratio constraint is used to quantitatively verify the aspect ratio of the elevator in the candidate feature region through a size ratio verification formula, which is as follows: in, The aspect ratio of the candidate feature region; The width of the candidate feature region; The height of the candidate feature region; Verification rules: Based on the elevator's intended use, a preset aspect ratio threshold range is established. ,when If the size ratio of the candidate feature region meets the constraints, it is determined that the feature region is a distractor and is therefore removed. The state association constraint uses a state association verification function to quantify the state association of candidate feature regions. The state association verification function is as follows: in, This refers to the operating status of the car; The state of the gate; For the verification results, 1 indicates that the state association meets the constraints, and 0 indicates that it does not meet the constraints and is judged as an interference feature and is removed.

3. The elevator video retrieval and recognition method based on an improved neural network according to claim 1, characterized in that, In step S5, the retrieval and identification results of the twin liquid hybrid neural network are visualized and output, and automatic anomaly warning is given based on the fault status in the retrieval and identification results. At the same time, the feature sample library is dynamically updated based on the fault status in the retrieval and identification results.

4. The elevator video retrieval and recognition method based on an improved neural network according to claim 1, characterized in that, In step S2, video frame data is obtained by performing noise reduction, deblurring, frame extraction, and grayscale conversion on the original video data, as follows: S201. Noise Reduction Processing: Gaussian filtering algorithm is used to denoise the original video data. The filter kernel size is set to 3×3. The Gaussian filtering formula is as follows: in, The standard deviation is Gaussian, with a value of 1.

0. The coordinates of the filter kernel; S202 Deblurring: Adaptive histogram equalization algorithm is used to remove video frame blurring in the original video data; S203, Frame Extraction: Using an interval frame extraction strategy, one video frame is extracted from the elevator scene in the original video data at intervals of one frame, resulting in a standardized video frame sequence. The frame size is uniformly adjusted to 224×224 pixels. S204. Grayscale Conversion: Convert the original color video frames into grayscale frames. The grayscale conversion formula is as follows: in, , , These are the pixel values ​​of the red, green, and blue channels of a color video frame, respectively. These are the pixel values ​​of the grayscale frame.

5. An elevator video retrieval and recognition system based on an improved neural network, characterized in that, The system can execute an elevator video retrieval and recognition method based on an improved neural network as described in any one of claims 1-4; The system includes a video acquisition module, a preprocessing module, a physical constraint screening module, a hybrid neural network module, a feature sample library module, and an output feedback module. The video acquisition module collects real-time raw video data through monitoring equipment deployed in the elevator scene, and transmits it to the preprocessing module for data preprocessing in real time. The preprocessing module performs noise reduction, deblurring, frame extraction, and grayscale conversion on the acquired raw video data to obtain video frame data. The physical constraint filtering module combines physical constraint conditions to perform feature filtering on video frame data, obtain effective feature regions, and normalize the effective feature regions. The hybrid neural network module inputs the filtered effective feature regions into the twin liquid hybrid neural network to complete feature extraction, similarity matching and temporal optimization, and obtain the final retrieval and recognition results; The feature sample library module stores elevator fault status features, providing a data source for the hybrid neural network module to perform similarity comparison, and is dynamically updated based on the retrieval and recognition results of the hybrid neural network module; The output feedback module visualizes the retrieval and recognition results of the hybrid neural network module and provides automatic anomaly warnings based on the retrieval and recognition results.

Citation Information

Patent Citations

  • Non-intrusive elevator monitoring method, device and system

    CN113233278A

  • Construction elevator standard operation monitoring and warning system and method based on video analysis

    CN114529868A