An elevator indicator light state determination method, electronic device, and medium
By using a Bayesian model that combines video and sensor data to identify the status of elevator indicator lights, the problem of low efficiency and waste of manpower and resources in traditional elevator door lock short-circuit detection is solved, realizing automated and real-time monitoring of elevator door lock short-circuit detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-20
- Publication Date
- 2026-03-31
AI Technical Summary
Traditional elevator door lock short circuit detection methods rely on manual inspection, which is inefficient, prone to omissions, and cannot achieve real-time monitoring. In addition, they require a lot of manpower and resources, affecting the judgment of results and increasing equipment wear.
By acquiring the changes in the indicator lights caused by the elevator door lock switch, and combining video data, physical sensor data, and environmental sensor data, the elevator indicator light status is identified using a preset model. A Bayesian model is then used to fuse multiple sensor data to determine whether the door lock is short-circuited.
It has achieved automation and real-time monitoring of elevator door lock short-circuit detection, reduced the risk of misjudgment, adapted to the diverse operating environment of elevators and the dynamic changes of indicator lights, and reduced the input of manpower and material resources.
Smart Images

Figure CN121542898B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of special equipment technology, and more specifically, to a method for determining the status of elevator indicator lights, an electronic device, and a medium. Background Technology
[0002] Elevators, as an indispensable vertical transportation tool in modern buildings, are directly related to the safety of public life and property. In daily elevator use, the door lock system, as a crucial safety protection device, not only plays a vital role in verifying the complete closure of the elevator landing and car doors, but also serves as a crucial line of defense against fatal accidents such as doors opening and the elevator moving unnecessarily. It is worth noting that intentionally short-circuiting the door lock circuit will directly lead to the doors opening, potentially causing a shearing accident between the car and the landing door. Therefore, regular safety inspections of the door lock system are essential. Elevator door locks include main door locks and auxiliary door locks. The traditional method for testing whether elevator landing door locks are short-circuited is as follows: Two certified personnel are on the car top. One person operates the inspection run button to start the elevator for inspection. When the elevator reaches the door lock position, the other person disconnects the main door lock (at this time, it must be ensured that the inspection run button is in the running state). If the elevator stops, it is determined that the main door lock is not short-circuited; otherwise, the main door lock is short-circuited. Then, the main door lock is manually short-circuited, and the auxiliary door lock is disconnected during the inspection run (at this time, it must be ensured that the inspection run button is in the running state). If the elevator stops, it is determined that the auxiliary door lock is not short-circuited; otherwise, the auxiliary door lock is short-circuited. This process is repeated for each floor.
[0003] Traditional detection methods have the following drawbacks: First, if maintenance personnel release the maintenance operation while the door lock is open, it can also cause the elevator to stop. In this case, it is impossible to determine whether the stop was caused by the door lock opening or the maintenance operation being interrupted, affecting the judgment. Second, it requires checking the elevator's operation each time, resulting in frequent starting and stopping, which is time-consuming and increases equipment wear. Third, at least two people are required to operate the elevator on the car top, resulting in a significant manpower investment. Therefore, it is clear that traditional detection methods are a considerable waste of both manpower and resources.
[0004] Therefore, this application proposes a method to determine whether the door lock is short-circuited by obtaining the changes in the indicator light caused by the elevator door lock switch and combining the opening and closing of the door lock. The key is how to identify the state of the elevator indicator light. Summary of the Invention
[0005] In view of this, the purpose of this application is to provide a method for determining the status of elevator indicator lights, which can improve the problems of traditional detection methods that rely on manual inspection or simple electrical testing, resulting in low efficiency, easy omissions, and inability to achieve real-time monitoring.
[0006] To achieve the above technical objectives, the technical solution adopted in this application is as follows:
[0007] In a first aspect, embodiments of this application provide a method for determining the status of an elevator indicator light, the method comprising:
[0008] Acquire current video data, N physical sensor data, and N environmental sensor data, wherein the current video data includes N video images, and the timestamp of each physical sensor data and environmental sensor data corresponds to the timestamp of the video image, and N is a natural number greater than or equal to 2;
[0009] The current video data is input into a preset model, and the first state of the elevator indicator light in each video image, the corresponding first confidence level, and the detection box of the elevator indicator light are obtained through the preset model.
[0010] Based on all the physical sensor data, determine the second state of the elevator indicator light and the corresponding second confidence level;
[0011] Based on all the physical sensor data, physical features are obtained; based on all the detection boxes, visual features are obtained; and based on all environmental sensor data, environmental features are obtained.
[0012] Based on the physical features and video features, and using a first algorithm, the feature matching degree between the physical features and video features is determined. The first algorithm is used to take the physical features and video features as input and output the feature matching degree.
[0013] Based on the first state, the second state, the first confidence level, the second confidence level, the feature matching degree, and the environmental features, the attention weights of the preset Bayesian model are determined using the second algorithm, and the attention weights are used as the model parameters of the Bayesian model. The second algorithm is used to fuse the first state, the second state, the first confidence level, the second confidence level, the feature matching degree, and the environmental features to determine the attention weights of the Bayesian model.
[0014] The visual features, physical features, environmental features, first state, first confidence level, second state, and second confidence level are all input into the Bayesian model, and the final state of the elevator indicator light in each frame of the video image is obtained through the Bayesian model.
[0015] Furthermore, after obtaining visual features based on all the detection boxes, the method further includes:
[0016] Based on a preset template, the visual features are verified to obtain a verification result. The preset template includes a preset visual feature range obtained from several historical video data. In the several historical video data, the occlusion rate of the area where the elevator indicator light is located is less than a first preset value and the inter-frame brightness variance is less than a second preset value.
[0017] When the verification result is used to characterize that the video features do not match the preset visual feature range, the first state and the first confidence level are corrected based on the video features.
[0018] Furthermore, determining the feature matching degree between the physical features and video features based on the first algorithm includes:
[0019] Obtain the physical features and the video features that have the same timestamp;
[0020] The physical features and the video features with the same timestamp are used as input values and fed into the first algorithm. The first algorithm is used to obtain the output value, which is used to characterize the feature matching degree.
[0021] Furthermore, the step of determining the preset attention weights of the Bayesian model based on the first state, the second state, the first confidence level, the second confidence level, the feature matching degree, and environmental features, using the second algorithm, includes:
[0022] A first state consistency score is determined based on the total number of frames in the current video data and the number of frames in the video image for the second state; and a second state consistency score is determined based on all the physical sensor data and the number of physical sensor data representing the first state.
[0023] Based on the aforementioned environmental characteristics, visual and physical basis weights are determined.
[0024] Based on the visual basis weights, physical basis weights, first confidence level, second confidence level, and feature confidence level, physical attention weights and visual attention weights are determined so that the physical attention weights and visual attention weights serve as the attention weights of the Bayesian model.
[0025] Furthermore, the step of inputting the visual features, physical features, environmental features, first state, first confidence level, second state, and second confidence level into the Bayesian model, and obtaining the final state of the elevator indicator light through the Bayesian model, includes:
[0026] The visual features, physical features, and environmental features are mapped to the same space to obtain visual feature codes, physical feature codes, and environmental feature codes with the same dimensions.
[0027] The visual feature encoding, physical feature encoding, environmental feature encoding, first state, first confidence level, second state, and second confidence level are all input into the Bayesian model to obtain several third states of the elevator indicator light and the third confidence level corresponding to the third state.
[0028] When the highest third confidence level is less than the first confidence level threshold, or the difference between the highest and second-highest third confidence levels is less than a preset difference, the third state is corrected based on a preset strategy to obtain the final state.
[0029] Furthermore, the step of correcting the third state based on a preset strategy to obtain the final state includes:
[0030] Based on the aforementioned environmental characteristics, the visual interference coefficient and the physical interference coefficient were calculated.
[0031] When there is a unique high-reliability state, the high-reliability state is set as the final state; when there are two or more high-reliability states, the first high-reliability state that is the same as the first state or the second state with a feature matching degree greater than the first matching degree threshold is set as the final state. The first high-reliability state is the third state corresponding to the third confidence degree that is greater than the second confidence degree threshold and whose visual interference coefficient and physical interference coefficient are both greater than the interference coefficient threshold.
[0032] When the first high-reliability state does not exist, the second high-reliability state is set as the final state. The second high-reliability state represents the first state or the second state where the matching degree threshold is greater than the second matching degree threshold. Specifically, when the corresponding visual basic weight is greater than the physical basic weight, the second high-reliability state is the first state, and when the corresponding visual basic weight is less than the physical basic weight, the second high-reliability state is the second state.
[0033] When there is no second highly reliable state, let the third reliable state be the final state, and the third reliable state represents the state with the largest number of states.
[0034] Furthermore, the preset model includes a backbone network and a neck network;
[0035] The backbone network includes a first GhostConv module group, a second GhostConv module group, and a third GhostConv module group. The first GhostConv module group includes several first GhostConv modules, the second GhostConv module group includes several second GhostConv modules, and the third GhostConv module group includes several third GhostConv modules. The first GhostConv module group is used to extract basic visual features from the video image to obtain a basic visual feature map. The second GhostConv module group is used to generate a ghost feature map based on the basic visual feature map. The third GhostConv module group is used to concatenate the ghost feature map and the basic visual feature map to obtain an output feature map.
[0036] The neck network includes a REC3 module and is equipped with a SIMAM attention mechanism.
[0037] Secondly, embodiments of this application provide an electronic device, which includes a processor and a memory coupled to each other. The memory stores a computer program, and when the computer program is executed by the processor, the electronic device performs the above-described method.
[0038] Thirdly, embodiments of this application provide a computer-readable storage medium, characterized in that the computer-readable storage medium stores a computer program, which, when run on a computer, causes the computer to perform the above-described method.
[0039] The invention employing the above technical solution has the following advantages:
[0040] The technical solution provided in this application outputs the first state, first confidence level, and detection box of the elevator indicator light in each frame through a preset model, and simultaneously determines the second state and second confidence level based on physical sensor data. Then, physical features, visual features, and environmental features are extracted from the three types of sensor data and the detection box, respectively. A first algorithm is used to calculate the cross-modal matching degree between physical and visual features, and a second algorithm is used to fuse the first / second states, confidence levels, feature matching degrees, and environmental features to dynamically determine the attention weights of the Bayesian model and use them as model parameters. Finally, the three types of features, the bimodal state, and the confidence level are input into the Bayesian model, and the final state of the elevator indicator light in each frame of the video image is output through probabilistic inference. This solution reduces the risk of single-modal misjudgment and adapts to the diverse environments of elevator operation and the dynamic changes of the indicator lights. Attached Figure Description
[0041] This application can be further illustrated by the non-limiting embodiments given in the accompanying drawings. It should be understood that the following drawings only illustrate some embodiments of this application and should not be considered as limiting the scope. For those skilled in the art, other related drawings can be obtained from these drawings without any inventive effort.
[0042] Figure 1 A flowchart provided for an embodiment of this application.
[0043] Figure 2 This is a sub-flowchart of S150 provided in an embodiment of this application.
[0044] Figure 3 This is a sub-flowchart of S160 provided in an embodiment of this application.
[0045] Figure 4 This is a sub-flowchart of S170 provided in an embodiment of this application. Detailed Implementation
[0046] The present application will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that similar or identical parts are referred to by the same reference numerals in the drawings or description. Implementations not shown or described in the drawings are forms known to those skilled in the art. In the description of this application, terms such as "first" and "second" are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0047] This application provides an electronic device that may include a processing module and a storage module. The storage module stores a computer program, which, when executed by the processing module, enables the electronic device to perform the corresponding steps in the elevator indicator light status determination method described below. The electronic device in this embodiment is not limited to desktop computers, laptops, tablets, etc.
[0048] Please refer to Figure 1 This application also provides a method for determining the status of an elevator indicator light. The method for determining the status of an elevator indicator light may include the following steps:
[0049] S110, acquire current video data, N physical sensor data and N environmental sensor data, wherein the current video data includes N video images, and the timestamp of each physical sensor data and environmental sensor data corresponds to the timestamp of the video image, and N is a natural number greater than or equal to 2;
[0050] S120, The current video data is input into a preset model, and the first state of the elevator indicator light in each video image, the corresponding first confidence level, and the detection box of the elevator indicator light are obtained through the preset model;
[0051] S130, based on all the physical sensor data, determine the second state of the elevator indicator light and the corresponding second confidence level;
[0052] S140, obtain physical features based on all the physical sensor data, obtain visual features based on all the detection frames, and obtain environmental features based on all the environmental sensor data;
[0053] S150, Based on the physical features and video features, and using the first algorithm, determine the feature matching degree between the physical features and video features;
[0054] S160, based on the first state, the second state, the first confidence level, the second confidence level, the feature matching degree, and the environmental features, the attention weight of the preset Bayesian model is determined according to the second algorithm, and the attention weight is used as the model parameter of the Bayesian model.
[0055] S170, the visual features, physical features, environmental features, first state, first confidence level, second state and second confidence level are all input into the Bayesian model, and the final state of the elevator indicator light in each frame of the video image is obtained through the Bayesian model.
[0056] In the above implementation, the final state of the elevator indicator light is determined based on the comparison of visual and physical features, the results of mutual verification, and the influence of environmental features on visual and physical features.
[0057] The steps for determining the status of elevator indicator lights will be explained in detail below:
[0058] In the S110, video data is captured by a dedicated embedded high-definition smart camera for elevators. This camera is installed on the top inner wall of the elevator control cabinet or in a reserved mounting position next to the main board, with the lens facing the indicator light area on the main board. The camera can simultaneously acquire frame data and perform timestamp calibration with the door lock circuit status detection terminal inside the control cabinet to ensure the continuity and validity of N frames of video images, and to be completely synchronized with the timestamps of the physical sensor data, thus meeting the multimodal data timing alignment requirements of the solution.
[0059] In this embodiment, N is 5, meaning that 5 frames of video images are acquired in each cycle, and each cycle can be 5 seconds.
[0060] In this embodiment, the physical sensor data is based on measurements taken by physical sensors, including vibration sensors. The vibration sensors are synchronized with the timestamps of the video data. The vibration sensors are mounted on an electromagnetic relay. When the main door lock is closed, the corresponding relay coil is energized, and the sensor collects a composite vibration of "closing impact + contact contact." When the main door lock is open, it collects a vibration of "release rebound + contact separation." If there is a short circuit fault in the door lock, the relay will remain in the closed state, and the sensor will not send a subsequent release vibration signal. By extracting features from the timing and amplitude of the vibration signal, the working state of the relay can be accurately deduced, thereby indirectly determining the on / off state and short circuit status of the door lock circuit, achieving an effective mapping of vibration signals to the door lock state.
[0061] In this embodiment, the environmental sensor data is based on measurements taken by environmental sensors, including a light sensor and a mechanical noise sensor. The light sensor is installed on the side of the lens of an embedded high-definition smart camera to collect the light intensity of the shooting area that accurately covers the indicator lights on the control cabinet motherboard. It collects data at a sampling frequency of 10Hz, and N=5 frames of data are collected synchronously every 5 seconds. The timestamps of the collected data are strictly aligned with those of the video images. After collection, a moving average is calculated first, and then light abrupt changes deviating from the mean by ±3σ (such as strong light during machine room maintenance or momentary occlusion) are removed. Finally, the effective average light data is obtained, which is used to calculate the visual interference coefficient. The mechanical noise sensor is installed on the mounting plate at the bottom of the elevator control cabinet to collect the decibels of background mechanical noise transmitted from the machine room equipment to the control cabinet. The timestamps of the collected data are strictly aligned with those of the vibration sensor data. N=5 frames of data are collected synchronously every 5 seconds. After collection, a moving average is calculated, and based on an 80dB sensor anti-interference threshold, the effective average noise data is obtained, which is used to calculate the physical interference coefficient.
[0062] In this embodiment, the timestamps of environmental sensor data, physical sensor data, and visual sensor data are first used to construct a globally unified clock via NTP (Network Time Protocol) or IEEE 1588PTP Precision Time Protocol, synchronizing the local clocks of the camera, physical sensor, and environmental sensor to the same UTC time base with error controlled within milliseconds. Secondly, a hardware trigger signal (such as a PPS second pulse) synchronously triggers the acquisition actions of the three types of devices, ensuring consistent data acquisition start times. Finally, at the software level, a sliding window and linear interpolation algorithm are used to compensate for data with slight time deviations, while abnormal timestamp samples are removed, ultimately achieving precise matching between each frame of video image and the corresponding physical sensor data and environmental sensor data. In S120, the preset model is the YOLO8 model, which consists of a backbone network, a neck network, and a detection head. The preset model includes a backbone network and a neck network. The backbone network includes a first GhostConv module group, a second GhostConv module group, and a third GhostConv module group. The first GhostConv module group includes several first GhostConv modules, the second GhostConv module group includes several second GhostConv modules, and the third GhostConv module group includes several third GhostConv modules. The first GhostConv module group is used to extract the basic visual features of the video image to obtain a basic visual feature map. The second GhostConv module group is used to generate a ghost feature map based on the basic visual feature map through a cheap linear transformation, and to mine feature redundancy to reduce computation. The third GhostConv module group is used to concatenate the ghost feature map and the basic visual feature map to obtain an output feature map with complete dimensions and strong expressive power, which is compressed by the SPPF module and then fed into the neck network. The neck network includes a RepC3 module, which deploys a SimAM attention mechanism and employs a PANet structure to achieve bidirectional fusion of multi-scale features. The RepC3 module enhances feature extraction capabilities through a reparameterized design with multiple branches during the training phase and a single branch during the inference phase. The SimAM attention mechanism has no additional parameters and highlights key features of the indicator light region through an energy function while suppressing background interference. The detection head adopts a decoupled design, divided into a classification branch and a bounding box regression branch, which output the indicator light category (status), classification confidence (first confidence), and detection box coordinates, respectively.
[0063] After the video image is input into the preset model, it is first preprocessed (scaled to 640×640 and normalized) and then fed into the backbone network: The first GhostConv module group performs preliminary convolution operations on the input image to extract basic visual features such as brightness and contour, generating a basic visual feature map; the second GhostConv module group applies a lightweight linear transformation to the basic visual feature map to generate redundant ghost feature maps, expanding the feature dimension with low computational cost; the third GhostConv module group concatenates the basic visual feature map and the ghost feature map along the channel axis to obtain a fused output feature map, which is then fed into the neck network after multi-scale information is aggregated by the SPPF module. In the neck network, the RepC3 module performs deep extraction and cross-stage fusion of the features output by the backbone network to enhance fine-grained texture features. The SimAM attention mechanism allocates spatial weights by calculating the neuron energy value to highlight the features of the indicator light area and suppress background noise. Then, the PANet structure realizes multi-scale feature interaction from top to bottom and bottom to top to generate a fused feature map adapted to indicator lights of different sizes. Finally, the feature map is input to the detection head, the classification branch outputs the state category of the indicator light (first state, such as shorted / non-shorted) and the corresponding classification confidence (first confidence) through the fully connected layer, and the regression branch outputs the bounding box coordinates (detection box) of the indicator light, thus completing the extraction of indicator light related information in a single frame video image.
[0064] The preset model in this embodiment can be trained based on the PyTorch framework. The specific process is as follows: First, an elevator door lock indicator dataset is constructed, collecting thousands of images covering both short-circuit and non-short-circuit states, including different lighting and background interference scenes. The LabelImg tool is used to annotate the detection boxes and class labels, and the dataset is divided into training, validation, and test sets in an 8:1:1 ratio. The training parameters are set as follows: epochs=150, batchsize=8, initial learning rate lr0=0.001, weight decay 0.005, using the SGD optimizer (momentum 0.937) and a warmup strategy (warm-up for the first 3 epochs). The loss function is CIoU loss (boundary box regression) + cross-entropy loss (classification). During the training phase, the GhostConv module of the backbone network generates features according to the process of "primary feature extraction - ghost feature expansion - feature fusion". The RepC3 module of the neck network is trained with a multi-branch structure, and the SimAM attention mechanism dynamically adjusts the feature response weights. The model parameters are iteratively optimized through backpropagation, and the accuracy changes are monitored using the validation set. If there is no improvement in accuracy when patience=50, training is stopped. After training, the multi-branch convolutions of the RepC3 module are equivalently fused to generate a single-branch inference structure. Finally, a lightweight model file is exported to ensure high efficiency during deployment. At the same time, ablation experiments are used to verify the effectiveness of the combination of GhostConv, SimAM, and RepC3 modules and determine the optimal model structure.
[0065] S130, as the physical mode state determination step, aims to extract the anti-interference second state and its corresponding second confidence level from N frames of time-series vibration sensor data based on the inherent linkage law of elevator door lock relays: "energized and engaged → generating characteristic vibration → indicator lights energized and illuminated synchronously". The specific implementation process is as follows:
[0066] Step 1: Data Time Series Preprocessing. Perform statistical aggregation processing on N frames of vibration sensor data (including vibration trigger time and vibration amplitude data), and calculate the mean ΔT of the difference between the vibration trigger time and the elevator door lock action time, the moving average of the vibration amplitude Amp, and the variance of the vibration amplitude fluctuation. This process eliminates noise in single-frame data (such as mechanical vibration in the computer room background and vibration noise caused by electromagnetic interference), ensuring the stability of input features.
[0067] Step 2: N-frame state voting to determine the second state. For each frame of preprocessed vibration data, output the supported state for each frame according to preset rules: If a composite vibration of "closing impact + contact contact" is detected → support the "door lock closed / indicator light on" state; if a vibration of "release rebound + contact separation" is detected → support the "door lock open / indicator light off" state; if only continuous closing vibration is detected and no release rebound vibration is collected → support the "door lock short-circuited / indicator light always on" state; count the frequency of each supported state in N frames, and determine the state with the highest frequency as the second state (Example scenario: when N=5 frames, 3 frames support the "door lock short-circuited" state, then the second state is determined as "door lock short-circuited", and the corresponding elevator indicator light state is "always on"). If there are ambiguous scenarios with multiple states having the same frequency, the state with a higher vibration trigger time difference and higher vibration amplitude consistency is given priority.
[0068] Step 3: Second Confidence Measurement Calculation. A three-dimensional weighted fusion model is used to calculate the confidence level, with the formula: conf = 0.4 × S1 + 0.3 × match + 0.3 × S3. Where: S1 is the physical feature's own reliability score, obtained by weighted summation of vibration trigger time difference consistency (weight 0.5) and vibration amplitude stability (weight 0.5); match is the vibration-visual feature matching degree, output from step S140; S3 is the environmental interference resistance score, calculated based on the physical interference coefficient. The final result is normalized and mapped to the 0-1 interval.
[0069] As an exemplary implementation, N=5 frames are set (vibration sensor data is aligned frame-by-frame with video image timestamps), and the physical sensor used is only a vibration sensor (sampling frequency 500Hz). Preset parameters are as follows: vibration amplitude stability variance threshold. =0.02g² (based on sensor accuracy calibration), physical interference coefficient Kvib=0.08, vibration-visual feature matching degree match=0.92, standard time difference ΔT0=0.3s between door lock action and vibration trigger (calibrated during initial elevator calibration).
[0070] N frames of vibration sensor data (after timing preprocessing):
[0071] The average difference between vibration trigger time and door lock action time: ΔT = 0.1s (small deviation from standard time difference, high consistency);
[0072] Vibration amplitude sliding average: Amp = 0.22g;
[0073] Variance of vibration amplitude fluctuation: =0.005g² (stable, less than the threshold of 0.02g²).
[0074] Continuous engagement vibration was detected in each frame, but no release rebound vibration was collected → "Door lock short-circuited / indicator light always on" status was supported for all 5 frames.
[0075] Status voting: All 5 frames support "door lock shorted / indicator light always on", so the second state = elevator indicator light always on (corresponding to door lock shorted).
[0076] Physical reliability score:
[0077] S1=0.5×(1- / )+0.5×(1-ΔT / ΔT0)≈0.5×(1-0.005 / 0.02)+0.5×(1-0.1 / 0.3)≈0.5×0.75+0.5×0.667≈0.375+0.333≈0.708;
[0078] Environmental interference resistance score:
[0079] S3 = 1 × (1 - Kvib) = 1 - 0.08 = 0.92;
[0080] The second confidence level conf = 0.4 × 0.708 + 0.3 × 0.92 + 0.3 × 0.92 ≈ 0.283 + 0.276 + 0.276 ≈ 0.835.
[0081] The final output of this step is: Elevator indicator light second state = always on (corresponding to door lock short circuit), second confidence level conf = 0.835.
[0082] The multimodal feature definition (physical / visual / environment) in S140 can be as follows:
[0083] In this embodiment, the physical features only include the vibration feature group. All physical features are calculated based on data collected by vibration sensors installed on the door lock relay. Specifically:
[0084] Vibration trigger time difference mean ΔT (unit: s): The average difference between the relay action vibration trigger time collected by the vibration sensor in N frames and the elevator door lock action command time, used to characterize the timing synchronization between the relay action and the door lock command;
[0085] Vibration amplitude moving average Amp (unit: g): The moving average of the vibration amplitude of the relay engaging / releasing collected by the vibration sensor for N frames, used to characterize the mechanical strength stability of the relay operation;
[0086] Vibration amplitude fluctuation variance (Unit: g²): The variance of the vibration amplitude of N frames is used to characterize the noise interference level of the vibration signal and to help determine the validity of the vibration data.
[0087] The visual features are extracted from the indicator light area of the control cabinet motherboard based on images captured by an embedded high-definition smart camera. The visual features may include the following:
[0088] (1) Appearance and color characteristics:
[0089] Average brightness L (grayscale value, 0-255): The average pixel value of the grayscale image of the indicator light's ROI area, used to characterize the luminous intensity of the indicator light;
[0090] Saturation mean S (HSV space, 0-255): The pixel average value of the S channel of the HSV color space in the ROI area of the indicator light, used to distinguish the color purity of the indicator light and avoid interference from the metal reflection of the control cabinet.
[0091] Mean brightness V (HSV color space, 0-255): The average pixel value of the V channel in the HSV color space of the indicator light ROI area, which helps to verify the actual luminous state of the indicator light.
[0092] (2) Dynamic time series characteristics:
[0093] Indicator flashing frequency F (unit: Hz): The main peak frequency extracted by fast Fourier transform of the brightness data of the indicator light ROI area of N frames, used to characterize the flashing pattern of the indicator light as the door lock circuit is turned on and off.
[0094] (3) Detection box morphological features:
[0095] The aspect ratio of the detection frame (dimensionless): the ratio of the width to the height of the indicator light ROI detection frame, used to verify whether the detection frame accurately defines the indicator light area and avoid false detection of other electrical components in the control cabinet;
[0096] The ratio_area (dimensionless) of the detection frame area is the ratio of the indicator light ROI area to the total area of the video image. It is used to characterize the proportion of the indicator light in the shooting field of view and to help determine the effectiveness of visual acquisition.
[0097] Environmental characteristics are used to quantify the degree of interference from the computer room control cabinet environment on multimodal acquisition, and are used to dynamically adjust the weights of each mode. These characteristics may include the following:
[0098] (1) Visual interference coefficient Kvis (dimensionless, value 0-1): It is only related to the ambient light intensity of the camera's field of view. It quantifies the degree of interference of light anomalies such as strong direct light during computer room maintenance, weak light during night inspection, and light and shadow disturbance of airflow in the control cabinet on visual feature acquisition. The higher the value, the stronger the light interference. It is the core parameter for subsequent visual modality weight adjustment.
[0099] (2) Physical interference coefficient Kvib (dimensionless, value 0-1): It is only for the acquisition interference of the door lock relay vibration sensor. It quantifies the influence of electromagnetic interference in the control cabinet and background mechanical vibration transmitted by the machine room traction machine on the acquisition of vibration signal. The higher the value, the stronger the vibration interference. It is specifically used for dynamic calibration of vibration mode weight.
[0100] For example, the method for obtaining multimodal features is as follows:
[0101] 1. Methods for obtaining physical characteristics.
[0102] Step 101: Data Synchronization Acquisition. The vibration sensor (sampling frequency 500Hz) installed on the door lock relay housing collects N frames of vibration data. Through the timestamp calibration module built into the control cabinet, the data is aligned frame by frame with the video images captured by the embedded high-definition smart camera, with a time synchronization error ≤10ms.
[0103] Step 102: Time-series aggregation processing. Perform 3-frame sliding window aggregation processing on the raw vibration data: calculate the mean ΔT of the difference between the vibration trigger time (the moment the relay engages / disengages) and the door lock action command time, the moving average of the vibration amplitude Amp, and the variance of the vibration amplitude fluctuation. Eliminate single-frame noise caused by mechanical vibration and electromagnetic interference in the computer room background, and improve feature stability;
[0104] Step 103: Feature Quantization and Normalization. Map the aggregated parameters to the 0-1 interval as follows:
[0105] 1. The moving average value of vibration amplitude, Amp, is normalized according to the 0-1g range;
[0106] 2. Variance of vibration amplitude fluctuation Divide by the preset vibration amplitude stability variance threshold Normalize;
[0107] 3. The vibration trigger time difference ΔT is normalized according to the reasonable response range of 0 - 0.5 s of the elevator equipment to obtain a standardized physical feature.
[0108] 2. Visual feature acquisition process
[0109] Step 201: Accurately extract the ROI region. Based on the detection box coordinates (x1, y1, x2, y2) output by S120, frame by frame, crop the target indicator light region (ROI) on the control cabinet main board, and exclude background interference such as wiring terminals and other small indicator lights in the control cabinet through the aspect ratio verification of the detection box;
[0110] Step 202: Extract single-frame features. Convert the ROI region to the HSV color space, and extract the pixel means of the three channels of L (mean brightness), S (mean saturation), and V (mean value); perform a fast Fourier transform on the brightness data of N frames of the ROI region, and take the frequency with the largest amplitude as the indicator light flashing frequency F; synchronously calculate the aspect ratio ratio_wh of the detection box and the area ratio ratio_area of the detection box;
[0111] Step 203: Temporal robust processing. Calculate the average values of L, S, V, and F of N frames, and at the same time calculate the variance of the detection box size (characterizing the positioning stability of the detection box), filter out the feature anomalies caused by single-frame illumination mutations, and finally output the robust visual features.
[0112] Environmental feature acquisition process.
[0113] (1) Obtain the visual interference coefficient Kvis
[0114] Step 301: Collect illumination data. The illumination sensor samples N frames of illumination intensity (unit: lux) at 10 Hz, which is strictly aligned with the video image timestamp;
[0115] Step 302: Filter out illumination anomalies. Calculate the moving average value L of N frames of illumination,剔除 the mutation data deviating from the mean value by ±3σ, and recalculate the mean value;
[0116] Step 303: Quantify the interference coefficient. Execute the following rules based on the illumination mean value:
[0117] ① Strong light interference: L ≥ 5000 lux → K1 = L / 10000 (value upper limit 1.0);
[0118] ② Weak light interference: L ≤ 100 lux → K1 = (100 - L) / 100;
[0119] ③ Normal illumination: 100 < L < 5000 lux → K1 = 0.1 (low interference reference value).
[0120] (2) Obtaining the physical interference coefficient from Kvib
[0121] Step 401: Interference Data Acquisition. A mechanical noise sensor (20-2000Hz band) acquires N frames of noise decibel values (unit: dB), aligning the timestamps of the mechanical noise sensor data with those of the physical sensor.
[0122] Step 402: Data smoothing. Calculate the moving average value N of the mechanical noise.
[0123] Step 403: Interference Coefficient Calculation. Based on the sensor's anti-interference threshold calibration results, perform the following quantization:
[0124] Physical disturbance coefficient Kvib (vibration): Critical threshold 80dB → K=min(N / 100,1.0).
[0125] For example, multimodal features can be obtained in the following way:
[0126] 1. Experimental setup and prerequisites
[0127] (1) Hardware configuration: Vibration sensor (range 0-5g, accuracy ±0.001g);
[0128] (2) Environmental sensors: light sensor (range 0-10000 lux), mechanical noise sensor (20-2000 Hz);
[0129] (3) Preset threshold: normal illumination range 100-5000 lux, vibration amplitude stability variance threshold =0.02g², all based on sensor performance and data center environment characteristics; the acquisition rule is N=5 frames / group, acquisition cycle 5S, and the synchronization error of timestamps for each modality data ≤10ms.
[0130] 2. Feature acquisition process and result output
[0131] (1) Physical feature acquisition and quantification
[0132] Exemplary raw data (N=5 frames):
[0133] Vibration time difference: 0.1s, 0.08s, 0.12s, 0.09s, 0.11s;
[0134] Vibration amplitude: 0.2g, 0.22g, 0.19g, 0.21g, 0.20g;
[0135] Timing processing results:
[0136] ΔT (mean vibration time difference) = (0.1 + 0.08 + 0.12 + 0.09 + 0.11) / 5 = 0.1s;
[0137] Amp (mean amplitude of vibration) = (0.2 + 0.22 + 0.19 + 0.21 + 0.20) / 5 = 0.204g;
[0138] (Variance of vibration amplitude fluctuation) = Σ[(Amp_i-Amp)²] / 5 = 0.00012g²;
[0139] The standardized physical characteristics include the following three dimensions (all normalized and mapped to the 0-1 range):
[0140] The mean vibration time difference ΔT: after normalization to a reasonable range of 0-0.5s for the equipment response, the value is 0.2. The closer the value is to 1, the higher the timing synchronization between the relay action and the door lock command.
[0141] The average vibration amplitude Amp is 0.204 after normalization to the 0-1g range. The closer the value is to 1, the more stable the mechanical strength of the relay operation.
[0142] Vibration amplitude fluctuation variance After normalization by a preset threshold (0.02g²), the value is 0.006. The closer the value is to 0, the less noise interference the vibration signal has and the higher its stability.
[0143] (2) Visual feature acquisition and quantification
[0144] Exemplary raw data (N=5 frames):
[0145] The coordinates of the detection box are all [120, 230, 170, 280] (x1, y1, x2, y2).
[0146] ROI region L values: 200, 205, 210, 208, 202;
[0147] ROI region S values: 70, 72, 75, 73, 71;
[0148] ROI region V values: 180, 185, 190, 188, 182;
[0149] Timing processing results:
[0150] L=205, S=72.2, V=185;
[0151] Fourier transform yields F=0.0Hz (indicator light does not flicker);
[0152] The detection frame ratio = (170-120) / (280-230) = 1.0, and the ratio = (50×50) / (1920×1080) ≈ 0.02;
[0153] Standardized visual features: [L=205,S=72.2,V=185,F=0.0,ratio=1.0,ratio=0.02].
[0154] Feature parameter description:
[0155] ① Brightness L: The average pixel value of the grayscale image of the ROI region, reflecting the luminous intensity of the indicator light;
[0156] ②Saturation S: The mean value of the S channel in the HSV space, representing the color purity of the indicator light;
[0157] ③ Brightness V: Mean value of V channel in HSV space, used to help determine the luminous state;
[0158] ④Flicker frequency F: The main frequency of brightness timing changes, used to distinguish between normal illumination and faulty flickering.
[0159] Saturation S: Convert the detection box to HSV color format (easier to extract color features), and take the average pixel value of the "S channel" (between 0 and 255, such as 72.1).
[0160] Brightness V: Also in HSV format, take the average pixel value of the "V channel" (between 0 and 255, such as 183.5).
[0161] Flicker frequency F: Take a small area of the indicator light for 5 consecutive frames, calculate the change in brightness for each frame, and use a simple algorithm (Fourier transform) to find the "regular frequency" of the brightness change (between 0-5Hz, 0.0 for no flicker).
[0162] Feature output format: The above visual features, the first state and the first confidence level output by S120 are encapsulated into structured data (such as JSON format) for subsequent matching degree calculation with physical features.
[0163] (3) Acquisition and quantification of environmental characteristics
[0164] Exemplary raw data (N=5 frames):
[0165] Illumination intensity: 2000 lux, 2100 lux, 1900 lux, 2200 lux, 2050 lux;
[0166] Mechanical noise: 60dB, 62dB, 58dB, 65dB, 61dB;
[0167] Raw data: 5 frames of illumination (2000 lux, 2100 lux, 1900 lux, 2200 lux, 2050 lux), mechanical noise (60dB, 62dB, 58dB, 65dB, 61dB);
[0168] Quantification results and stability score:
[0169] ① Average illumination L = 2050 lux (normal illumination) → K = 0.1;
[0170] ② Noise mean N = 61.2 dB → K = 61.2 / 100 = 0.612;
[0171] ③ Environmental stability score E = 1 - (interference coefficient time series variance / 0.01) = 0.85 (the smaller the variance, the higher the stability, which means that the fluctuation of environmental interference in the computer room is smaller and the feature acquisition stability is higher).
[0172] Standardized environmental characteristics: [K=0.1, K=0.612, E=0.85].
[0173] The standardized environmental features output in this embodiment will serve as the core input data for subsequent feature matching degree calculation and Bayesian model attention weight determination, thereby achieving effective fusion of multimodal information.
[0174] In this embodiment, the first state and its first confidence level need to be verified based on a preset template to determine whether correction is needed. Specifically:
[0175] Based on a preset template, the visual features are verified to obtain a verification result. The preset template includes a preset visual feature range obtained from several historical video data. In the several historical video data, the occlusion rate of the area where the elevator indicator light is located is less than a first preset value and the inter-frame brightness variance is less than a second preset value. When the verification result is used to characterize that the video features do not match the preset visual feature range, the first state and the first confidence level are corrected based on the video features.
[0176] In this embodiment, several historical video data of elevator indicator lights are collected, and valid data with an occlusion rate less than a first preset value (e.g., 10%) and an inter-frame brightness variance less than a second preset value (e.g., 10) are selected. Visual features (brightness, saturation, flicker frequency, etc.) of these data are extracted, and the normal range of each feature is statistically determined to form a preset visual feature range, thus obtaining a preset template. After acquiring the visual features of the current video image, they are compared one by one with the feature range of the preset template: if the brightness of the visual feature is within the template brightness range, the saturation / brightness is within the template color gamut, and the flicker frequency is within the template frequency range, the verification result is "match"; if any feature exceeds the range, the verification result is "mismatch". When the verification result is "mismatch", the first state and the first confidence score are corrected based on visual features: for example, if the model outputs "light on" but the visual brightness is lower than the template lower limit, the first state is corrected to "light off"; at the same time, the confidence score is adjusted according to the formula "corrected first confidence score = original first confidence score × template matching degree" (the matching degree is 0-1, and it is reduced according to the degree of deviation when there is a mismatch), and the matching degree = number of features that match the template / total number of verification features.
[0177] In this embodiment, visual feature distortion caused by lighting interference and slight occlusion is filtered out by using a preset template based on historical valid data, thus preventing the model from outputting incorrect states due to abnormal output of a single frame image. Correction of states and confidence levels for mismatched scenes improves the reliability of visual modal output, providing more accurate "visual evidence" for subsequent cross-modal fusion, reducing the impact of misjudgments by a single visual model on the overall solution, and enhancing the anti-interference capability and robustness of visual recognition.
[0178] In S150, it is necessary to provide state verification for the visual judgment results of the preset model based on physical signals, and at the same time, measure the degree of fit with visual features through feature matching degree, providing a basis for cross-modal fusion. Therefore, as Figure 2 As shown, S150 may specifically include the following steps:
[0179] S151: Obtain the physical feature and the video feature with the same timestamp;
[0180] S152: The physical features and the video features with the same timestamp are used as input values and fed into the first algorithm. The first algorithm is used to obtain an output value, which is used to characterize the feature matching degree.
[0181] In this embodiment, all physical sensor data and visual features are first traversed, and feature pairs that correspond one-to-one are selected based on the timestamp field: that is, each physical sensor data and its visual features that are completely consistent with its timestamp form a set of matching data (ensuring time synchronization). The physical features (such as vibration) and visual features (such as flicker frequency and brightness) in each set of matching data are used as inputs and fed into the first algorithm.
[0182] For example:
[0183] Association mapping rule: The time of "door lock relay action vibration" detected by the vibration sensor is defined as the trigger time T_phys, which should be synchronized with the "indicator light state change time" (T_vis) detected by the vision node (the door lock relay action trigger circuit is turned on and off, and the indicator light completes the state switch after a lag of ΔT0=0.2s, and there is a fixed time difference between the two).
[0184] Implementation steps:
[0185] Physical signal extraction: The vibration sensor collects the vibration of the door lock area at a sampling frequency of 500Hz. The ambient noise is filtered by wavelet filtering. When the vibration amplitude is ≥0.3g (gravitational acceleration), the trigger time T_phys is recorded.
[0186] Visual feature association: Obtain the "state change time" of the visual node output (e.g., the time T_vis when the light changes from "off" to "on").
[0187] The feature matching degree calculation is based on the first algorithm, which is as follows: First, obtain a fixed time difference ΔT0=0.2s through the initial calibration of the elevator (0.2s after the relay action, the indicator light completes the state switch), and then calculate the actual time difference ΔT_sync=|T_vis-(T_phys+ΔT0)|, and the matching degree=1-min(ΔT_sync / ΔT0,1.0) (if the difference does not exceed ΔT0, the matching degree is ≥0).
[0188] Verification of pair validity: Matching degree ≥ 0.6 → valid (action and state changes are synchronized); < 0.6 → invalid (may indicate abnormal door lock action or visual misjudgment).
[0189] Example: Relay action time T_phys = 1.0s, indicator light illumination time T_vis = 1.2s (T_phys + ΔT0 = 1.2s), ΔT = 0 → matching degree = 1.0 (valid, timing is completely synchronized); if T_phys = 1.0s, T_vis = 1.15s, ΔT = 0.05s → matching degree = 1 - 0.05 / 0.2 = 0.75 (valid, the slight time difference is a normal response delay); in the original example, T_vis = 1.15s resulted in a matching degree of 0.25, which should actually be because ΔT0 was not added during the calculation. After adjustment, it can be corrected to be valid, which is more in line with the actual response logic of the device.
[0190] The calculation logic for the overall matching degree has been adjusted to a multi-dimensional weighted fusion based on the dual physical features of vibration and visual features. The specific formula and dimension definitions are as follows:
[0191] Overall matching degree = (Vibration - State-time matching degree) ) + (Vibration amplitude - Detection frame stability matching degree) 5)
[0192] Vibration-state time matching degree: The binding matching degree between the physical modal vibration trigger time and the visual modal state switching time (value 0-1), which measures the time synchronization between the vibration signal and the visual state change of the indicator light, such as 0.85 in the exemplary scenario in the previous section.
[0193] Vibration amplitude - detection frame stability matching degree: If the variance of visual detection frame size stability ≤ preset threshold (indicating no offset of the detection frame), then when the mean vibration amplitude Amp ≥ 0.2g, the matching degree = 1.0; when the mean vibration amplitude Amp < 0.2g, the matching degree = 0.5; if the variance of detection frame stability > threshold, then the matching degree is uniformly taken as 0.3.
[0194] In S160, the first state, second state, first confidence level, second confidence level, feature matching degree, and environmental features need to be combined. The second algorithm is used to fuse the first state, second state, first confidence level, second confidence level, feature matching degree, and environmental features to determine the attention weights of the Bayesian model. Therefore, as... Figure 3 As shown, S160 specifically includes the following steps:
[0195] S161: Determine a first state consistency score based on the total number of frames of the current video data and the number of frames of the video image for the second state; and determine a second state consistency score based on all the physical sensor data and the number of physical sensor data representing the first state.
[0196] S162: Based on the environmental characteristics, determine the visual basic weights and physical basic weights;
[0197] S163: Determine the physical attention weight and visual attention weight based on the visual basic weight, physical basic weight, first confidence level, second confidence level, and feature confidence level, so that the physical attention weight and visual attention weight serve as the attention weights of the Bayesian model.
[0198] In this embodiment, the state consistency score is: First state consistency score = number of video image frames consistent with the first state / total number of current video data frames (quantifying the temporal stability of the visual state); Second state consistency score = number of physical sensor data consistent with the first state / total amount of all physical sensor data (quantifying the cooperative stability of the physical state and the visual state).
[0199] In this embodiment, the visual interference coefficient and physical interference coefficient are first calculated based on environmental characteristics (quantifying the impact of the environment on the two modalities). Specifically:
[0200] Visual interference coefficient, which is used to quantify the interference degree of the light in the machine room on the visual acquisition of the control cabinet indicator lights, and is completed based on the acquisition data of the light sensor:
[0201] (1) Data acquisition: Use a light sensor (range 0 - 10000 lux) installed on the side of the lens of an embedded high-definition intelligent camera. With a sampling frequency of 10 Hz, synchronously acquire N = 5 frames of light intensity data (unit: lux) within a 5S acquisition period. The acquired data is strictly aligned with the video image timestamp to ensure the timing matching of the light data and visual acquisition;
[0202] (2) Data preprocessing: First calculate the moving average of the N-frame light intensity, then calculate the standard deviation σ of the light data,剔除 the mutated data that deviates from the mean by ±3σ, and recalculate the filtered light mean Lavg;
[0203] (3) Interference quantification: Based on the calibration threshold of the light in the machine room (the normal light range is 100 - 5000 lux), map the result to the 0 - 1 interval:
[0204] Strong light interference:
[0205] Lavg≥5000 lux → Kvis = min(Lavg / 10000, 1.0), (the upper limit of the value is 1.0)
[0206] Weak light interference: Lavg≤100 lux → Kvis = (100 - Lavg) / 100
[0207] Normal light:
[0208] 100 < Lavg < 5000 lux → Kvis = 0.1 (low interference reference value)
[0209] The physical interference coefficient is used to quantify the interference degree of the mechanical noise in the machine room on the acquisition of the door lock relay vibration sensor, and is completed based on the acquisition data of the mechanical noise sensor:
[0210] (1) Data acquisition: Use a mechanical noise sensor (frequency band 20 - 2000 Hz, adapted to the noise frequency band of the machine room traction machine and ventilation equipment) installed on the bottom mounting board of the elevator control cabinet. Within a 5S acquisition period, synchronously acquire N = 5 frames of mechanical noise decibel values (unit: dB). The acquired data is strictly aligned with the vibration sensor data timestamp;
[0211] (2) Data preprocessing: Calculate the moving average Navg of the N-frame mechanical noise to smooth the instantaneous noise fluctuations caused by the start and stop of the traction machine, etc.;
[0212] (3) Interference quantification: Based on the anti-interference critical threshold of 80 dB of the vibration sensor, map the result to the 0 - 1 interval:
[0213] Kvib=min(Navg / 100,1.0), if Navg≥80dB, then Kvib=1.0 (interference upper limit).
[0214] The basic weights are determined based on environmental characteristics: visual interference coefficients and physical interference coefficients within the environmental characteristics, calculated using the formula: "Visual basic weight = (1 - visual interference coefficient) / [(1 - visual interference coefficient) + (1 - physical interference coefficient)]".
[0215] The calculation is as follows: "Physical base weight = (1 - physical interference coefficient) / [(1 - visual interference coefficient) + (1 - physical interference coefficient)]", ensuring that the base weight is negatively correlated with the degree of interference.
[0216] The final attention weights are calculated using a second algorithm to determine the attention weights for the corresponding Bayesian model. This second algorithm can be:
[0217] Visual attention weight = visual base weight × first confidence level × feature matching degree × first state consistency score; physical attention weight = physical base weight × second confidence level × feature matching degree × second state consistency score; after normalizing the two (the sum is 1), they are used as the attention weights of the Bayesian model.
[0218] In this embodiment, as Figure 4 As shown, S170 specifically includes the following steps:
[0219] S171: Map the visual features, physical features, and environmental features to the same space to obtain visual feature codes, physical feature codes, and environmental feature codes with the same dimension;
[0220] S172: Input the visual feature encoding, physical feature encoding, environmental feature encoding, first state, first confidence level, second state and second confidence level into the Bayesian model to obtain several third states of the elevator indicator light and the third confidence level corresponding to the third state;
[0221] S173: When the highest third confidence level is less than the first confidence level threshold, or the difference between the highest third confidence level and the second highest third confidence level is less than a preset difference, the third state is corrected based on a preset strategy to obtain the final state.
[0222] In this embodiment, three independent lightweight MLPs (corresponding to visual, physical, and environmental features respectively) are used to map the original features of different dimensions to the same high-dimensional space (e.g., 256-dimensional), resulting in visual feature encoding, physical feature encoding, and environmental feature encoding, thus eliminating the dimensional differences of heterogeneous features. Then, the three types of encoding, the first state, the first confidence level, the second state, and the second confidence level are input into the Bayesian model. The Bayesian network calculates the posterior probability of each candidate state (e.g., on, off, abnormal) using preset prior probabilities, likelihood probabilities, and attention weights, and outputs several third states and their corresponding third confidence levels. Finally, it determines whether ambiguity correction is triggered: if the highest third confidence level is less than the first confidence level threshold (e.g., 0.7), or the difference between the highest and second-highest third confidence levels is less than a preset difference (e.g., 0.2), then a preset strategy correction is initiated; otherwise, the third state corresponding to the highest third confidence level is directly used as the final state.
[0223] Prior probability ( The candidate states (such as on / off / abnormal) are initial values obtained based on the statistical frequency of a large amount of labeled historical data, and can be dynamically calibrated through real-time running data.
[0224] The specific steps to obtain it are as follows:
[0225] Collection Scope: Video and physical sensor data covering elevators of different brands, indicator lights of varying aging levels, and various daily scenarios (normal operation, minor interference, equipment maintenance). The total data volume needs to reach tens of thousands of frames (e.g., 10,000 frames) to ensure statistical representativeness.
[0226] Data filtering: Only retain valid data -- occlusion rate <10%, visual / physical confidence level ≥0.7, no extreme interference (interference coefficient K≤0.3) to avoid noisy data affecting statistical accuracy.
[0227] Manual annotation: For the filtered valid data, annotate the corresponding real state (light on / light off / abnormal) for each frame to form an annotated dataset of "data-real state".
[0228] Based on the labeled dataset, the frequency proportion of each state is statistically analyzed and directly used as the prior probability, as shown in the formula:
[0229]
[0230] To label the central state of the dataset The number of frames that appear, This represents the total number of valid frames in the labeled dataset.
[0231] Example: In 1000 frames of labeled data, 3200 frames have lights on, 6500 frames have lights off, and 300 frames have anomalies. The initial prior probability is:
[0232]
[0233]
[0234]
[0235] The third confidence score is calculated based on Bayesian equilibrium, combined with attention weights to weighted neutralize the bimodalities. The posterior probability formula is as follows:
[0236]
[0237] in: Used to represent prior probability;
[0238] Joint likelihood probability (Fusing visual and physical modalities, with attention weights as weighted coefficients):
[0239]
[0240] P(D) is the evidence factor (normalization constant), ensuring that the sum of the posterior probabilities of all states is 1.
[0241]
[0242] Third state set: .
[0243] Indicates visual attention weights, The physical attention weights are calculated in step S163.
[0244] Feature mapping solves the problem of multimodal feature heterogeneity, laying the foundation for unified reasoning of Bayesian models; the probabilistic reasoning capability of Bayesian models effectively quantifies the uncertainty of each modality and integrates multi-source evidence to output candidate states; the ambiguity correction triggering mechanism can accurately identify uncertain scenarios, avoid directly outputting unreliable results, and further improve the reliability of the final state through subsequent preset strategy correction, so that the solution can still maintain stable recognition accuracy in complex scenarios.
[0245] In this embodiment, the preset strategy can be:
[0246] When there is a unique high-reliability state, the high-reliability state is set as the final state; when there are two or more high-reliability states, the first high-reliability state that is the same as the first state or the second state with a feature matching degree greater than the first matching degree threshold is set as the final state. The first high-reliability state is the third state corresponding to the third confidence degree that is greater than the second confidence degree threshold and whose visual interference coefficient and physical interference coefficient are both greater than the interference coefficient threshold.
[0247] When the first high-reliability state does not exist, the second high-reliability state is set as the final state. The second high-reliability state represents the first state or the second state where the matching degree threshold is greater than the second matching degree threshold. Specifically, when the corresponding visual basic weight is greater than the physical basic weight, the second high-reliability state is the first state, and when the corresponding visual basic weight is less than the physical basic weight, the second high-reliability state is the second state.
[0248] When there is no second highly reliable state, let the third reliable state be the final state, and the third reliable state represents the state with the largest number of states.
[0249] The first step is to select the first highly reliable state: find the third state with a third confidence level greater than the second confidence level threshold (e.g., 0.6) and both the visual interference coefficient and the physical interference coefficient less than the interference coefficient threshold (e.g., 0.4). If there is only one such state, it is directly used as the final state; if there are two or more, select the one that matches the first state or the second state with a feature matching degree greater than the first matching degree threshold (e.g., 0.7) as the final state.
[0250] The second step, if no first highly reliable state exists, select the second most reliable state: find the first or second state whose feature matching degree is greater than the second matching degree threshold (e.g., 0.6). If the visual basis weight is greater than the physical basis weight, then the second most reliable state is the first state; otherwise, it is the second state, which is the final state. The third step, if no second most reliable state exists, select the third most reliable state: count the states that appear most frequently among all the first, second, and third states, and use this as the final state.
[0251] In this embodiment, a prioritization strategy of "first high-reliability state → second high-reliability state → third high-reliability state" systematically resolves the ambiguity problem in the Bayesian model output. Prioritizing states with low interference, high confidence, and high consensus ensures the reliability of the correction result; intermediate levels combine basic weights to determine the dominant mode, taking into account modality adaptability; and a fallback majority voting strategy avoids infinite loops where no valid states are available. The entire correction process is logically rigorous and progressive, improving the accuracy of state recognition in ambiguous scenarios and enhancing the robustness and fault tolerance of the solution.
[0252] The overall solution in this embodiment can be as follows:
[0253] 1. Synchronous acquisition and time calibration of multimodal data
[0254] Deploy dedicated equipment and collect data according to rules: elevator-specific high-definition cameras capture video of indicator lights, vibration sensors collect vibration data of door lock relays, and light / mechanical noise sensors collect environmental data; with a 5-second acquisition cycle, 5 frames of data are collected in each cycle. The timing synchronization of the three types of data is achieved through NTP / PTP global clock + hardware trigger + software interpolation compensation (error ≤10ms), ensuring the timing consistency of subsequent cross-modal processing.
[0255] 2. Single-modal feature extraction and preliminary state / confidence output
[0256] Visual modality: The improved YOLO8 model is used to extract the visual features of the indicator lights, and the first state of the indicator lights (such as on / off / abnormal) and the first confidence level are output. The visual features are then verified based on a preset template of historical valid data. If the features exceed the normal range, the first state and confidence level are corrected, and lighting and occlusion interference are filtered.
[0257] Vibration physical modes: Process vibration sensor data, aggregate vibration temporal synchronization and amplitude stability characteristics, obtain the second state (corresponding to the door lock action state) through N-frame state voting, and then calculate the second confidence level by weighting vibration reliability, cross-modal matching degree, and environmental anti-interference.
[0258] Environmental modalities: Based on illumination / mechanical noise data, visual interference coefficients (quantifying the interference of illumination on visual acquisition) and physical interference coefficients (quantifying the interference of noise on vibration acquisition) are calculated separately for subsequent weight adjustment.
[0259] 3. Cross-modal feature matching
[0260] Vibration and visual features are matched by timestamp, and the comprehensive matching degree is calculated by weighted algorithm: the temporal synchronization of vibration and indicator light status and the stability of vibration amplitude and visual detection box are integrated to quantify the synergistic consistency of the two types of modal features, and provide a basis for cross-modal fusion.
[0261] 4. Bayesian Fusion Attention Weight Calculation
[0262] By combining the temporal stability of visual / vibration states, the degree of environmental interference, the confidence of each modality, and the cross-modal matching degree, the visual / physical attention weights of the Bayesian model are calculated: the basic weights are negatively correlated with environmental interference, and the final weights are adjusted by combining confidence, matching degree, and state consistency, and then normalized to serve as the basis for the model's attention allocation.
[0263] 5. Bayesian Model Inference and Final State Determination
[0264] After mapping visual, vibration, and environmental features to the same high-dimensional space, the data is input into a Bayesian model. Combining prior probabilities and attention weights, the model outputs candidate states and their corresponding confidence levels. If a candidate state is ambiguous (confidence is too low or the difference in confidence levels between multiple states is too small), it is corrected according to a three-level priority strategy: first, select a high-reliability state with low interference and high confidence; if none is found, select the dominant modality state with high matching degree; if still none is found, select the state with the highest frequency of occurrence. Finally, the safety status of the elevator door lock / indicator light is output.
[0265] In this embodiment, the processing module can be an integrated circuit chip with signal processing capabilities. The processing module can be a general-purpose processor. For example, the processor can be a Central Processing Unit (CPU), a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application.
[0266] The storage module can be, but is not limited to, random access memory, read-only memory, programmable read-only memory, erasable programmable read-only memory, electrically erasable programmable read-only memory, etc.
[0267] It should be noted that those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the electronic device described above can be referred to the corresponding steps in the aforementioned method, and will not be elaborated further here.
[0268] This application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program that, when run on a computer, causes the computer to execute the elevator indicator light status determination method described in the above embodiments.
[0269] Based on the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by hardware or by using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application can be embodied in the form of a software product. This software product can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard drive, etc.) and includes several instructions to cause a computer device (such as a personal computer, electronic device, or network device, etc.) to execute the methods described in the various implementation scenarios of this application.
[0270] In the embodiments provided in this application, it should be understood that the disclosed apparatus, systems, and methods can also be implemented in other ways. The apparatus, systems, and methods embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code, which includes one or more executable instructions for implementing a specified logical function. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions. Furthermore, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.
[0271] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A method for determining the state of an elevator indicator light, characterized by The method comprises: acquiring current video data, N physical sensor data, and N environment sensor data, wherein the current video data comprises N frames of video images, each of the physical sensor data and environment sensor data has a timestamp corresponding to a timestamp of a video image, and N is a natural number greater than or equal to 2; inputting the current video data into a preset model to obtain, by the preset model, a first state of an elevator indicator light in each of the video images, a corresponding first confidence, and a detection box of the elevator indicator light; determining, according to all the physical sensor data, a second state of the elevator indicator light and a corresponding second confidence; obtaining physical features according to all the physical sensor data, visual features according to all the detection boxes, and environment features according to all the environment sensor data; determining, according to the physical features and the video features, a feature matching degree of the physical features and the video features based on a first algorithm, wherein the first algorithm is used to take the physical features and the video features as inputs and output the feature matching degree; determining, according to the first state, the second state, the first confidence, the second confidence, the feature matching degree, and the environment features, an attention weight of a preset Bayesian model based on a second algorithm, and taking the attention weight as a model parameter of the Bayesian model, wherein the second algorithm is used to fuse the first state, the second state, the first confidence, the second confidence, the feature matching degree, and the environment features to determine the attention weight of the Bayesian model; inputting the visual features, the physical features, the environment features, the first state, the first confidence, the second state, and the second confidence into the Bayesian model to obtain, by the Bayesian model, a final state of the elevator indicator light in each of the video images; the preset model comprises a backbone network and a neck network; the backbone network comprises a first GhostConv module group, a second GhostConv module group, and a third GhostConv module group, the first GhostConv module group comprises a plurality of first GhostConv modules, the second GhostConv module group comprises a plurality of second GhostConv modules, and the third GhostConv module group comprises a plurality of third GhostConv modules, the first GhostConv module group is used to extract basic visual features of the video images to obtain a basic visual feature map, the second GhostConv module group is used to generate a ghost feature map based on the basic visual feature map, and the third GhostConv module group is used to splice the ghost feature map and the basic visual feature map to obtain an output feature map; the neck network comprises a REC3 module, and the neck network is deployed with a SIMAM attention mechanism; the determining, according to the physical features and the video features, of the feature matching degree of the physical features and the video features based on the first algorithm comprises: acquiring the physical features and the video features having the same timestamp; inputting the physical feature and the video feature with the same timestamp as input values into a first algorithm, and obtaining an output value through the first algorithm, the output value being used to represent the feature matching degree; determining, based on a second algorithm, an attention weight of a preset Bayesian model according to the first state, the second state, the first confidence, the second confidence, the feature matching degree, and the environmental feature, including: determining a first state consistency score according to a total frame number of the current video data and a frame number of the video image of the second state, and determining a second state consistency score according to all the physical sensor data and a number of physical sensor data representing the first state; determining a visual basic weight and a physical basic weight based on the environmental feature; determining a physical attention weight and a visual attention weight according to the visual basic weight, the physical basic weight, the first confidence, the second confidence, and the feature confidence, so that the physical attention weight and the visual attention weight are used as the attention weight of the Bayesian model.
2. The method of claim 1, wherein, After obtaining the visual feature according to all the detection frames, the method further includes: verifying the visual feature based on a preset template to obtain a verification result, the preset template including a preset visual feature range obtained according to a plurality of historical video data, and a blocking rate of a region where the elevator indicator light is located in the plurality of historical video data being less than a first preset value and an inter-frame brightness variance being less than a second preset value; when the verification result is used to represent that the video feature does not match the preset visual feature range, then correcting the first state and the first confidence based on the video feature.
3. The method of claim 1, wherein, inputting the visual feature, the physical feature, the environmental feature, the first state, the first confidence, the second state, and the second confidence into the Bayesian model, and obtaining the final state of the elevator indicator light through the Bayesian model, including: mapping the visual feature, the physical feature, and the environmental feature into the same space to obtain visual feature encoding, physical feature encoding, and environmental feature encoding with the same dimension; inputting the visual feature encoding, the physical feature encoding, the environmental feature encoding, the first state, the first confidence, the second state, and the second confidence into the Bayesian model to obtain a plurality of third states of the elevator indicator light and third confidence corresponding to the third states; when the highest third confidence is less than a first confidence threshold or a difference between the highest third confidence and a second highest third confidence is less than a preset difference, correcting the third state based on a preset strategy to obtain the final state.
4. The method of claim 3, wherein, the correcting the third state based on the preset strategy to obtain the final state, including: calculating a visual interference coefficient and a physical interference coefficient based on the environmental feature; When there is only one high-reliability state, the high-reliability state is the final state; when there are two or more high-reliability states, a first high-reliability state that is the same as the first state or the second state with a feature matching degree greater than a first matching degree threshold and a third state corresponding to a third confidence degree greater than a second confidence degree threshold and both a visual interference coefficient and a physical interference coefficient greater than an interference coefficient threshold is the final state; When there is no first high-reliability state, a second high-reliability state is the final state, the second high-reliability state representing a first state or a second state with a matching degree greater than a second matching degree threshold, wherein the second high-reliability state is the first state when a corresponding visual basic weight is greater than a physical basic weight, and the second high-reliability state is the second state when the corresponding visual basic weight is less than the physical basic weight; When there is no second high-reliability state, a third high-reliability state is the final state, the third high-reliability state representing the state with the largest number.
5. An electronic device, comprising: The electronic device includes a processor and a memory coupled to each other, and the memory stores a computer program, which, when executed by the processor, causes the electronic device to perform the method of any one of claims 1-4.
6. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a computer program, which, when executed on a computer, causes the computer to perform the method of any one of claims 1-4.
Citation Information
Patent Citations
Method, device and system for intelligently detecting target based on YOLO model
CN120877167A
Elevator safety monitoring system fused with computer vision and method thereof
CN120943081A