Escalator comb tooth foreign matter detection and early warning method and system
By combining the preprocessing and feature fusion of visual and audio signals, the problem of low detection accuracy of foreign objects in the escalator comb plate area is solved, and high-precision detection in multiple scenarios is achieved.
Patent Information
- Application Number
- CN202510689231.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-07-08
AI Technical Summary
In the detection of foreign objects in the escalator comb plate area, light changes and camera occlusion lead to low detection accuracy and difficult to identify masked foreign objects.
Combining visual and audio signals, foreign object detection is performed, features are extracted by preprocessing image and voiceprint signals, weighted fusion is performed, comprehensive detection scores are generated and warning level division is performed.
In the strong, dark light and camera occlusion scenarios, the accuracy and reliability of foreign object detection are significantly improved, and the detection blind spots are made up.
Smart Images

Figure CN120270888A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of object recognition, and in particular to a method and system for detecting and warning foreign objects in escalator comb teeth. Background Art
[0002] As the core equipment for efficiently transporting passengers in public places, the safety of escalators is directly related to public safety in large-scale passenger flow scenarios. In recent years, with the dense deployment of escalators in shopping malls, airports, high-speed rail stations and other scenarios, foreign object entrapment accidents in the comb plate area (the meshing point between the steps and the fixed end) have become one of the main risk sources.
[0003] At present, vision-based detection solutions are widely used as the main warning method for comb plate foreign objects. However, when the lighting conditions change, such as the strong overhead light in subway stations and the dark light environment in shopping malls, it is easy to cause overexposure or underexposure of the image, reducing the recognition rate of foreign object contours; secondly, when there are dense passengers, luggage, handbags and other items may temporarily block the camera's field of view, causing detection interruption; and if the foreign object is partially covered by the step tooth groove, such as a coin with only the edge exposed, it is difficult for visual detection to identify the foreign object, resulting in low detection accuracy of escalator comb foreign objects. Summary of the invention
[0004] Based on this, the object of the present invention is to provide an escalator comb tooth foreign object detection and early warning method which improves the efficiency of escalator comb tooth foreign object detection and early warning.
[0005] An escalator comb foreign body detection and early warning method, comprising:
[0006] Acquire a detection image of the comb plate area of the escalator, and pre-process the detection image to obtain a pre-processed image of the comb plate area;
[0007] Acquire an audio signal of the comb plate area of the escalator, and pre-process the audio signal to obtain a pre-processed voiceprint of the comb plate area;
[0008] Based on a preset visual foreign body detection model, foreign body recognition is performed on the preprocessed image of the comb plate area to obtain a foreign body visual confidence level;
[0009] Perform multi-branch audio feature extraction on the pre-processed voiceprint in the comb plate area to obtain a voiceprint fusion feature;
[0010] According to the voiceprint fusion feature, based on a preset foreign object sound matching mechanism, a foreign object voiceprint matching degree is obtained;
[0011] Based on the preset visual confidence weight and voiceprint matching weight, the foreign object visual confidence and foreign object voiceprint matching are weightedly fused to obtain a comprehensive foreign object detection score;
[0012] Based on the comprehensive foreign object detection score, according to the preset foreign object detection warning level division rule, the warning level of the escalator comb plate area is divided;
[0013] According to the warning level division result of the escalator comb plate area, a warning instruction corresponding to the warning level is generated, and the warning instruction is transmitted to the escalator warning device.
[0014] This application also provides an escalator comb foreign object detection and warning system, including:
[0015] Comb plate area preprocessing image acquisition module: used to acquire the detection image of the escalator comb plate area, and preprocess the detection image to obtain the preprocessed image of the comb plate area;
[0016] Comb plate area preprocessing voiceprint acquisition module: used to acquire the audio signal of the escalator comb plate area, and preprocess the audio signal to obtain the preprocessed voiceprint of the comb plate area;
[0017] Foreign object visual confidence acquisition module: used to identify foreign objects in the preprocessed image of the comb plate area based on a preset visual foreign object detection model, and obtain the foreign object visual confidence;
[0018] Voiceprint fusion feature acquisition module: used to extract multi-branch audio features from the preprocessed voiceprint of the comb plate area to obtain the voiceprint fusion feature;
[0019] Foreign object voiceprint matching degree acquisition module: used to obtain the foreign object voiceprint matching degree according to the voiceprint fusion feature based on a preset foreign object sound matching mechanism;
[0020] Foreign object comprehensive detection score acquisition module: used to perform weighted fusion on the foreign object visual confidence and the foreign object voiceprint matching degree respectively based on the preset visual confidence weight and voiceprint matching degree weight to obtain the foreign object comprehensive detection score;
[0021] Warning level division module: used to divide the warning level of the escalator comb plate area according to the foreign object comprehensive detection score based on the preset foreign object detection warning level division rule;
[0022] Warning instruction transmission module: used to generate a warning instruction corresponding to the warning level according to the warning level division result of the escalator comb plate area, and transmit the warning instruction to the escalator warning device.
[0023] Compared with the prior art, the present application performs foreign object detection by simultaneously preprocessing the image and the voiceprint in the comb plate area of the escalator, effectively making up for the detection blind spots in scenes with strong top light, low light, and blocked cameras, and significantly improving the detection accuracy and reliability of foreign objects in the comb of the escalator.
[0024] In order to understand the present application more clearly, the following will describe the specific implementation manners of the present application in conjunction with the accompanying drawings. Description of the Drawings
[0025] Figure 1 It is a flowchart of a method for detecting and warning foreign objects in the comb of an escalator according to the present application;
[0026] Figure 2 It is a flowchart of a method for obtaining a preprocessed image of the comb plate area in a method for detecting and warning foreign objects in the comb of an escalator according to the present application;
[0027] Figure 3 It is a schematic installation diagram of a wide-dynamic camera in a method for detecting and warning foreign objects in the comb of an escalator according to the present application;
[0028] Figure 4 It is a flowchart of a method for dynamically adjusting the exposure time in a method for detecting and warning foreign objects in the comb of an escalator according to the present application;
[0029] Figure 5 It is a flowchart of a method for obtaining a preprocessed voiceprint of the comb plate area in a method for detecting and warning foreign objects in the comb of an escalator according to the present application;
[0030] Figure 6 It is a schematic installation diagram of a microphone in a method for detecting and warning foreign objects in the comb of an escalator according to the present application;
[0031] Figure 7 It is a flowchart of a method for noise cancellation processing in a method for detecting and warning foreign objects in the comb of an escalator according to the present application;
[0032] Figure 8 It is a flowchart of a method for obtaining the visual confidence of foreign objects in a method for detecting and warning foreign objects in the comb of an escalator according to the present application;
[0033] Figure 9 It is a flowchart of a method for obtaining a weighted feature map in a method for detecting and warning foreign objects in the comb of an escalator according to the present application;
[0034] Figure 10 It is a flowchart of a method for obtaining a voiceprint fusion feature in a method for detecting and warning foreign objects in the comb of an escalator according to the present application;
[0035] Figure 11 It is a flowchart of a method for obtaining the voiceprint matching degree of foreign objects in a method for detecting and warning foreign objects in the comb of an escalator according to the present application;
[0036] Figure 12 This is a schematic diagram of an escalator comb foreign object detection and warning system for the present application. DETAILED DESCRIPTION
[0037] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.
[0038] It should be understood that the flow chart drawings are not drawn in real proportion. The flow charts used in this application illustrate the operations implemented according to some embodiments of the application. It should be understood that the operations of the flow charts may be implemented out of order, and steps without logical contextual relationships may be reversed in order or implemented simultaneously. In addition, those skilled in the art may add one or more other operations to the flow charts, or may remove one or more operations from the flow charts, under the guidance of the content of this application.
[0039] Reference to "embodiments" herein means that a particular feature, structure, or characteristic described in conjunction with the embodiments may be included in at least one embodiment of the present application. The appearance of the phrase in various locations in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment that is mutually exclusive with other embodiments. It is explicitly and implicitly understood by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0040] Example 1
[0041] See also Figure 1 , Figure 1 This is a flow chart of a method for detecting and warning foreign objects in escalator comb teeth. This application provides a method for detecting and warning foreign objects in escalator comb teeth, including:
[0042] S1: Acquire a detection image of the comb plate area of the escalator, and pre-process the detection image to obtain a pre-processed image of the comb plate area;
[0043] S2: Acquire an audio signal of the comb plate area of the escalator, and pre-process the audio signal to obtain a pre-processed voiceprint of the comb plate area;
[0044] S3: Based on a preset visual foreign body detection model, foreign body recognition is performed on the pre-processed image of the comb plate area to obtain a visual confidence level of the foreign body;
[0045] S4: Perform multi-branch audio feature extraction on the preprocessed voiceprint in the comb plate area to obtain voiceprint fusion features;
[0046] S5: Based on the voiceprint fusion features, obtain the foreign object voiceprint matching degree based on a preset foreign object sound matching mechanism;
[0047] S6: Based on preset visual confidence weights and voiceprint matching degree weights, perform weighted fusion on the foreign object visual confidence and the foreign object voiceprint matching degree respectively to obtain a comprehensive foreign object detection score;
[0048] S7: Based on the comprehensive foreign object detection score, perform warning level division on the comb plate area of the escalator based on a preset foreign object detection warning level division rule;
[0049] S8: According to the warning level division result of the comb plate area of the escalator, generate a warning instruction corresponding to the warning level, and transmit the warning instruction to the escalator warning device.
[0050] Compared with the prior art, the present application performs foreign object detection by simultaneously using the preprocessed image and the preprocessed voiceprint of the comb plate area of the escalator, effectively making up for the detection blind areas in scenes with strong top light, low light, and blocked cameras, and significantly improving the detection accuracy and reliability of foreign objects in the escalator comb.
[0051] In this embodiment, for step S1, please refer to Figure 2 , Figure 2 is the flowchart of the method for obtaining the preprocessed image of the comb plate area in a foreign object detection and warning method for an escalator comb of the present application. The method of obtaining the detection image of the comb plate area of the escalator and preprocessing the detection image to obtain the preprocessed image of the comb plate area includes:
[0052] S11: Obtain the detection image taken by a camera module with a preset exposure time whose field of view angle can cover the meshing area of the comb and the step of the escalator;
[0053] S12: Based on the detection image, determine the area coordinates of the area including the meshing area of the comb and the step in the escalator, and crop the detection image to obtain a target area detection image;
[0054] S13: Perform a normalization operation on the target area detection image to obtain the preprocessed image of the comb plate area.
[0055] For step S11, the camera module selects a wide-dynamic-range camera that supports a resolution of 1920×1080 and a frame rate of ≥24fps. The wide-dynamic-range camera is connected to a camera bracket and installed 1 meter directly above the comb plate, with the optical axis maintaining an angle of about 45 degrees with the ground. The field of view can cover the meshing area between the comb teeth and the steps, as Figure 3 shown.
[0056] In one embodiment, please refer to Figure 4 simultaneously. Figure 4 which is a flowchart of the method for dynamically adjusting the exposure time in a method for detecting and warning foreign objects on the comb teeth of an escalator according to this application. Dynamically adjusting the exposure time includes:
[0057] S111: Obtain the step movement speed of the escalator;
[0058] S112: According to the step movement speed, based on motion compensation, dynamically adjust the exposure time according to the following formula to obtain the adjusted exposure time:
[0059]
[0060] In the formula, T exp is the adjusted exposure time, v is the step movement speed, and p is the preset pixel size.
[0061] For steps S111 - S112, the step movement speed can be obtained by installing a rotary encoder on the step sprocket or the drive motor shaft, calculating the rotational speed through pulse counting, and then converting it into the step linear speed in combination with the transmission ratio, or by installing a tachometer to calculate the speed of the step surface. By adjusting the exposure time of the camera module through motion compensation, it is avoided that when the step movement speed is relatively fast, the image of the comb plate area captured is prone to motion blur, which interferes with foreign object detection. And when adjusting the exposure time of the camera module, the minimum time interval for each adjustment is 0.001s, so that the adjusted exposure time better meets the actual step movement speed requirements, avoiding excessive or too small adjustment amplitude, thereby better ensuring that the motion blur length is less than 1 pixel and improving the quality of the detected image.
[0062] For step S12, for the detected image, through image processing techniques such as edge detection, etc., locate the position of the meshing area between the comb teeth and the steps in the escalator as the region of interest (ROI), and extract the center coordinates, width, and height of the ROI area as the region coordinates. In other embodiments, the region coordinates of the ROI area can also be predicted based on the step movement trajectory of the escalator and the Kalman filtering algorithm.
[0063] For step S13, the normalization operation on the target region detection image includes: performing gray normalization, geometric normalization, pixel normalization, etc. on the target region detection image.
[0064] For step S2, please refer to Figure 5 , Figure 5 which is a flowchart of a method for obtaining a preprocessed acoustic fingerprint of a comb plate region in an escalator foreign object detection and warning method according to the present application. The method for obtaining an audio signal of an escalator comb plate region and preprocessing the audio signal to obtain a preprocessed acoustic fingerprint of the comb plate region includes:
[0065] S21: Obtain the audio signal collected by a microphone module that forms a microphone array in the escalator comb plate region;
[0066] S22: Calculate the signal-to-noise ratio of the audio signal. If the signal-to-noise ratio is lower than a preset signal-to-noise ratio threshold, perform amplification and / or attenuation processing on the audio signal to obtain an optimized audio signal;
[0067] S23: Based on a preset noise fingerprint library, perform noise cancellation processing on the optimized audio signal to obtain the preprocessed acoustic fingerprint of the comb plate region.
[0068] For steps S21 - S22, install 1 MEMS microphone at each of the four corners of the escalator comb plate, and all are 20 mm away from the edge of the comb plate; install 2 MEMS microphones at the positions of the quarter lines of the escalator comb plate, that is, at 1 / 4 and 3 / 4 of the length of the comb plate, to form a 6-microphone array, as specifically shown in Figure 6 shown.
[0069] Each of the MEMS microphones integrates a 20 Hz - 20 kHz band-pass filter and a PGA programmable gain amplifier, with a dynamic range of 80 dB. Even in an environment with 100 dB of ambient noise, an effective signal with a signal-to-noise ratio of 20 dB can be extracted. The signal-to-noise ratio threshold is set to 30 dB. When the signal-to-noise ratio of the audio signal is less than 30 dB, the gain of the microphone array can be increased through the programmable gain amplifier to perform amplification and / or attenuation processing on the audio signal to obtain the optimized audio signal. Of course, in other embodiments, the signal-to-noise ratio threshold can be adaptively modified according to the actual scenario.
[0070] For step S23, the inherent noise of the escalator such as motor harmonics and step vibrations and other periodic noises can be collected. Based on the hidden Markov model, the time-frequency features of the inherent noise such as Mel frequency spectrum coefficients or sub-band energies are used as the observation sequence, and the model parameters such as the state transition matrix and the observation probability matrix are adjusted through unsupervised learning such as the Baum-Welch algorithm.
[0071] Store the set of model parameters after adjustment as the noise fingerprint library, and any one of the model parameters in the set of model parameters can uniquely identify a type of noise.
[0072] Please also refer to Figure 7 , Figure 7 which is the method flowchart of noise elimination processing in a method for detecting and warning foreign objects in a comb of an escalator according to this application. Based on the preset noise fingerprint library, perform noise elimination processing on the optimized audio signal to obtain the preprocessed soundprint of the comb plate area, including:
[0073] S231: Perform short-time frame segmentation and windowing processing on the optimized audio signal, and extract the noise characteristics of each frame;
[0074] S232: Input the noise characteristics into the set of model parameters in the noise fingerprint library, and determine the type of current noise through a decoding algorithm;
[0075] S233: Determine the corresponding noise spectrum template according to the type of the noise, and based on spectral subtraction, weaken the noise spectrum components in the optimized audio signal to obtain the preprocessed soundprint of the comb plate area.
[0076] For steps S231 - S232, after performing short-time frame segmentation processing on the continuous optimized audio signal, perform windowing processing such as applying a window function such as a Hamming window, which can reduce spectral leakage and improve spectral resolution. Perform short-time Fourier transform on the optimized audio signal of each frame to obtain the noisy signal spectrum. Calculate the Mel-frequency cepstral coefficients or sub-band energy of the noisy signal spectrum as the noise characteristics.
[0077] For step S233, the decoding algorithm can be an algorithm such as Viterbi. Subtract the noise spectrum template from the noisy signal spectrum to obtain the denoised soundprint signal spectrum, and perform inverse Fourier transform on the soundprint signal spectrum to obtain the preprocessed soundprint of the comb plate area.
[0078] For step S3, the visual foreign object detection model is an improved YOLO model, including: a backbone network, a neck network, and a detection head integrated with an attention module. The backbone network adopts the YOLO-Fastest-XL network architecture, the neck network adopts the FPN structure, and the detection head integrates the SE-Net attention module.
[0079] Please refer to Figure 8 , Figure 8 which is the method flowchart of obtaining the visual confidence of foreign objects in a method for detecting and warning foreign objects in a comb of an escalator according to this application. Based on the preset visual foreign object detection model, perform foreign object recognition on the preprocessed image of the comb plate area to obtain the visual confidence of foreign objects, including:
[0080] S31: Extract features from the preprocessed image of the comb plate region based on several convolutional layers of the skeletal network to obtain several large-scale feature maps and small-scale feature maps;
[0081] S32: Respectively perform downsampling operations on the large-scale feature maps and upsampling operations on the small-scale feature maps through the neck network to obtain downsampled feature maps and upsampled feature maps, and perform convolutional fusion on the downsampled feature maps and upsampled feature maps to obtain fused feature maps;
[0082] S33: Perform channel weighting on the fused feature maps through the detection head to obtain weighted feature maps, and generate foreign object recognition results according to the weighted feature maps, where the foreign object recognition results include: foreign object bounding box coordinates, foreign object category information, and the visual confidence of the foreign object.
[0083] For step S31, the preprocessed image of the comb plate region sequentially passes through the convolutional layers of the skeletal network, and the convolutional kernels of the convolutional layers slide and convolve on the preprocessed image of the comb plate region to extract the features in the preprocessed image of the comb plate region to obtain the large-scale feature maps and small-scale feature maps. Among them, the output size of the large-scale feature maps is 100×75×192, and the receptive field range is 16×16 pixels. The output size of the small-scale feature maps is 50×38×96, and the receptive field range is 32×32 pixels. By outputting the large-scale feature maps, more details of the preprocessed image of the comb plate region can be retained, which is suitable for detecting foreign objects with a size of more than 50×50 pixels; by outputting the small-scale feature maps, high-level semantic features can be retained, which is suitable for detecting foreign objects with a size in the range of 20×20 - 50×50 pixels, thereby jointly improving the accuracy of detecting foreign objects.
[0084] For step S33, perform upsampling operations on the small-scale feature maps through methods such as bilinear interpolation and / or transposed convolution to restore the lost semantic information of the small-scale feature maps. Perform downsampling operations on the large-scale feature maps through methods such as max pooling and / or strided convolution to align the spatial resolutions of the downsampled feature maps and upsampled feature maps, specifically including: spatial alignment and channel alignment. Among them, the spatial alignment means that the spatial dimensions of the downsampled feature maps and upsampled feature maps are made consistent through upsampling and downsampling operations, such as both being 100×75. The channel alignment means that the number of channels of the downsampled feature maps and upsampled feature maps is made consistent, such as both being 256.
[0085] The ways of performing convolutional fusion on the downsampled feature maps and upsampled feature maps include additive fusion and concatenation fusion, and the size of the obtained fused feature maps is 100×75×256.
[0086] For step S34, please refer to Figure 9 , Figure 9 which is the flowchart of the method for obtaining the weighted feature map in an escalator comb foreign object detection and warning method of this application. The detection head performs channel weighting on the fusion feature map to obtain the weighted feature map, including:
[0087] S341: According to the following formula, perform global averaging on the fusion feature map to generate a channel global feature vector:
[0088]
[0089] In the formula, z c is the channel global feature vector under the c-th channel, H and W are the spatial dimensions of the fusion feature map, and u c (i, j) is the feature value of the c-th channel in the fusion feature map at the coordinate (i, j);
[0090] S342: According to the channel global feature vector, according to the following formula, learn the channel weights through a fully connected network to generate a channel attention vector formula:
[0091] s c = σ(W2 × δ(W1 × z c ))
[0092] In the formula, s c is the channel attention vector under the c-th channel, σ is the ReLU activation function, δ is the Sigmoid activation function, W1 is the first weight matrix of the first layer of the fully connected network, and W2 is the second weight matrix of the second layer of the fully connected network;
[0093] S343: Multiply the corresponding channels of the channel attention vector and the fusion feature map to obtain the weighted feature map.
[0094] For steps S341 - S343, the number of layers of the fully connected network is two, and the first weight matrix and the second weight matrix can be updated through the backpropagation algorithm according to actual needs.
[0095] The SE-Net module first performs global average pooling on the fusion feature map to obtain a global feature description at the channel level. Then, it learns the channel dependence relationship through two layers of fully connected networks to generate a channel attention vector, weights each channel of the fusion feature map, enhances the key feature channels related to foreign objects, and suppresses irrelevant channels.
[0096] For step S4, please refer to Figure 10 , Figure 10This is a flow chart of a method for obtaining voiceprint fusion features in an escalator comb foreign body detection and warning method in this application. The multi-branch audio feature extraction of the comb plate area pre-processed voiceprint to obtain the voiceprint fusion feature includes:
[0097] S41: dividing the continuous pre-processed voiceprints of the comb plate area into a number of short time frames, and calculating the spectrum features of the short time frames;
[0098] S42: extracting a number of frequency points with the highest energy from the frequency spectrum features of each of the short-time frames to form a time-frequency energy distribution feature of a corresponding dimension;
[0099] S43: performing time-frequency dynamic analysis on the frequency spectrum characteristics of each of the short-time frames through a filter bank to obtain time-frequency dynamic distribution characteristics;
[0100] S44: performing nonlinear feature extraction on the preprocessed voiceprint of the comb plate region through a deep residual feature extraction network to obtain a global voiceprint representation feature;
[0101] S45: The time-frequency energy distribution feature, the time-frequency dynamic distribution feature and the global voiceprint representation feature are concatenated and then subjected to a dimensionality reduction operation to obtain the voiceprint fusion feature.
[0102] For steps S41-S42, the continuous pre-processed voiceprint of the comb plate area is divided into several short-time frames through 512-point FFT and 75% overlapping Hann window, specifically: frame length = 512 / sampling rate, frame shift = 512×(1-0.75) = 128 points, and the sampling rate can be adaptively adjusted according to actual needs. Fourier transform and single-sided amplitude spectrum calculation are performed on each of the short-time frames to obtain the spectrum characteristics of the short-time frame.
[0103] The 128 frequency points with the highest energy are selected from the frequency spectrum features of each of the short-time frames for feature encoding to obtain the time-frequency energy distribution features of 128 dimensions.
[0104] For step S43, in this embodiment, the filter group includes 40 Mel filters. The spectral features of the short-time frame are processed by the 40 Mel filters to obtain 40-dimensional Mel energy. The 40-dimensional Mel energy is logarithmically processed to obtain logarithmic energy. The logarithmic energy is discrete cosine transformed to convert the signal from the time domain to the cepstrum domain, and the first 13 coefficients after the discrete cosine transformation are selected as 13-dimensional Mel frequency cepstrum coefficients (Mel Frequency Cepstrum Coefficient, MFCC). The 13-dimensional Mel frequency cepstrum coefficients capture the static spectral envelope characteristics of the voiceprint and are sensitive to the inherent properties of the comb plate (such as structural stiffness, tooth shape error) and operating status (such as vibration, resonance).
[0105] Calculate the first-order difference of MFCC between adjacent short-time frames to obtain 13-dimensional Mel-frequency cepstral difference coefficients. The 13-dimensional Mel-frequency cepstral coefficients capture the time-domain dynamic characteristics of the voiceprint, reflect the time-varying law of the running state of the comb plate, and can capture the transient changes of the voiceprint features, such as sudden changes in vibration energy, and can effectively capture the voiceprint fluctuations of foreign objects stuck in the meshing area between the comb teeth and the steps of the escalator. Combine the 13-dimensional Mel-frequency cepstral coefficients and 13-dimensional Mel-frequency cepstral difference coefficients to form the 26-dimensional time-frequency dynamic distribution features.
[0106] Of course, in other embodiments, the number of Mel filters can be adjusted according to actual needs.
[0107] For step S44, the deep residual feature extraction network is the ResNet-18 model. The first 4 residual blocks of the ResNet-18 model extract multi-scale features of the preprocessed voiceprint in the comb plate area layer by layer through stacked convolutional layers, batch normalization (BN), and ReLU activation functions. After a 512-dimensional feature vector is output from the fourth residual block, after global averaging of the 512-dimensional feature vector, the 512-dimensional feature vector is reduced to 128 dimensions to obtain the global voiceprint characterization feature, which is sensitive to the chaotic signals (with wide-spectrum and low-dimensional attractor characteristics) generated by friction when foreign objects are stuck in the meshing area between the comb teeth and the steps of the escalator.
[0108] For step S45, after splicing the 128-dimensional time-frequency energy distribution features, 26-dimensional time-frequency dynamic distribution features, and 128-dimensional global voiceprint characterization features, a 282-dimensional feature vector is formed. Remove the noise dimensions through the principal component analysis method, and reduce the 282-dimensional feature vector to 256 dimensions to obtain the voiceprint fusion feature. The high-frequency friction signal is located through the time-frequency energy distribution feature, the frequency modulation caused by the vibration of foreign objects stuck in the meshing area between the comb teeth and the steps of the escalator is captured by the time-frequency dynamic distribution feature, and the matching degree between the abnormal mode and the known foreign object voiceprint template is confirmed by the global voiceprint characterization feature, which can effectively improve the accuracy of detecting and identifying foreign objects stuck in the meshing area between the comb teeth and the steps.
[0109] For step S5, in this embodiment, please refer to Figure 11 , Figure 11 is the flowchart of the method for obtaining the foreign object voiceprint matching degree in an escalator comb foreign object detection and warning method of the present application. Based on the voiceprint fusion feature and a preset foreign object sound matching mechanism, the foreign object voiceprint matching degree is obtained, including:
[0110] S51: Obtain a number of preset foreign object voiceprint templates;
[0111] S52: Match the voiceprint fusion features with each of the foreign object voiceprint templates respectively to obtain corresponding voiceprint matching degrees;
[0112] S53: Select the maximum of the voiceprint matching degrees as the foreign object voiceprint matching degree.
[0113] For step S51, constructing the foreign object voiceprint template includes: obtaining the standard annotated voiceprints of foreign objects made of different materials such as copper, nickel, and aluminum hitting at different angles such as perpendicular, 30 degrees, and 60 degrees, etc., extracting the standard voiceprint features from the annotated voiceprints based on the steps described in S41 - S45 above. Each combination of a foreign object of any material and an impact angle is uniquely associated with the standard voiceprint feature, thereby constructing the foreign object voiceprint template.
[0114] For steps S52 and S53, compare the voiceprint fusion features with the standard voiceprint features in each of the foreign object voiceprint templates to obtain each corresponding voiceprint matching degree, and select the maximum of the voiceprint matching degrees as the foreign object voiceprint matching degree.
[0115] For step S6, the visual confidence weight and the voiceprint matching weight are 0.6 and 0.4 respectively. In one embodiment, dynamically adjusting the voiceprint matching weight includes: dynamically adjusting the voiceprint matching degree weight according to the difference between the signal - to - noise ratio of the audio signal and the signal - to - noise ratio threshold. Specifically, if the signal - to - noise ratio of the audio signal is lower than 30 db, increase the voiceprint matching degree weight to 0.7.
[0116] Dynamically adjusting the visual confidence weight includes: calculating the current light intensity by obtaining the grayscale histogram of the camera module; if the light intensity is less than a preset first light threshold, then increase the visual confidence weight; if the light intensity is greater than a preset second light threshold, then decrease the visual confidence weight or keep it unchanged. Specifically, the first light threshold is 50, and the second light threshold is 200. If the light intensity is less than 50, then increase the visual confidence weight to 0.8; if the light intensity is greater than 200, then the visual confidence weight remains unchanged at 0.6 or decreases to 0.4.
[0117] For steps S7 and S8, the rules for classifying the foreign object detection warning levels are as follows: Multiply the foreign object visual confidence by the visual confidence weight to obtain the visual foreign object detection score. Multiply the foreign object voiceprint matching degree by the voiceprint matching degree weight to obtain the voiceprint foreign object detection score. If the visual foreign object detection score is greater than or equal to the preset visual foreign object detection threshold, or if the voiceprint foreign object detection score is greater than or equal to the preset voiceprint foreign object detection threshold, then classify the escalator comb plate area as a first-level warning level; if the comprehensive foreign object detection score is greater than or equal to the comprehensive foreign object detection threshold, then classify the escalator comb plate area as a second-level warning level.
[0118] The warning instruction includes the generation time, warning level, warning suggestions, etc. The escalator warning device includes the red LED, buzzer, PLC, and safety relay module of the escalator. When the warning level is a first-level warning, the corresponding warning suggestion is to control the red LED of the escalator to flash at a frequency of 2 Hz, and at the same time control the buzzer to emit a 1 kHz sound, lasting for 1 second each time, with an interval of 2 seconds; when the warning level is a second-level warning, the corresponding warning suggestion is to send a Modbus-TCP instruction to the PLC and send an emergency stop instruction (function code 0x06, address 0x0001) to trigger the safety relay module to cut off the power supply of the escalator motor (standby channel) and store the current 10s data to the local SSD.
[0119] In other embodiments, the warning instruction can be adaptively modified according to the actual foreign object detection requirements.
[0120] Embodiment 2
[0121] Please refer to Figure 12 , Figure 12 which is a schematic diagram of an escalator comb foreign object detection and warning system of the present application. The present application also provides an escalator comb foreign object detection and warning system, which is characterized by including:
[0122] Comb plate area preprocessing image acquisition module 1: used to acquire the detection image of the escalator comb plate area and preprocess the detection image to obtain the preprocessed image of the comb plate area;
[0123] Comb plate area preprocessing voiceprint acquisition module 2: used to acquire the audio signal of the escalator comb plate area and preprocess the audio signal to obtain the preprocessed voiceprint of the comb plate area;
[0124] Foreign object visual confidence acquisition module 3: used to identify foreign objects in the preprocessed image of the comb plate area based on a preset visual foreign object detection model to obtain the foreign object visual confidence;
[0125] Voiceprint fusion feature acquisition module 4: used to perform multi-branch audio feature extraction on the preprocessed voiceprint in the comb plate area to obtain voiceprint fusion features;
[0126] Foreign object voiceprint matching degree acquisition module 5: used to obtain the foreign object voiceprint matching degree based on the preset foreign object sound matching mechanism according to the voiceprint fusion features;
[0127] Foreign object comprehensive detection score acquisition module 6: used to perform weighted fusion on the foreign object visual confidence and the foreign object voiceprint matching degree respectively based on the preset visual confidence weight and voiceprint matching degree weight to obtain the foreign object comprehensive detection score;
[0128] Early warning level division module 7: used to divide the early warning level of the escalator comb plate area according to the foreign object comprehensive detection score based on the preset foreign object detection early warning level division rules;
[0129] Early warning instruction transmission module 8: used to generate an early warning instruction corresponding to the early warning level according to the early warning level division result of the escalator comb plate area and transmit the early warning instruction to the escalator early warning device.
[0130] It should be noted that the data obtained when an escalator comb foreign object detection and early warning system provided by this application implements an escalator comb foreign object detection and early warning method are stored in the storage of this system one by one. When relevant calculations are required, the data required for the calculation can be directly obtained from the storage corresponding to it.
[0131] It should also be noted that when an escalator comb foreign object detection and early warning system provided in the above embodiment implements an escalator comb foreign object detection and early warning method, only the above-mentioned division of each functional module is used as an example for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, an escalator comb foreign object detection and early warning system provided in the above embodiment and an escalator comb foreign object detection and early warning method in Embodiment 1 belong to the same concept, and the implementation process is detailed in the method embodiment, which will not be repeated here.
[0132] Based on the same inventive concept, this application also provides an electronic device, which can be a server, a desktop computing device or a mobile computing device (for example, a laptop computing device, a handheld computing device, a tablet computer, a netbook, etc.) and other terminal devices. The device includes one or more processors and a memory. The processor is used to execute a program to implement the above-mentioned escalator comb foreign object detection and early warning method; the memory is used to store a computer program executable by the processor.
[0133] This application may be in the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain program code. The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the escalator comb foreign object detection and warning method. Computer-usable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include but are not limited to: phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.
[0134] This application is not limited to the above embodiments. If various changes or deformations to this application do not depart from the spirit and scope of this application, and if these changes and deformations fall within the scope of the claims of this application and equivalent technical scope, then this application also intends to include these changes and deformations.
Claims
1. An escalator comb foreign body detection and early warning method, characterized in that: Including: Obtain a detection image of the escalator comb plate area, and preprocess the detection image to obtain a preprocessed image of the comb plate area; Obtain an audio signal of the escalator comb plate area, and preprocess the audio signal to obtain a preprocessed voiceprint of the comb plate area; Based on a preset visual foreign object detection model, perform foreign object recognition on the preprocessed image of the comb plate area to obtain a foreign object visual confidence level; Extract multi-branch audio features from the preprocessed voiceprint of the comb plate area to obtain a voiceprint fusion feature; According to the voiceprint fusion feature, based on a preset foreign object sound matching mechanism, obtain a foreign object voiceprint matching degree; Based on a preset visual confidence level weight and a voiceprint matching degree weight, perform weighted fusion on the foreign object visual confidence level and the foreign object voiceprint matching degree respectively to obtain a comprehensive foreign object detection score; According to the comprehensive foreign object detection score, based on a preset foreign object detection warning level division rule, divide the warning level of the escalator comb plate area; According to the warning level division result of the escalator comb plate area, generate a warning instruction corresponding to the warning level, and transmit the warning instruction to the escalator warning device.
2. The escalator comb foreign body detection and early warning method according to claim 1 is characterized in that: The visual foreign object detection model includes a backbone network, a neck network, and a detection head integrated with an attention module; The performing foreign object recognition on the preprocessed image of the comb plate area based on a preset visual foreign object detection model to obtain a foreign object visual confidence level includes: Perform feature extraction on the preprocessed image of the comb plate area based on several convolutional layers of the backbone network to obtain several large-scale feature maps and small-scale feature maps; Perform downsampling operations on the large-scale feature maps and upsampling operations on the small-scale feature maps respectively through the neck network to obtain downsampled feature maps and upsampled feature maps, and perform convolutional fusion on the downsampled feature maps and the upsampled feature maps to obtain a fused feature map; Perform channel weighting on the fused feature map through the detection head to obtain a weighted feature map, and generate a foreign object recognition result according to the weighted feature map, where the foreign object recognition result includes: foreign object bounding box coordinates, foreign object category information, and the foreign object visual confidence level.
3. The escalator comb foreign body detection and early warning method according to claim 2 is characterized in that: The performing channel weighting on the fused feature map by the detection head to obtain a weighted feature map includes: Perform global averaging on the fused feature map according to the following formula to generate a channel global feature vector: where z c is the channel global feature vector under the c-th channel, H and W are the spatial dimensions of the fused feature map, and u c (i,j) is the eigenvalue of the c-th channel in the fused feature map at the coordinate (i,j); According to the channel global feature vector, learn channel weights through a fully connected network according to the following formula to generate a channel attention vector formula: s c = σ(W2 × δ(W1 × z c )) where s c is the channel attention vector under the c-th channel, σ is the ReLU activation function, δ is the Sigmoid activation function, W1 is the first weight matrix of the first fully connected network, and W2 is the second weight matrix of the second fully connected network; Multiply the channel attention vector and the fused feature map in corresponding channels to obtain the weighted feature map.
4. The escalator comb foreign body detection and early warning method according to claim 1 is characterized in that: The extracting multi-branch audio features from the preprocessed voiceprint of the comb plate area to obtain a voiceprint fusion feature includes: Segment the continuous preprocessed voiceprint of the comb plate area into several short-time frames, and calculate the spectral features of the short-time frames; Extract several frequency points with the highest energy from the spectral features of each short-time frame to form a time-frequency energy distribution feature of the corresponding dimension; Perform time-frequency dynamic analysis on the spectral features of each short-time frame through a filter bank to obtain a time-frequency dynamic distribution feature; Nonlinear feature extraction is performed on the preprocessed voiceprint in the comb plate area through a deep residual feature extraction network to obtain a global voiceprint characterization feature; The time-frequency energy distribution feature, the time-frequency dynamic distribution feature, and the global voiceprint characterization feature are concatenated and then dimensionality reduction operation is performed to obtain the voiceprint fusion feature.
5. The escalator comb foreign body detection and early warning method according to claim 2 is characterized in that: The obtaining of the foreign object voiceprint matching degree based on the preset foreign object sound matching mechanism according to the voiceprint fusion feature includes: Obtaining a number of preset foreign object voiceprint templates; Matching the voiceprint fusion feature with each of the foreign object voiceprint templates respectively to obtain corresponding voiceprint matching degrees; Selecting the maximum of the voiceprint matching degrees as the foreign object voiceprint matching degree.
6. The escalator comb foreign body detection and early warning method according to claim 1 is characterized in that: The obtaining of the detection image of the escalator comb plate area and the preprocessing of the detection image to obtain the preprocessed image of the comb plate area includes: Obtaining the detection image captured by a camera module with a field of view angle capable of covering the meshing area of the comb and the step of the escalator at a preset exposure time; Based on the detection image, determining the area coordinates of the area including the meshing area of the comb and the step in the escalator, and cropping the detection image to obtain a target area detection image; Performing a normalization operation on the target area detection image to obtain the preprocessed image of the comb plate area.
7. The escalator comb foreign body detection and early warning method according to claim 6 is characterized in that: The dynamic adjustment of the exposure time includes: Obtaining the step movement speed of the escalator; According to the step movement speed, dynamically adjusting the exposure time based on motion compensation according to the following formula to obtain the adjusted exposure time: Where T exp is the adjusted exposure time, v is the stepped movement speed, and p is the preset pixel size.
8. The escalator comb foreign body detection and early warning method according to claim 1 is characterized in that: The obtaining of the audio signal in the escalator comb plate area and the preprocessing of the audio signal to obtain the preprocessed voiceprint in the comb plate area includes: Obtaining the audio signal collected by a microphone module forming a microphone array in the escalator comb plate area; Calculating the signal-to-noise ratio of the audio signal, and if the signal-to-noise ratio is lower than a preset signal-to-noise ratio threshold, performing amplification and / or attenuation processing on the audio signal to obtain an optimized audio signal; Performing noise cancellation processing on the optimized audio signal based on a preset noise fingerprint library to obtain the preprocessed voiceprint in the comb plate area.
9. An escalator comb foreign body detection and warning system, characterized in that: Including: Preprocessed image acquisition module for the comb plate area: used to obtain the detection image of the escalator comb plate area and preprocess the detection image to obtain the preprocessed image of the comb plate area; Preprocessed voiceprint acquisition module for the comb plate area: used to obtain the audio signal in the escalator comb plate area and preprocess the audio signal to obtain the preprocessed voiceprint in the comb plate area; Foreign object visual confidence acquisition module: used to perform foreign object recognition on the preprocessed image of the comb plate area based on a preset visual foreign object detection model to obtain the foreign object visual confidence; Voiceprint fusion feature acquisition module: used to perform multi-branch audio feature extraction on the preprocessed voiceprint in the comb plate area to obtain the voiceprint fusion feature; Foreign object voiceprint matching degree acquisition module: used to obtain the foreign object voiceprint matching degree based on the preset foreign object sound matching mechanism according to the voiceprint fusion feature; Foreign object comprehensive detection score acquisition module: It is used to perform weighted fusion on the foreign object visual confidence and foreign object voiceprint matching degree respectively based on the preset visual confidence weight and voiceprint matching degree weight, so as to obtain the foreign object comprehensive detection score; Early warning level division module: It is used to divide the early warning level of the escalator comb plate area according to the foreign object comprehensive detection score based on the preset foreign object detection early warning level division rule; Early warning instruction transmission module: It is used to generate an early warning instruction corresponding to the early warning level according to the early warning level division result of the escalator comb plate area, and transmit the early warning instruction to the escalator early warning device.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the escalator comb foreign object detection and early warning method according to any one of claims 1-8.