A voiceprint recognition method and system for elevator operating state monitoring
By constructing a sound identification spectrum and comparing the sound differences under elevator operating conditions, the problems of insufficient adaptability and sensitivity of existing elevator monitoring methods have been solved, achieving higher monitoring accuracy and system stability, reducing false alarm rate, and ensuring elevator safety and maintenance efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- GUANGZHOU DAYIN ZHIYUAN DIGITAL TECH CO LTD
- Filing Date
- 2025-01-15
- Publication Date
- 2026-07-24
AI Technical Summary
Existing elevator operation status monitoring methods are insufficient in terms of adaptability and sensitivity. They are difficult to adapt to changes in sound under different operating conditions and are easily affected by environmental noise, resulting in a high false alarm rate and reducing the reliability and practicality of the system.
By collecting sound samples of the elevator under different operating conditions, selecting the stable period as the benchmark sound segment, preprocessing to remove noise interference and enhance sound characteristics, constructing a sound identification spectrum, and comparing the newly collected sound with the identification spectrum to record the differences, the elevator status is assessed to determine whether it deviates from the normal range, and corresponding maintenance or alarm signals are triggered.
This improved the accuracy and reliability of elevator operation status monitoring, reduced the false alarm rate, enhanced the system's noise resistance and stability, and ensured the safe operation and maintenance efficiency of the elevator.
Smart Images

Figure CN120039731B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of voiceprint recognition technology, specifically relating to a voiceprint recognition method and system for monitoring elevator operation status. Background Technology
[0002] In modern buildings, elevators serve as crucial vertical transportation tools, and their safety and reliability are paramount. To ensure safe elevator operation, real-time monitoring of their operational status is becoming increasingly important. Traditional elevator monitoring methods primarily rely on mechanical sensors and electronic monitoring systems, which can detect changes in physical parameters such as speed, position, and vibration. However, for some non-contact anomalies, such as incomplete door closure, internal component wear, or insufficient lubrication, traditional methods often struggle to detect them in a timely manner.
[0003] In recent years, with the development of voiceprint recognition technology, a new method based on sound feature analysis has been introduced into the field of elevator operation status monitoring. Existing technologies typically use fixedly installed sound sensors to collect the sounds generated during elevator operation and determine the presence of anomalies based on pre-set thresholds. While this method can monitor elevator status to some extent, it has the following limitations:
[0004] Poor adaptability: The sound emitted by elevators varies significantly under different working conditions, and existing methods mostly rely on fixed voiceprint feature libraries, which are difficult to adapt to the sound changes of elevators under various working conditions.
[0005] Insufficient sensitivity: Due to the lack of an effective noise filtering mechanism, existing methods are easily affected by environmental noise, resulting in a high false alarm rate and reducing the reliability and practicality of the system.
[0006] Therefore, how to establish a robust voiceprint feature model to adapt to the sound changes under different elevator operating conditions, thereby improving the accuracy and reliability of the monitoring system, is an urgent problem to be solved. Summary of the Invention
[0007] The purpose of this invention is to provide a voiceprint recognition method and system for monitoring elevator operation status, which not only improves the accuracy and reliability of the monitoring system, but also significantly enhances the system's noise resistance and stability, and reduces the false alarm rate, thereby solving the problems mentioned in the background art.
[0008] To achieve the above objectives, this invention proposes a voiceprint recognition method for monitoring elevator operation status, comprising:
[0009] Sound samples generated by the elevator under various working conditions are collected, and a stable period is selected from the collected sound samples as a reference sound segment.
[0010] The reference sound segment is preprocessed to remove noise interference and enhance sound features. The processed reference sound segment is then transformed and its inherent frequency components are extracted.
[0011] A sound identification spectrum is constructed using the inherent frequency components. The newly acquired elevator sound is compared with the constructed sound identification spectrum, and the differences are recorded.
[0012] Based on the differences mentioned, assess whether the current state of the elevator deviates from the normal range, and trigger the corresponding maintenance or alarm signal according to the assessment results.
[0013] Preferably, the sound samples collected by the elevator during various operating conditions include:
[0014] When the elevator is in different working conditions, a sensor array is used to capture sound signals and record the corresponding environmental variables. The average amplitude and frequency distribution of the captured sound signals are calculated.
[0015] Based on the calculated average amplitude and frequency distribution, a voiceprint vector is constructed for each working condition, and the degree of state change is quantified by comparing the differences between adjacent working conditions.
[0016] Preferably, selecting a stable period as a reference sound segment from the collected sound samples includes:
[0017] In the acquired sound samples, identify and mark the stable segments in each sample to ensure that the amplitude fluctuations within the stable segments do not exceed a preset threshold.
[0018] For each marked stable segment, calculate the energy of the sound signal within that segment, and select the stable segment with the highest energy value as the reference sound segment;
[0019] The selected reference sound segment is standardized so that the maximum amplitude is normalized to a unit value.
[0020] Preferably, the preprocessing of the reference sound segment to remove noise interference and enhance sound features includes:
[0021] For a selected reference sound segment, it is converted to the frequency domain, and low-energy regions in the spectrum are identified and labeled;
[0022] Design a filter function to reduce the impact of noise, apply the filter function to the frequency domain representation to achieve noise reduction and feature enhancement, and then convert it back to the time domain to obtain the preprocessed reference sound segment.
[0023] Preferably, the step of transforming the processed reference sound segment and extracting its inherent frequency components includes:
[0024] Generate a time-frequency representation for the preprocessed reference audio segment;
[0025] Based on the time-frequency representation, identify and record significant frequency points, construct a frequency distribution histogram, and count the number of times each frequency occurs;
[0026] The most frequently occurring frequency is selected as the inherent frequency component to ensure that the frequency can reflect the core sound characteristics of the elevator under normal operating conditions.
[0027] Preferably, the step of constructing a sound identifier spectrum using the inherent frequency components includes:
[0028] Based on the determined inherent frequency components, a frequency-time matrix is constructed, and the importance of different frequencies is evaluated based on the constructed frequency-time matrix.
[0029] The core features representing the sound of normal elevator operation are selected from the frequency-time matrix, and a sound identification spectrum is drawn to ensure that the spectrum clearly shows the sound characteristics of the elevator during normal operation.
[0030] Preferably, the step of comparing the newly acquired elevator sound with the constructed sound identifier spectrum and recording the differences includes:
[0031] The newly collected elevator sound samples are preprocessed to generate new preprocessed sound segments;
[0032] Based on the obtained preprocessed new sound segment, its time-frequency representation is calculated to ensure that it has the same time and frequency resolution as the sound identifier map.
[0033] Based on the obtained time-frequency representation, a frequency-time matrix of the new sound segment is constructed. The frequency-time matrix is compared element-by-element with the sound identifier map to generate visual elements. The difference is calculated, and all frequency points exceeding the set threshold are recorded as difference points.
[0034] Preferably, the assessment of whether the current state of the elevator deviates from the normal range based on the difference points includes:
[0035] Based on the generated visual elements, extract the color intensity and position coordinates of each element;
[0036] Based on the extracted information, weights are assigned according to the importance of different areas of the elevator to form a weighted color-position matrix.
[0037] Based on the resulting weighted color-position matrix, all visual elements are integrated to create a multi-layered overlay image, which is then combined with a timeline to generate a dynamic composite view that displays the real-time status of various parts of the elevator as time changes.
[0038] Preferably, triggering the corresponding maintenance or alarm signal based on the evaluation result includes:
[0039] Receive the dynamic composite view, analyze the color changes and positional shifts within it, and form a change vector;
[0040] Based on the obtained change vector, the difference between the change vector and the preset threshold is compared to determine whether there is a state that exceeds the normal range;
[0041] If an anomaly is confirmed, a priority index is calculated based on the degree of anomaly, and a response action is selected based on the calculated priority index.
[0042] If the priority index exceeds the critical value, an alarm signal is triggered; otherwise, a maintenance notification is issued.
[0043] On the other hand, the present invention proposes a voiceprint recognition system for monitoring elevator operation status, comprising:
[0044] The sound sample acquisition and reference segment selection module is used to acquire sound samples generated by the elevator during various working conditions, and select a stable period from the acquired sound samples as a reference sound segment.
[0045] The preprocessing and intrinsic frequency component extraction module is used to preprocess the reference sound segment, remove noise interference and enhance sound features, transform the processed reference sound segment and extract intrinsic frequency components.
[0046] The sound identifier spectrum construction and comparison module is used to construct a sound identifier spectrum using the inherent frequency components, compare the newly acquired elevator sound with the constructed sound identifier spectrum, and record the differences.
[0047] The status assessment and response triggering module is used to assess whether the current status of the elevator deviates from the normal range based on the difference points, and trigger the corresponding maintenance or alarm signals according to the assessment results.
[0048] Technical effects and advantages of the present invention: The voiceprint recognition method and system for monitoring elevator operation status proposed in this invention have the following advantages compared with the prior art:
[0049] This invention dynamically adjusts the voiceprint feature model to adapt to sound changes under different operating conditions, effectively removing noise interference and enhancing sound characteristics. The method constructs a sound identification spectrum reflecting the sound characteristics of the elevator during normal operation. By comparing newly acquired sounds with the identification spectrum and recording differences, the method assesses whether the elevator's status deviates from the normal range based on these differences, and ultimately triggers maintenance or alarm signals based on the assessment results. This method not only improves the accuracy and reliability of the monitoring system but also significantly enhances the system's noise immunity and stability, reduces the false alarm rate, and enables rapid response to abnormal states, ensuring the safe operation of the elevator and improving maintenance efficiency and safety. Attached Figure Description
[0050] Figure 1 This is a flowchart of the voiceprint recognition method for monitoring elevator operation status according to the present invention;
[0051] Figure 2 This is a block diagram of the voiceprint recognition system for monitoring elevator operation status according to the present invention. Detailed Implementation
[0052] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0053] This invention provides a voiceprint recognition method for monitoring elevator operation status. By dynamically adjusting the voiceprint feature model to adapt to sound changes under different operating conditions, it effectively removes noise interference and enhances sound characteristics. The method constructs a sound identifier spectrum reflecting the sound characteristics of the elevator during normal operation. Differences are recorded by comparing newly acquired sounds with the identifier spectrum, and these differences are used to assess whether the elevator status deviates from the normal range. Finally, maintenance or alarm signals are triggered based on the assessment results. Details are as follows:
[0054] like Figure 1 As shown, a voiceprint recognition method for monitoring elevator operation status according to the present invention includes the following steps:
[0055] Step 1: Collect sound samples generated by the elevator under various operating conditions; this further includes the following sub-processes:
[0056] When the elevator is in different operating conditions, a sensor array is used to capture sound signals and record the corresponding environmental variables. By using multiple sensors (sensor array) to capture sound signals, it is ensured that all sounds that may be generated during elevator operation are covered. Simultaneously recording environmental variables (such as temperature and humidity) helps to eliminate the influence of environmental factors on the sound data and improves the accuracy of subsequent analysis.
[0057] Based on the captured sound signals, the average amplitude A and frequency distribution F of each signal are calculated using the formulas: A = sum(abs(s(t))) / N and F = fft(s(t)) / length(s), where s(t) represents the sound signal sample at time t, and N is the total number of samples; the average amplitude A represents the overall intensity of the sound, while the frequency distribution F reflects the spectral characteristics of the sound. These two parameters together describe the main features of the sound, providing the basic data for subsequent construction of the speaker vector. The Fast Fourier Transform (FFT) is used to convert the time-domain signal into a frequency-domain representation, thereby allowing for clearer identification of the frequency components in the sound.
[0058] Based on the obtained A and F values, a voiceprint vector V is constructed for each working condition. This vector consists of a series of discrete points, represented as: V = [A1, F1, A2, F2, ..., An, Fn], where n represents the number of samples. The voiceprint vector V combines the average amplitude and frequency distribution to form a data structure that comprehensively describes the sound characteristics of the elevator under its working conditions. This multi-dimensional representation makes the sound differences between different working conditions more obvious, facilitating subsequent comparative analysis.
[0059] The degree of state change is quantified by comparing the difference D between adjacent operating conditions. The formula is: D = sqrt(sum((V_current - V_reference).^2)), where V_current is the voiceprint vector under the current operating condition, and V_reference is the voiceprint vector under the selected reference operating condition. The Euclidean distance D is calculated to measure the difference between the two voiceprint vectors, reflecting the degree of change from one operating condition to another. A smaller D value means that the two conditions are similar, while a larger D indicates a significant change. This step is crucial for detecting subtle changes in elevator state and helps to identify potential problems in a timely manner.
[0060] For example, suppose there is an elevator system that collects sound signals under three conditions: normal operation (condition 1), incomplete door closure (condition 2), and internal parts wear (condition 3). Five samples are collected for each condition, i.e., n = 5.
[0061] For each sample s(t), first calculate its average amplitude A and frequency distribution F.
[0062] Then, these values are combined into a voiceprint vector V, for example, for condition 1, V_1 = [A1,F1,A2,F2,...,A5,F5].
[0063] Next, using condition 1 as a reference, the difference D between the voiceprint vectors under other conditions and this condition is calculated. If the D value for condition 2 or condition 3 is significantly higher than a certain preset threshold, it indicates that the elevator's state has deviated and further inspection is required.
[0064] This method allows for effective monitoring of different elevator operating conditions and enables the identification of anomalies at an early stage, allowing for preventative maintenance measures to be taken.
[0065] Step two involves selecting a stable period from the sound samples obtained in step one as a reference sound segment; this further includes the following sub-processes:
[0066] In the sound samples obtained in step one, stable segments within each sample are identified and marked, where the amplitude fluctuation within a segment does not exceed a preset threshold T_v. By setting the amplitude fluctuation threshold T_v, sound segments indicating that the elevator is in a relatively stable operating state can be filtered out. This step ensures that the basic data for subsequent analysis is representative, eliminates noise interference caused by unstable factors (such as start-up, shutdown, or abnormal vibration), and improves the reliability of the monitoring results.
[0067] For each marked stable segment, the energy E_s of the sound signal within that segment is calculated using the formula: E_s = sum(s(t).^2), where s(t) represents the sound signal sample at time t. Energy E_s is an important indicator of sound intensity, reflecting the cumulative effect of the sound signal throughout the segment. By quantifying the energy of each stable segment, the differences in intensity between different segments can be further evaluated, helping to select the most representative benchmark sound segment. The principle of this formula is based on the definition of energy, that is, the sum of squares of the signal represents its total energy.
[0068] Selecting the stable segment with the highest energy value as the reference sound segment S_ref ensures that S_ref represents the sound characteristics of the elevator under ideal operating conditions. Choosing the stable segment with the highest energy as the reference sound segment ensures that the selected segment contains the most typical sound characteristics of the elevator during normal operation. High energy usually means a stronger sound signal, which helps to improve the sensitivity and accuracy of subsequent analysis, making the reference sound segment more reflective of the elevator's ideal operating state.
[0069] The selected benchmark audio segment S_ref is normalized to a maximum amplitude A_max of 1, using the formula: s_normalized(t) = s_ref(t) / A_max. Normalization eliminates the influence of amplitude differences under varying acquisition conditions, ensuring all audio segments are compared at the same scale. A normalized maximum amplitude of 1 guarantees consistency and comparability in subsequent analyses and also helps improve the algorithm's ability to distinguish different sound features.
[0070] For example, suppose we have collected several sound samples from different elevator operating conditions and have already completed the voiceprint vector construction in the first step. Now, in the second step, we use these samples to select a stable reference sound segment:
[0071] Identifying stable segments: First, set the amplitude fluctuation threshold T_v = 0.1. Then, iterate through all time periods in each sample, find the continuous time periods with amplitude fluctuations less than T_v, and mark them as stable segments.
[0072] Calculating the energy E_s: For each segment marked as stable, the energy of the sound signal within that segment is calculated using the formula E_s = sum(s(t).^2). For example, for a stable segment in the first sample, if it contains sound signal samples at time point t of [0.5, 0.6, 0.7, 0.8], then the calculated energy E_s is:
[0073] E_s=(0.5^2+0.6^2+0.7^2+0.8^2)=1.74
[0074] Selecting the reference audio segment S_ref: Compare the energy values of all stable segments and select the one with the highest energy as the reference audio segment S_ref. For example, if there is one segment with the highest energy value among all stable segments, reaching E_s_max = 2.5, then this segment becomes the reference audio segment.
[0075] Standardization Processing: Finally, the reference audio segment is standardized. Assuming the maximum amplitude A_max of the segment is 0.9, the original amplitude s_ref(t) at each time point t in the segment is normalized using the formula s_normalized(t) = s_ref(t) / A_max. For example, for the original amplitude s_ref(t) = [0.5, 0.6, 0.7, 0.8], the normalized result is:
[0076] s_normalized(t)=[0.5 / 0.9,0.6 / 0.9,0.7 / 0.9,0.8 / 0.9]≈
[0077] [0.556, 0.667, 0.778, 0.889].
[0078] Through the above steps, a benchmark sound segment representing the sound characteristics of an elevator under ideal working conditions was successfully selected and standardized, providing a reliable foundation for subsequent voiceprint recognition and status monitoring.
[0079] Step 3 involves preprocessing the baseline audio segment to remove noise interference and enhance audio features; this further includes the following sub-processes:
[0080] For the selected reference sound segment S_ref, apply the Fourier transform (FT) to convert it to the frequency domain to obtain the frequency domain representation F_S_ref = FT(S_ref);
[0081] Identify and label low-energy regions L_f in the spectrum, where the energy E_f of the region is below a set threshold T_E, using the formula: L_f={f|E_f(f)} <T_E};
[0082] Design a filter function G_f that takes a value close to 0 in the L_f region and a value close to 1 in other regions;
[0083] The filter function G_f is applied to the frequency domain representation F_S_ref, and noise reduction and feature enhancement are achieved through multiplication operations to generate the optimized frequency domain representation F_optimized = F_S_ref * G_f. Then, the inverse Fourier transform (IFT) is used to transform it back to the time domain to obtain the preprocessed reference sound segment S_optimized = IFT(F_optimized).
[0084] Step four, based on the reference sound segment processed in step three, extracts its inherent frequency components through mathematical transformation; this further includes the following sub-processes:
[0085] For the preprocessed reference audio segment S_optimized, a Short-Time Fourier Transform (STFT) is applied to generate a time-frequency representation T_F_S = STFT(S_optimized). Converting the time-domain signal to a frequency-domain representation via Fourier transform allows for clearer observation and analysis of the frequency components in the signal. This step makes subsequent spectrum analysis and filtering operations more intuitive and effective, facilitating the accurate identification and processing of noise within different frequency ranges.
[0086] Based on the obtained time-frequency representation T_F_S, a set of significant frequency points P_f is identified and recorded. The intensity I_f of these points in the time-frequency graph exceeds a threshold T_I, as defined by the formula: P_f = {(t,f)|I_f(t,f)>T_I}. Low-energy regions typically represent noise or unimportant background signals. By setting an energy threshold T_E, these regions can be accurately identified, providing a basis for subsequent filtering. The principle behind this formula is based on the energy distribution of each frequency point in the spectrum, selecting frequency bands with significantly lower energy than the sound characteristics under normal operating conditions as the targets for filtering.
[0087] A frequency distribution histogram H_f is constructed, and the frequency N_f of each frequency is counted to quantify the importance of different frequencies. The formula is: H_f(f) = sum(P_f(t,f)). The filter function G_f is designed to selectively attenuate the influence of low-energy regions (i.e., noise) while preserving important sound features in the high-frequency range. This filter can effectively separate the useful and useless parts of the signal, improving the quality of the sound signal. The filter design is based on an understanding of the spectral characteristics, ensuring that noise interference is minimized without affecting key features.
[0088] The top M most frequent frequencies are selected as the intrinsic frequency components F_intrinsic to ensure that the frequencies reflect the core sound characteristics under normal elevator operation. The formula is: F_intrinsic = topM(H_f), where topM represents the frequencies corresponding to the M highest entries in the histogram. A filter function is applied to the frequency domain representation through multiplication, achieving selective filtering of the original signal. The optimized frequency domain representation F_optimized is more focused on the key frequency components of the elevator's operating state, reducing interference from irrelevant noise. Finally, the optimized frequency domain signal is converted back to the time domain using inverse Fourier transform to obtain the preprocessed reference sound segment S_optimized. This step ensures that the final output sound segment is both clean and rich in feature information, facilitating subsequent voiceprint recognition and status monitoring.
[0089] For example, assuming a stable reference sound segment S_ref has been selected from different elevator operating conditions, it will now be preprocessed to remove noise interference and enhance sound features:
[0090] Applying Fourier Transform: First, apply Fast Fourier Transform (FFT) to the reference sound segment S_ref to transform it into the frequency domain, obtaining the frequency domain representation F_S_ref = FT(S_ref). For example, if the sound signal samples at time point t contained in S_ref are [0.556, 0.667, 0.778, 0.889], then the calculated frequency domain representation F_S_ref will be a series of complex numbers, representing the amplitude and phase information at different frequencies.
[0091] Identifying low-energy regions: Next, set the energy threshold T_E = 0.1, traverse each frequency point f in the frequency domain representation F_S_ref, calculate its energy E_f(f), and compare it with T_E. For example, for a certain frequency point f, if its energy E_f(f) = 0.08, then mark it as a low-energy region L_f. The specific formula is:
[0092] L_f = {f|E_f(f)} <T_E};
[0093] Design a filter function: Based on the identified low-energy region L_f, design a filter function G_f. Within the L_f region, G_f takes a value close to 0; in other regions, G_f takes a value close to 1. For example, if L_f contains frequencies f1 and f2, then G_f(f1)≈0 and G_f(f2)≈0, while G_f(f3)≈1 and G_f(f4)≈1 for frequencies outside the L_f region.
[0094] Denoising and Feature Enhancement: The filter function G_f is applied to the frequency domain representation F_S_ref. Denoising and feature enhancement are achieved through multiplication, generating an optimized frequency domain representation F_optimized = F_S_ref * G_f. For example, if the value of F_S_ref at a certain frequency point f is 0.5 + 0.3i, and G_f(f) = 0.9, then F_optimized(f) = (0.5 + 0.3i) * 0.9 = 0.45 + 0.27i. Finally, the optimized frequency domain representation is converted back to the time domain using the inverse Fourier transform (IFT), resulting in the preprocessed reference sound segment S_optimized = IFT(F_optimized). For example, the time domain signal obtained after the inverse transform of F_optimized might be [0.56, 0.67, 0.78, 0.89], representing a clean and feature-enhanced sound segment after preprocessing.
[0095] Through the above steps, the baseline audio segment was successfully preprocessed, effectively removing noise interference and enhancing audio features, providing high-quality basic data for subsequent voiceprint recognition and status monitoring.
[0096] Step 5: Construct a sound identification spectrum using the inherent frequency components obtained in Step 4. This spectrum reflects the sound characteristics of the elevator during normal operation. This further includes the following sub-processes:
[0097] Based on the determined intrinsic frequency component F_intrinsic, a frequency-time matrix M_f_t is constructed, where each element m_ft(t,f) represents the sound intensity at frequency f at time t, as defined by the formula: m_ft(t,f) = I_ft(t,f), where I_ft(t,f) represents the intensity at time t and frequency f in the time-frequency diagram. The frequency-time matrix M_f_t provides a two-dimensional representation, combining the time and frequency dimensions, allowing for a direct observation of the changes in sound intensity at different frequencies during elevator operation. This step lays the foundation for subsequent evaluation of the importance of each frequency.
[0098] Based on the constructed frequency-time matrix M_f_t, the average intensity A_f of each frequency f over all time periods is calculated to assess the importance of different frequencies. The formula is: A_f(f) = sum(m_ft(t,f)) / T_t, where T_t is the total time length; the average intensity A_f reflects the overall performance of a frequency throughout the entire monitoring period. By calculating the average intensity of each frequency, its importance in the normal operation of the elevator can be quantified, helping to identify representative frequency components. The principle behind this formula is to average the intensity over the time dimension to obtain a stable intensity value, thereby better understanding the behavior of the frequency throughout the entire time period.
[0099] Based on the calculated average intensity A_f, a set of frequency points P_high with intensity higher than a preset threshold T_A is selected. These points constitute the core characteristics of the elevator's normal operating sound. The formula is: P_high = {f|A_f(f)>T_A}. By setting the threshold T_A, frequency points that exhibit significant intensity during normal elevator operation can be selected from a large number of frequencies. These frequency points typically represent key characteristics of the elevator's operating state, facilitating subsequent analysis and comparison, ensuring that the focus is on the most representative and stable sound characteristics.
[0100] A sound signature map, S_map, is generated, visualizing the location and relative intensity of each frequency point. The formula is: S_map(f) = normalize(A_f(f)), where normalize normalizes the intensity to the [0,1] interval, ensuring the map clearly displays the sound characteristics of the elevator during normal operation. Normalization allows intensity values at different frequencies to be compared on the same scale, avoiding visual bias caused by large differences in the original intensity values. The resulting sound signature map, S_map, not only visually displays the sound characteristics of the elevator during normal operation but also facilitates subsequent comparative analysis with newly collected data, improving the accuracy and reliability of the monitoring system.
[0101] For example, assuming that the intrinsic frequency components F_intrinsic have been extracted from different operating conditions of the elevator, these frequency components will now be used to construct a sound identifier map:
[0102] Constructing the frequency-time matrix: First, construct the frequency-time matrix M_f_t based on the determined intrinsic frequency components F_intrinsic. For example, if F_intrinsic contains frequencies f1, f2, f3, and their intensities at times t1, t2, and t3 are [0.5, 0.6, 0.7] respectively, then the constructed matrix elements are:
[0103] m_ft(t1,f1)=0.5,m_ft(t2,f1)=0.6,m_ft(t3,f1)=0.7
[0104] Calculate the average intensity A_f: Next, calculate the average intensity A_f of each frequency over all time periods. Assuming a total time length T_t = 3 (i.e., three time points), the average intensity A_f(f1) for frequency f1 is calculated as follows:
[0105] A_f(f1)=(0.5+0.6+0.7) / 3=0.6
[0106] Filtering the core feature frequency set P_high: Set a threshold T_A = 0.5, and filter out the frequency set P_high whose average intensity is higher than T_A. For example, if the average intensity of frequency f1 is A_f(f1) = 0.6, and the average intensity of frequency f2 is A_f(f2) = 0.4, then P_high contains frequency f1.
[0107] P_high={f1}
[0108] Finally, the sound identifier map S_map is drawn, visualizing the position and relative intensity of each frequency point. The normalized intensity values are:
[0109] S_map(f1)=normalize(0.6)=0.6 / max(A_f)=0.6 / 0.6=1
[0110] Assuming max(A_f) is the maximum average intensity across all frequencies, then S_map(f1) is normalized to 1. For other frequency points, similar processing and normalization to the [0,1] interval are performed.
[0111] Through the above steps, a sound signature map S_map reflecting the sound characteristics of the elevator during normal operation was successfully constructed, providing an important reference for subsequent status monitoring and anomaly detection. This map not only visually displays the intensity distribution of key frequency points but also facilitates comparison with newly acquired data, improving the accuracy and reliability of the monitoring system.
[0112] Step six involves comparing the newly collected elevator sounds with the sound identification map established in step five and recording the differences; this further includes the following sub-processes:
[0113] In the newly acquired elevator sound samples, the same preprocessing method as in step three is applied to generate a new preprocessed sound segment S_new. Preprocessing ensures that the new sound segment shares the same characteristics as the data previously used to construct the sound identifier map, making the comparison between the two more accurate and meaningful. Preprocessing may include noise reduction, standardization, and other operations to improve the quality of subsequent analysis.
[0114] Based on the preprocessed new sound segment S_new, its time-frequency representation T_F_S_new = STFT(S_new) is calculated using Short Time Fourier Transform (STFT) to ensure that it has the same time and frequency resolution as the sound identification map S_map in step five. Through STFT, the signal in the time domain can be converted into a representation in both time and frequency dimensions. This helps to understand the signal characteristics more deeply and ensures the consistency between the new sound segment and the time-frequency map S_map in time and frequency, thus enabling effective comparison.
[0115] Based on the obtained time-frequency representation T_F_S_new, a frequency-time matrix M_f_t_new is constructed for the new sound segment, where m_new(t,f) represents the sound intensity at frequency f at time t, and the formula is: m_new(t,f) = I_new(t,f), where I_new(t,f) is the intensity of the new sample in the time-frequency map. Constructing the frequency-time matrix M_f_t_new is to create a structured dataset that can be directly compared with the sound identifier map S_map. This allows for comparison of old and new data within the same framework, identifying any significant changes or anomalies.
[0116] The constructed frequency-time matrix M_f_t_new is compared element-wise with the sound identifier map S_map established in step five to generate visual elements. The difference degree D_f is calculated using the formula: D_f(f) = abs(m_new(t,f) - S_map(f)). All frequency points exceeding the set threshold T_D are recorded as the difference point set P_diff. Through element-wise comparison, the difference between the new sound segment and the sound identifier map under normal operating conditions can be quantified. The difference degree D_f represents the absolute difference between the new and old data at each frequency point, while the set threshold T_D helps to determine which differences are significant and require further attention.
[0117] For example, assuming the preliminary work has been completed according to the steps above, step six will now be implemented:
[0118] Preprocessing new sound segments: First, the newly acquired elevator sound samples are preprocessed to generate new sound segments S_new. For example, background noise is removed to make the sound segments more suitable for subsequent analysis.
[0119] Calculate the time-frequency representation T_F_S_new: Then, use STFT to calculate the time-frequency representation T_F_S_new of the preprocessed new sound segment. Assuming S_new contains three time points [t1,t2,t3] and three frequency points [f1,f2,f3], then T_F_S_new will reflect the changes in sound intensity at these time and frequency points.
[0120] Constructing the frequency-time matrix M_f_t_new: Next, based on T_F_S_new, construct the frequency-time matrix M_f_t_new for the new sound segment. For example, if the intensity of frequency f1 in T_F_S_new at time point t1 is 0.8, then the matrix elements are:
[0121] m_new(t1,f1)=0.8
[0122] Comparison and recording of differences: Finally, M_f_t_new is compared element-by-element with the sound identifier map S_map, and the difference degree D_f is calculated. For example, if the difference degree of frequency f1 at time point t1 is:
[0123] D_f(f1)=abs(0.8-S_map(f1))=abs(0.8-1)=0.2
[0124] Assuming the threshold T_D = 0.1 is set, then since D_f(f1) > T_D, the frequency point f1 is recorded as part of the difference point set P_diff.
[0125] By following the steps described above, the differences between newly acquired elevator sounds and the sound signature patterns under normal operating conditions can be effectively identified, thus enabling timely detection and location of potential problems. This method not only improves the accuracy of fault diagnosis but also provides important decision support for maintenance personnel.
[0126] Step 7: Based on the discrepancies recorded in Step 6, assess whether the current state of the elevator deviates from the normal range; this further includes the following sub-processes:
[0127] Based on the visual elements generated in step six, the color intensity C_v and position coordinates P_v of each element are extracted to construct a color-position matrix M_C_P_v, with the formula: M_C_P_v(p) = C_v(p), where p represents the position coordinates. The color-position matrix M_C_P_v provides an intuitive representation that combines the location of sound difference points with color intensity. This step allows for a clearer visualization of the temporal and spatial distribution of different frequency components, laying the foundation for further analysis.
[0128] Based on the constructed color-position matrix M_C_P_v, weights W_v are assigned according to the importance of different areas of the elevator, forming a weighted color-position matrix M_W_C_P_v, with the formula: M_W_C_P_v(p) = M_C_P_v(p) * W_v(p). This ensures that visual elements in important areas are emphasized. By introducing weights W_v, changes in sound characteristics in key areas of the elevator can be highlighted. This weighted processing helps identify abnormal situations that significantly affect the elevator's operating status, improving the sensitivity and targeting of the monitoring system. The principle behind this formula is to multiply the color intensity of each position by its corresponding weight using a multiplication operation, thereby adjusting its importance.
[0129] Based on the weighted color-position matrix M_W_C_P_v, all visual elements are integrated to create a multi-layer overlay map L_stack_v, where each layer represents the state of a specific parameter or region. The formula is: L_stack_v = sum(M_W_C_P_layer_v), where M_W_C_P_layer_v is the weighted matrix corresponding to the specific parameter or region. The multi-layer overlay map L_stack_v integrates the state information of different parameters or regions into a comprehensive view, facilitating a complete understanding of the elevator's operating status. Each overlay map reflects the sound characteristics within a specific frequency or time period, allowing for the comparison of information from multiple dimensions in the same view, enhancing the interpretability and practicality of the data.
[0130] By utilizing the generated multi-layered overlay image L_stack_v and combining it with the time axis T_axis_v, a dynamic comprehensive view D_view_v is generated. This view displays the real-time status of various parts of the elevator as it changes over time. The formula is: D_view_v(t) = L_stack_v + T_axis_v(t), where t represents a point in time, ensuring that the view can intuitively reflect the elevator's operating status. The dynamic comprehensive view D_view_v, by introducing the time axis T_axis_v, transforms the static multi-layered overlay image into a dynamic image that changes over time, realistically reproducing the state evolution of various parts during elevator operation. This method not only improves the real-time performance of monitoring but also provides important decision support for fault diagnosis and preventative maintenance.
[0131] For example, assuming step six has been completed and a series of visual elements showing differences have been obtained, step seven will now be performed:
[0132] Constructing the color-position matrix: First, extract the color intensity C_v and position coordinates P_v of each element from the visual elements generated in step six, and construct the color-position matrix M_C_P_v. For example, if a difference point is located at coordinates (x1, y1) and has a color intensity of 0.8, then the matrix elements are:
[0133] M_C_P_v(x1,y1)=0.8
[0134] Forming a weighted color-position matrix: Next, weights W_v are assigned based on the importance of different areas of the elevator, forming a weighted color-position matrix M_W_C_P_v. For example, if the importance weight of the elevator door area is 2, while that of other areas is 1, then for the difference point (x1, y1) located in the door area, its weighted color intensity is:
[0135] M_W_C_P_v(x1,y1)=0.8*2=1.6
[0136] Creating a multi-layer overlay image: Then, based on the resulting weighted color-position matrix M_W_C_P_v, all visual elements are integrated to create a multi-layer overlay image L_stack_v. Assuming M_W_C_P_v contains three layers, each corresponding to a different frequency or time period, the overlay image is calculated as follows:
[0137] L_stack_v=sum(M_W_C_P_layer_v)
[0138] Each layer, M_W_C_P_layer_v, represents a weighted matrix for a specific frequency or time period.
[0139] Generating the dynamic composite view: Finally, combining the time axis T_axis_v, a dynamic composite view D_view_v is generated. Assuming the time axis T_axis_v contains three time points [t1, t2, t3], the dynamic composite view at each time point t is as follows:
[0140] D_view_v(t)=L_stack_v+T_axis_v(t)
[0141] For example, at time point t1, the dynamic composite view is as follows:
[0142] D_view_v(t1)=L_stack_v+T_axis_v(t1)
[0143] Through the above steps, the recorded discrepancies successfully assessed whether the elevator's current state deviated from the normal range. The dynamic integrated view D_view_v not only visually displays the real-time status of each part of the elevator but also provides monitoring personnel with detailed reference information, helping them to promptly identify and resolve potential problems. This method significantly improves the safety and reliability of elevator operation.
[0144] Step eight, based on the evaluation results of step seven, triggers the corresponding maintenance or alarm signals to ensure the safe operation of the elevator; this further includes the following sub-processes:
[0145] The dynamic integrated view D_view_v(t) generated in step seven is received, and the color change ΔC_v and position change ΔP_v are analyzed to form a change vector V_change_v = [ΔC_v, ΔP_v], which represents the changes in the state of each part of the elevator. The change vector V_change_v provides a quantitative indicator for evaluating the changes in the state of each part of the elevator. By analyzing the changes in color and position, the location and extent of abnormal situations can be accurately located, providing a basis for subsequent decision-making.
[0146] Based on the obtained change vector V_change_v, the difference between it and preset thresholds T_C_v and T_P_v is compared to determine whether there are states exceeding the normal range. The formula used here is: Flag_v = (abs(ΔC_v)>T_C_v) OR (abs(ΔP_v)>T_P_v). When Flag_v is true, it indicates an abnormal situation; this formula is used to determine whether elevator components are malfunctioning. When Flag_v is true, it means that an abnormal change that may affect the safe operation of the elevator has been detected. The principle of this formula is to evaluate two conditions using the logical operator OR. If either condition is met (i.e., the color change exceeds the threshold or the position change exceeds the threshold), an abnormal situation is considered to exist.
[0147] Upon confirming an anomaly, a priority index I_priority_v is calculated based on the severity of the anomaly, using the formula: I_priority_v = log(abs(ΔC_v) + abs(ΔP_v) + 1). This index reflects the urgency of the situation. By introducing a logarithmic function, the effects of color changes and positional shifts are combined, ensuring a positive value even when both are zero (due to the addition of 1). This step helps the system determine the appropriate level of response, thereby improving efficiency.
[0148] Based on the calculated priority index I_priority_v, an appropriate response action A_response_v is selected. If I_priority_v exceeds the critical value L_critical_v, an alarm signal S_alert_v is triggered; otherwise, a maintenance notification N_maintenance_v is issued, following the rule: I_priority_v > L_critical_v; THENA_response_v = S_alert_v; ELSEA_response_v = N_maintenance_v. This response mechanism ensures that problems of varying severity are handled promptly and appropriately. For urgent issues, an alarm is immediately triggered to alert relevant personnel to take immediate action; for non-urgent issues, routine maintenance is scheduled to avoid unnecessary panic and waste of resources.
[0149] For example, assuming the dynamic composite view D_view_v(t) generated in step seven has been obtained, step eight will now be implemented:
[0150] Analyzing the dynamic composite view: From D_view_v(t), the color change ΔC_v = 0.5 and the position change ΔP_v = 0.3 at a certain time point t are obtained, forming a change vector:
[0151] V_change_v = [0.5, 0.3]
[0152] Determine if any conditions exceed the normal range: Set the color change threshold T_C_v = 0.4 and the position change threshold T_P_v = 0.2, and apply the formula:
[0153] Flag_v=(abs(0.5)>0.4)OR(abs(0.3)>0.2)
[0154] The result is Flag_v = True, indicating that there is an abnormal situation.
[0155] Calculate the priority index: Based on the abnormal situation, calculate the priority index:
[0156] I_priority_v=log(abs(0.5)+abs(0.3)+1)
[0157] =log(1.8)
[0158] ≈0.6107
[0159] Select an appropriate response action: Set the critical value L_critical_v = 0.5, and select the response action based on the value of I_priority_v:
[0160] IF0.6107>0.5THENA_response_v=S_alert_vELSEA_response_v=N_maintenance_v
[0161] Because I_priority_v is greater than L_critical_v, the alarm signal S_alert_v is selected to be triggered.
[0162] Through the above steps, the corresponding maintenance or alarm signals were successfully triggered based on the assessment results, ensuring the safe operation of the elevator. This method not only improves the speed and accuracy of fault response but also enhances the system's automation level and reduces the need for manual intervention.
[0163] On the other hand, this invention proposes a voiceprint recognition system for monitoring elevator operating status, such as... Figure 2 As shown, it includes:
[0164] The sound sample acquisition and reference segment selection module is used to acquire sound samples generated by the elevator during various working conditions, and select a stable period from the acquired sound samples as a reference sound segment.
[0165] The preprocessing and intrinsic frequency component extraction module is used to preprocess the reference sound segment, remove noise interference and enhance sound features, transform the processed reference sound segment and extract intrinsic frequency components.
[0166] The sound identifier spectrum construction and comparison module is used to construct a sound identifier spectrum using the inherent frequency components, compare the newly acquired elevator sound with the constructed sound identifier spectrum, and record the differences.
[0167] The status assessment and response triggering module is used to assess whether the current status of the elevator deviates from the normal range based on the difference points, and trigger the corresponding maintenance or alarm signals according to the assessment results.
[0168] In addition, the aforementioned sound sample acquisition and reference segment selection module, preprocessing and intrinsic frequency component extraction module, sound identifier spectrum construction and comparison module, and state assessment and response triggering module are also used to implement other steps of the aforementioned voiceprint recognition method for monitoring elevator operation status, which will not be elaborated here.
[0169] In summary, this invention effectively removes noise interference and enhances sound characteristics by dynamically adjusting the voiceprint feature model to adapt to sound changes under different operating conditions. The method constructs a sound identification spectrum reflecting the sound characteristics of the elevator during normal operation. By comparing newly acquired sounds with the identification spectrum and recording differences, the method assesses whether the elevator's status deviates from the normal range based on these differences, and finally triggers maintenance or alarm signals based on the assessment results. This method not only improves the accuracy and reliability of the monitoring system but also significantly enhances the system's noise resistance and stability, reduces the false alarm rate, and enables rapid response to abnormal states, ensuring the safe operation of the elevator and improving maintenance efficiency and safety.
[0170] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A voiceprint recognition method for monitoring elevator operation status, characterized in that, include: Sound samples generated by the elevator under various working conditions are collected, and a stable period is selected from the collected sound samples as a reference sound segment. The reference sound segment is preprocessed to remove noise interference and enhance sound features. The processed reference sound segment is then transformed and its inherent frequency components are extracted. A sound identification spectrum is constructed using the inherent frequency components. The newly acquired elevator sound is compared with the constructed sound identification spectrum, and the differences are recorded. Based on the differences mentioned, assess whether the current state of the elevator deviates from the normal range, and trigger the corresponding maintenance or alarm signal according to the assessment results; The process of collecting sound samples generated by the elevator under various working conditions includes: capturing sound signals using a sensor array and recording corresponding environmental variables when the elevator is in different working conditions; calculating the average amplitude and frequency distribution of the captured sound signals; constructing a sound signature vector for each working condition based on the calculated average amplitude and frequency distribution; and quantifying the degree of state change by comparing the differences between adjacent working conditions. The step of selecting a stable period as a reference sound segment from the acquired sound samples includes: identifying and marking stable segments in each sample in the acquired sound samples, ensuring that the amplitude fluctuation within the stable segments does not exceed a preset threshold; for each marked stable segment, calculating the energy of the sound signal within the segment, and selecting the stable segment with the highest energy value as the reference sound segment; and standardizing the selected reference sound segment so that the maximum amplitude is normalized to a unit value. The preprocessing of the reference sound segment to remove noise interference and enhance sound features includes: for the selected reference sound segment, converting it to the frequency domain, identifying and marking low-energy regions in the spectrum; designing a filter function to reduce the influence of noise, applying the filter function to the frequency domain representation to achieve noise reduction and feature enhancement, and then converting it back to the time domain to obtain the preprocessed reference sound segment. The process of transforming the processed reference sound segment and extracting its inherent frequency components includes: generating a time-frequency representation for the preprocessed reference sound segment; identifying and recording significant frequency points based on the time-frequency representation, constructing a frequency distribution histogram, and counting the number of times each frequency occurs; selecting the frequency that occurs most frequently as the inherent frequency component to ensure that the frequency can reflect the core sound characteristics under normal elevator operation. The construction of the sound identifier map using the inherent frequency components includes: Based on the determined inherent frequency components, a frequency-time matrix is constructed, and the importance of different frequencies is evaluated based on the constructed frequency-time matrix. The core features representing the sound of normal elevator operation are selected from the frequency-time matrix, and a sound identification spectrum is drawn to ensure that the spectrum clearly shows the sound characteristics of the elevator during normal operation. The process of comparing the newly collected elevator sounds with the constructed sound identifier map and recording the differences includes: The newly collected elevator sound samples are preprocessed to generate new preprocessed sound segments; Based on the obtained preprocessed new sound segment, its time-frequency representation is calculated to ensure that it has the same time and frequency resolution as the sound identifier map. Based on the obtained time-frequency representation, a frequency-time matrix of the new sound segment is constructed. The frequency-time matrix is compared element by element with the sound identifier map to generate visual elements. The difference is calculated, and all frequency points exceeding the set threshold are recorded as difference points. The assessment of whether the current state of the elevator deviates from the normal range based on the aforementioned differences includes: Based on the generated visual elements, extract the color intensity and position coordinates of each element; Based on the extracted information, weights are assigned according to the importance of different areas of the elevator to form a weighted color-position matrix. Based on the resulting weighted color-position matrix, all visual elements are integrated to create a multi-layered overlay map, which is then combined with a timeline to generate a dynamic composite view that displays the real-time status of various parts of the elevator as time changes.
2. The voiceprint recognition method for monitoring elevator operation status according to claim 1, characterized in that, Based on the evaluation results, the corresponding maintenance or alarm signals are triggered, including: Receive the dynamic composite view, analyze the color changes and positional shifts within it, and form a change vector; Based on the obtained change vector, the difference between the change vector and the preset threshold is compared to determine whether there is a state that exceeds the normal range; If an anomaly is confirmed, a priority index is calculated based on the degree of anomaly, and a response action is selected based on the calculated priority index. If the priority index exceeds the critical value, an alarm signal is triggered; otherwise, a maintenance notification is issued.
3. A voiceprint recognition system for monitoring elevator operating status, used to perform the method as described in any one of claims 1-2, characterized in that, include: The sound sample acquisition and reference segment selection module is used to collect sound samples generated by the elevator during various operating conditions, and select a stable period as a reference sound segment from the collected sound samples. Specifically, it includes: capturing sound signals using a sensor array and recording corresponding environmental variables when the elevator is in different operating conditions; calculating the average amplitude and frequency distribution of the captured sound signals; constructing a sound signature vector for each operating condition based on the calculated average amplitude and frequency distribution, and quantifying the degree of state change by comparing the differences between adjacent operating conditions; identifying and marking stable segments in each acquired sound sample to ensure that the amplitude fluctuation within the stable segment does not exceed a preset threshold; calculating the energy of the sound signal within each marked stable segment, and selecting the stable segment with the highest energy value as the reference sound segment; and standardizing the selected reference sound segment to normalize the maximum amplitude to a unit value. The preprocessing and intrinsic frequency component extraction module is used to preprocess the reference sound segment, remove noise interference and enhance sound features, transform the processed reference sound segment and extract intrinsic frequency components. Specifically, it includes: for the selected reference sound segment, converting it to the frequency domain, identifying and marking low-energy regions in the spectrum; designing a filter function to reduce noise influence, applying the filter function to the frequency domain representation to achieve noise reduction and feature enhancement, and then converting it back to the time domain to obtain the preprocessed reference sound segment; generating a time-frequency representation for the preprocessed reference sound segment; based on the time-frequency representation, identifying and recording significant frequency points, constructing a frequency distribution histogram, and counting the frequency occurrences of each frequency; selecting the frequency with the most occurrences as the intrinsic frequency component to ensure that the frequency can reflect the core sound characteristics under normal elevator operation. The sound identifier spectrum construction and comparison module is used to construct a sound identifier spectrum using the inherent frequency components, compare the newly acquired elevator sound with the constructed sound identifier spectrum, and record the differences. The status assessment and response triggering module is used to assess whether the current status of the elevator deviates from the normal range based on the difference points, and trigger the corresponding maintenance or alarm signals according to the assessment results.