Voiceprint recognition method and system for elevator running state monitoring
By dynamically adjusting the voiceprint feature model, removing noise interference, building a sound identification map, comparing the newly collected sound and identification map, and evaluating the elevator status, the problems of poor adaptability and insufficient sensitivity in the existing technology are solved, the accuracy and reliability of the monitoring system are improved, and the rapid response to abnormal states is achieved.
Patent Information
- Application Number
- CN202510059377.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-15
AI Technical Summary
The prior art has poor adaptability and insufficient sensitivity in the monitoring of elevator operating status, making it difficult to effectively monitor non-contact abnormalities, and has a high false alarm rate, which reduces the reliability and practicality of the system.
By collecting sound samples of elevators under different working conditions, selecting stable periods as reference sound segments, pre-processing to remove noise interference, extracting natural frequency components, building sound identification maps, comparing newly collected sound and identification maps, recording differences points, evaluating whether the elevator state deviates from the normal range, and triggering corresponding maintenance or alarm signals.
It improves the accuracy and reliability of the monitoring system, enhances noise resistance and stability, reduces the false alarm rate, realizes rapid response to abnormal conditions, and ensures the safe operation and maintenance efficiency of the elevator.
Smart Images

Figure CN120039731A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of voiceprint recognition, and particularly relates to a voiceprint recognition method and system for elevator operation state monitoring. Background Art
[0002] In modern buildings, elevators, as important vertical transportation tools, their safety and reliability are crucial. To ensure the safe operation of elevators, it has become increasingly important to monitor the operation state of elevators in real time. Traditional elevator operation state monitoring methods mainly rely on mechanical sensors and electronic monitoring systems, which can detect changes in physical parameters such as speed, position, and vibration. However, for some non-contact abnormal situations, such as incomplete door closure, internal part wear, or insufficient lubrication, traditional methods often have difficulty in detecting them in a timely manner.
[0003] In recent years, with the development of voiceprint recognition technology, a new method based on sound feature analysis has been introduced into the field of elevator operation state monitoring. In the prior art, a fixedly installed sound sensor is usually used to collect the sounds generated during elevator operation, and whether there is an abnormality is judged by a preset threshold. Although this method can achieve the monitoring of elevator status to a certain extent, it has the following limitations:
[0004] Poor adaptability: The sounds emitted by elevators under different working conditions will change significantly, and the existing methods mostly rely on a fixed voiceprint feature library, making it difficult to adapt to the sound changes of elevators under various working conditions.
[0005] Insufficient sensitivity: Due to the lack of an effective noise filtering mechanism, the existing methods are easily affected by environmental noise, resulting in a high false alarm rate and reducing the reliability and practicality of the system.
[0006] Therefore, how to establish a robust voiceprint feature model to adapt to the sound changes of elevators under different working conditions, so as to improve the accuracy and reliability of the monitoring system is an urgent problem to be solved. Summary of the Invention
[0007] The purpose of the present invention is to provide a voiceprint recognition method and system for elevator operation state monitoring, which not only improves the accuracy and reliability of the monitoring system, but also significantly enhances the anti-noise ability and stability of the system, reduces the false alarm rate, so as to solve the problems proposed in the above background art.
[0008] To achieve the above purpose, the present invention proposes a voiceprint recognition method for elevator operation state monitoring, including:
[0009] Collecting sound samples generated by the elevator during various working conditions, and selecting a stable period in the collected sound samples as a reference sound segment;
[0010] Preprocess the reference sound segment to remove noise interference and enhance sound features, transform the processed reference sound segment and extract the inherent frequency components;
[0011] Use the inherent frequency components to construct a sound identification map, compare the newly collected elevator sound with the constructed sound identification map, and record the difference points;
[0012] Evaluate whether the current state of the elevator deviates from the normal range based on the difference points, and trigger corresponding maintenance or alarm signals according to the evaluation results.
[0013] Preferably, collecting sound samples generated by the elevator during various working conditions includes:
[0014] When the elevator is in different working conditions, use a sensor array to capture sound signals, record the corresponding environmental variables, and calculate the average amplitude and frequency distribution of the captured sound signals;
[0015] According to the calculated average amplitude and frequency distribution, construct a voiceprint vector for each working condition, and quantify the degree of state change by comparing the differences between adjacent working conditions.
[0016] Preferably, selecting a stable period as the reference sound segment from the collected sound samples includes:
[0017] In the obtained sound samples, identify and mark the stable segments in each sample, ensuring that the amplitude fluctuation within the stable segment does not exceed a preset threshold;
[0018] For each marked stable segment, calculate the energy of the sound signal within the segment, and select the stable segment with the highest energy value as the reference sound segment;
[0019] Normalize the selected reference sound segment so that the maximum amplitude is normalized to a unit value.
[0020] Preferably, preprocessing the reference sound segment to remove noise interference and enhance sound features includes:
[0021] For the selected reference sound segment, transform it to the frequency domain, identify and mark the low-energy regions in the spectrum;
[0022] Design a filter function to reduce the influence of noise, apply the filter function to the frequency-domain representation to achieve noise reduction and feature enhancement, and then transform it back to the time domain to obtain the preprocessed reference sound segment.
[0023] Preferably, transforming the processed reference sound segment and extracting the inherent frequency components includes:
[0024] Generate a time-frequency representation for the preprocessed reference sound segment;
[0025] Based on the time-frequency representation, identify and record significant frequency points, construct a frequency distribution histogram, and count the number of occurrences of each frequency;
[0026] Select the frequency with the most occurrences as the inherent frequency component to ensure that the frequency can reflect the core sound characteristics under normal elevator operation.
[0027] Preferably, constructing a sound identification map using the inherent frequency component includes:
[0028] Based on the determined inherent frequency component, construct a frequency-time matrix, and evaluate the importance of different frequencies based on the constructed frequency-time matrix;
[0029] Filter out the core features representing the normal elevator operation sound in the frequency-time matrix, draw a sound identification map, and ensure that the map clearly shows the sound characteristics during normal elevator operation.
[0030] Preferably, comparing the newly collected elevator sound with the constructed sound identification map and recording the difference points includes:
[0031] Preprocess the newly collected elevator sound sample to generate a preprocessed new sound segment;
[0032] Based on the obtained preprocessed new sound segment, calculate its time-frequency representation to ensure the same time and frequency resolution as the sound identification map;
[0033] According to the obtained time-frequency representation, construct a frequency-time matrix for the new sound segment, compare the frequency-time matrix with the sound identification map element by element, generate visual elements, calculate the difference degree, and record all frequency points exceeding the set threshold as difference points.
[0034] Preferably, evaluating whether the current state of the elevator deviates from the normal range based on the difference points includes:
[0035] Based on the generated visual elements, extract the color intensity and position coordinates of each element;
[0036] Based on the extracted information, assign weights according to the importance of different elevator areas to form a weighted color-position matrix;
[0037] According to the formed weighted color-position matrix, integrate all visual elements, create a multi-layer superimposed map, and generate a dynamic comprehensive view in combination with the time axis. The dynamic comprehensive view shows the real-time state of each part of the elevator over time.
[0038] Preferably, triggering corresponding maintenance or alarm signals according to the evaluation results includes:
[0039] Receiving a dynamic comprehensive view, parsing color changes and position changes therein, and forming a change vector;
[0040] Based on the obtained change vector, by comparing the difference with a preset threshold, determining whether there is a state beyond the normal range;
[0041] In the case of confirming an abnormality, calculating a priority index according to the degree of abnormality, and based on the calculated priority index, selecting a response action;
[0042] If the priority index exceeds the critical value, trigger an alarm signal; otherwise, send a maintenance notice.
[0043] On the other hand, the present invention proposes a voiceprint recognition system for elevator operation state monitoring, including:
[0044] A sound sample collection and reference segment selection module, configured to collect sound samples generated by the elevator during various working conditions, and select a stable period from the collected sound samples as a reference sound segment;
[0045] A preprocessing and natural frequency component extraction module, configured to preprocess the reference sound segment, remove noise interference and enhance sound features, transform the processed reference sound segment, and extract natural frequency components;
[0046] A sound identification map construction and comparison module, configured to construct a sound identification map using the natural frequency components, compare the newly collected elevator sound with the constructed sound identification map, and record the difference points;
[0047] A state evaluation and response trigger module, configured to evaluate whether the current state of the elevator deviates from the normal range based on the difference points, and trigger corresponding maintenance or alarm signals according to the evaluation results.
[0048] The technical effects and advantages of the present invention: A voiceprint recognition method and system for elevator operation state monitoring proposed by the present invention have the following advantages compared with the prior art:
[0049] The present invention effectively removes noise interference and enhances sound features by dynamically adjusting the voiceprint feature model to adapt to sound changes under different working conditions. This method constructs a sound identification map reflecting the sound characteristics during normal elevator operation, compares the differences between newly collected sounds and the records in the identification map, evaluates whether the elevator state deviates from the normal range based on these differences, and finally triggers maintenance or alarm signals according to the evaluation results. This method not only improves the accuracy and reliability of the monitoring system, but also significantly enhances the anti-noise ability and stability of the system, reduces the false alarm rate, and enables a rapid response to abnormal states, ensuring the safe operation of the elevator and improving the maintenance efficiency and safety. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 is a flowchart of the voiceprint recognition method for elevator operation status monitoring according to the present invention;
[0051] Figure 2 is a block diagram of the voiceprint recognition system for elevator operation status monitoring according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. The specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0053] The present invention provides a voiceprint recognition method for elevator operation status monitoring, which effectively removes noise interference and enhances sound features by dynamically adjusting the voiceprint feature model to adapt to sound changes under different working conditions. This method constructs a sound identification map reflecting the sound characteristics during normal elevator operation, compares the differences between newly collected sounds and the records in the identification map, evaluates whether the elevator state deviates from the normal range based on these differences, and finally triggers maintenance or alarm signals according to the evaluation results. Specifically as follows:
[0054] As Figure 1 shown, a voiceprint recognition method for elevator operation status monitoring in the present invention includes the following steps:
[0055] Step 1, collect sound samples generated by the elevator during various working conditions; further including the following sub-processes:
[0056] When the elevator is in different working conditions, use a sensor array to capture sound signals and record the corresponding environmental variables. By using multiple sensors (sensor array) to capture sound signals, it is ensured that all sounds that may be generated during the elevator operation can be covered. At the same time, recording environmental variables (such as temperature, humidity, etc.) helps to exclude the influence of environmental factors on sound data and improve the accuracy of subsequent analysis.
[0057] Based on the captured sound signals, calculate the average amplitude A and frequency distribution F of each signal. The formulas are: A = sum(abs(s(t))) / N and F = fft(s(t)) / length(s), where s(t) represents the sound signal sample at time t and N is the total number of samples. The average amplitude A represents the overall intensity of the sound, while the frequency distribution F reflects the spectral characteristics of the sound. These two parameters jointly describe the main features of the sound and provide basic data for constructing the voiceprint vector later. The fast Fourier transform (FFT) is used to convert the time-domain signal into a frequency-domain representation, so that the frequency components in the sound can be more clearly identified.
[0058] According to the obtained values of A and F, construct the voiceprint vector V for each working condition. This vector is composed of a series of discrete points and is expressed as: V = [A1, F1, A2, F2,..., An, Fn], where n represents the number of samples. The voiceprint vector V combines the average amplitude and frequency distribution to form a data structure that comprehensively describes the sound characteristics under the elevator working state. This multi-dimensional representation method makes the sound differences between different working conditions more obvious and is convenient for subsequent comparative analysis.
[0059] Quantify the degree of state change by comparing the difference D between adjacent working conditions. The formula is: D = sqrt(sum((V_current - V_reference).^2)), where V_current is the voiceprint vector under the current working condition and V_reference is the voiceprint vector under the selected reference working condition. The calculation of the Euclidean distance D is used to measure the difference between two voiceprint vectors and reflects the degree of change from one working condition to another. A smaller D value means that the two conditions are similar, while a larger D indicates a significant change. This step is very important for detecting subtle changes in the elevator state and helps to discover potential problems in a timely manner.
[0060] Exemplarily, assume there is an elevator system, and sound signals are collected under three conditions: normal operation (condition 1), incomplete door closing (condition 2), and internal part wear (condition 3). 5 samples are collected under each condition, that is, n = 5.
[0061] For each sample s(t), first calculate its average amplitude A and frequency distribution F.
[0062] Then, these values are combined into a voiceprint vector V. For example, for working condition 1, V_1 = [A1, F1, A2, F2,..., A5, F5].
[0063] Next, select working condition 1 as the reference and calculate the difference D between the voiceprint vectors under other working conditions. If the D value of working condition 2 or working condition 3 is significantly higher than a preset threshold, it indicates that the state of the elevator has deviated and further inspection is required.
[0064] Through this method, different working conditions of the elevator can be effectively monitored, and abnormal conditions can be identified at an early stage, so as to take preventive maintenance measures.
[0065] Step 2: In the sound samples obtained in Step 1, select a stable period as the reference sound segment; it further includes the following sub-processes:
[0066] In the sound samples obtained in Step 1, identify and mark the stable paragraphs in each sample, where the amplitude fluctuation within the paragraph does not exceed the preset threshold T_v; by setting the threshold T_v of the amplitude fluctuation, the sound paragraphs when the elevator is in a relatively stable operating state can be screened out. This step ensures that the basic data for subsequent analysis is representative, excludes noise interference caused by non-stable factors (such as startup, stop, or abnormal vibration), and improves the reliability of the monitoring results.
[0067] For each marked stable paragraph, calculate the energy E_s of the sound signal within the paragraph. The formula is: E_s = sum(s(t).^2), where s(t) represents the sound signal sample at time t; the energy E_s is an important indicator to measure the sound intensity and reflects the cumulative effect of the sound signal throughout the paragraph. By quantifying the energy of each stable paragraph, the differences in intensity between different paragraphs can be further evaluated to help select the most representative reference sound segment. The principle of this formula is based on the definition of energy, that is, the sum of the squares of the signals represents its total energy.
[0068] Select the stable paragraph with the highest energy value as the reference sound segment S_ref to ensure that S_ref can represent the sound characteristics of the elevator under the ideal working condition; selecting the stable paragraph with the highest energy as the reference sound segment can ensure that the selected segment contains the most typical sound characteristics of the elevator during normal operation. High energy usually means a stronger sound signal, which helps to improve the sensitivity and accuracy of subsequent analysis and makes the reference sound segment better reflect the ideal working condition of the elevator.
[0069] Normalize the selected reference sound segment S_ref so that its maximum amplitude A_max is normalized to the unit value 1. The formula is: s_normalized(t) = s_ref(t) / A_max. The normalization process is to eliminate the influence of amplitude differences under different acquisition conditions, enabling all sound segments to be compared on the same scale. The normalized maximum amplitude is 1, ensuring the consistency and comparability of subsequent analyses, and also helping to improve the algorithm's ability to distinguish different sound features.
[0070] Exemplarily, assume that several sound samples are collected from different operating conditions of an elevator, and the construction of the voiceprint vector in the first step has been completed. Now enter the second step, and use these samples to select a stable reference sound segment:
[0071] Identify stable segments: First, set the amplitude fluctuation threshold T_v = 0.1, then traverse all time periods in each sample, find the continuous time periods with amplitude fluctuations less than T_v, and mark them as stable segments.
[0072] Calculate the energy E_s: For each segment marked as stable, use the formula E_s = sum(s(t).^2) to calculate the energy of the sound signal within this segment. For example, for a certain stable segment in the first sample, if the sound signal samples at time points t it contains are [0.5, 0.6, 0.7, 0.8], then the calculated energy E_s is:
[0073] E_s = (0.5^2 + 0.6^2 + 0.7^2 + 0.8^2) = 1.74
[0074] Select the reference sound segment S_ref: Compare the energy values of all stable segments, and select the one with the highest energy as the reference sound segment S_ref. For example, if among all stable segments, there is a segment with the maximum energy value reaching E_s_max = 2.5, then this segment becomes the reference sound segment.
[0075] Normalization process: Finally, normalize this reference sound segment. Assume the maximum amplitude A_max of this segment is 0.9, then for the original amplitude s_ref(t) at each time point t in the segment, normalize it according to the formula s_normalized(t) = s_ref(t) / A_max. For example, for the original amplitude s_ref(t) = [0.5, 0.6, 0.7, 0.8], the normalized result is:
[0076] s_normalized(t) = [0.5 / 0.9, 0.6 / 0.9, 0.7 / 0.9, 0.8 / 0.9] ≈
[0077] [0.556, 0.667, 0.778, 0.889].
[0078] Through the above steps, the reference sound segment representing the sound characteristics in the ideal working state of the elevator is successfully selected and standardized, providing a reliable basis for subsequent voiceprint recognition and condition monitoring.
[0079] Step 3: Preprocess the reference sound segment to remove noise interference and enhance sound features; further includes the following sub-processes:
[0080] For the selected reference sound segment S_ref, apply the Fourier transform FT to convert it to the frequency domain, obtaining the frequency-domain representation F_S_ref = FT(S_ref);
[0081] Identify and mark the low-energy regions L_f in the spectrum, where the energy E_f of the region is lower than the set threshold T_E, and the formula is: L_f = {f|E_f(f) < T_E};
[0082] Design a filter function G_f that takes values close to 0 in the L_f region and close to 1 in other regions;
[0083] Apply the filter function G_f to the frequency-domain representation F_S_ref, and achieve noise reduction and feature enhancement through multiplication operations, generating the optimized frequency-domain representation F_optimized = F_S_ref * G_f, and then use the inverse Fourier transform IFT to convert it back to the time domain to obtain the preprocessed reference sound segment S_optimized = IFT(F_optimized).
[0084] Step 4: Based on the reference sound segment processed in Step 3, extract its natural frequency components through mathematical transformation; further includes the following sub-processes:
[0085] Apply the short-time Fourier transform STFT to the preprocessed reference sound segment S_optimized to generate the time-frequency representation T_F_S = STFT(S_optimized); converting the time-domain signal to the frequency-domain representation through the Fourier transform can more clearly observe and analyze the frequency components in the signal. This step makes the subsequent spectrum analysis and filtering operations more intuitive and effective, and helps to accurately identify and process the noise in different frequency ranges.
[0086] Based on the obtained time-frequency representation \(T_F_S\), identify and record the set of significant frequency points \(P_f\), where the intensity \(I_f\) of a point in the time-frequency diagram exceeds the threshold \(T_I\). The formula is: \(P_f=\{(t,f)|I_f(t,f)>T_I\}\); low-energy regions typically represent noise or unimportant background signals. By setting the energy threshold \(T_E\), these regions can be accurately marked, providing a basis for subsequent filtering. The principle of this formula is based on the energy distribution of each frequency point in the spectrum, and those frequency segments with energy significantly lower than the sound characteristics in the normal operating state are selected as the objects to be filtered.
[0087] Construct a frequency distribution histogram \(H_f\) and count the number of occurrences \(N_f\) of each frequency to quantify the importance of different frequencies. The formula is: \(H_f(f)=\sum(P_f(t,f))\); the design purpose of the filter function \(G_f\) is to selectively weaken the influence of low-energy regions (i.e., noise) while retaining the important sound characteristics in the high-frequency band. This filter can effectively separate the useful part and the useless part in the signal, improving the quality of the sound signal. The design of the filter is based on the understanding of the spectral characteristics, ensuring that noise interference is minimized without affecting the key features.
[0088] Select the top \(M\) frequencies with the most occurrences as the intrinsic frequency components \(F_{intrinsic}\) to ensure that the frequencies can reflect the core sound characteristics of the elevator in the normal operating state. The formula is: \(F_{intrinsic}=topM(H_f)\), where \(topM\) represents selecting the frequencies corresponding to the highest \(M\) entries in the histogram. By applying the filter function to the frequency-domain representation through multiplication, selective filtering of the original signal is achieved. The optimized frequency-domain representation \(F_{optimized}\) is more concentrated on the key frequency components of the elevator operating state, reducing the interference of irrelevant noise. Finally, the optimized frequency-domain signal is converted back to the time domain through the inverse Fourier transform to obtain the preprocessed reference sound segment \(S_{optimized}\). This step ensures that the final output sound segment is both clean and rich in characteristic information, facilitating subsequent voiceprint recognition and status monitoring.
[0089] Exemplarily, assume that a stable reference sound segment \(S_{ref}\) has been selected from different operating conditions of the elevator, and now it will be preprocessed to remove noise interference and enhance the sound characteristics:
[0090] Applying Fourier Transform: First, apply the Fast Fourier Transform (FFT) to the reference sound segment S_ref to convert it to the frequency domain, obtaining the frequency domain representation F_S_ref = FT(S_ref). For example, if the sound signal samples at time points t contained in S_ref are [0.556, 0.667, 0.778, 0.889] respectively, the calculated frequency domain representation F_S_ref will be a series of complex numbers representing the amplitude and phase information at different frequencies.
[0091] Identifying Low-Energy Regions: Next, set the energy threshold T_E = 0.1, traverse each frequency point f in the frequency domain representation F_S_ref, calculate its energy E_f(f) and compare it with T_E. For example, for a certain frequency point f, if its energy E_f(f) = 0.08, then mark it as a low-energy region L_f. The specific formula is:
[0092] L_f = {f|E_f(f) < T_E};
[0093] Designing a Filter Function: According to the identified low-energy region L_f, design a filter function G_f. Within the L_f region, G_f takes values close to 0; in other regions, G_f takes values close to 1. For example, if L_f contains frequencies f1, f2, then G_f(f1) ≈ 0 and G_f(f2) ≈ 0, while G_f(f3) ≈ 1 and G_f(f4) ≈ 1 for frequency points in the non-L_f region.
[0094] Denoising and Feature Enhancement: Apply the filter function G_f to the frequency domain representation F_S_ref, and achieve denoising and feature enhancement through multiplication operations, generating the optimized frequency domain representation F_optimized = F_S_ref * G_f. For example, if the value of F_S_ref at a certain frequency point f is 0.5 + 0.3i, and G_f(f) = 0.9, then F_optimized(f) = (0.5 + 0.3i) * 0.9 = 0.45 + 0.27i. Finally, use the Inverse Fourier Transform IFT to convert the optimized frequency domain representation back to the time domain, obtaining the preprocessed reference sound segment S_optimized = IFT(F_optimized). For example, the time domain signal obtained after the inverse transformation of F_optimized may be [0.56, 0.67, 0.78, 0.89], representing a clean and feature-enhanced sound segment after preprocessing.
[0095] Through the above steps, the reference sound segment has been successfully preprocessed, effectively removing noise interference and enhancing the sound features, providing high-quality basic data for subsequent voiceprint recognition and status monitoring.
[0096] Step 5: Construct a sound signature map using the intrinsic frequency components obtained in Step 4. This map reflects the sound characteristics during normal elevator operation. It further includes the following sub - processes:
[0097] Based on the determined intrinsic frequency components F_intrinsic, construct a frequency - time matrix M_f_t, where each element m_ft(t, f) represents the sound intensity at frequency f at time t. The formula is: m_ft(t, f) = I_ft(t, f), where I_ft(t, f) represents the intensity at time t and frequency f in the time - frequency diagram. The frequency - time matrix M_f_t provides a two - dimensional representation that combines the two dimensions of time and frequency, enabling the intuitive observation of the variation in sound intensity at different frequencies during elevator operation. This step lays the foundation for subsequent evaluation of the importance of each frequency.
[0098] Based on the constructed frequency - time matrix M_f_t, calculate the average intensity A_f of each frequency f over all time periods to evaluate the importance of different frequencies. The formula is: A_f(f) = sum(m_ft(t, f)) / T_t, where T_t is the total time length. The average intensity A_f reflects the overall performance of a certain frequency over the entire monitoring time period. By calculating the average intensity of each frequency, its importance in normal elevator operation can be quantified, helping to screen out representative frequency components. The principle of this formula is to average the intensity over the time dimension to obtain a stable intensity value, thus better understanding the behavior of this frequency over the entire time period.
[0099] According to the calculated average intensity A_f, select the set of frequency points P_high with intensities higher than the preset threshold T_A. These points constitute the core characteristics of the sound during normal elevator operation. The formula is: P_high = {f|A_f(f)>T_A}. By setting the threshold T_A, those frequency points that exhibit significant intensity during normal elevator operation can be selected from among many frequencies. These frequency points usually represent the key characteristics of the elevator operation state, facilitating subsequent analysis and comparison to ensure that the most representative and stable sound characteristics are being focused on.
[0100] Draw the sound signature map S_map, which visualizes the position and relative intensity of each frequency point. The formula is: S_map(f) = normalize(A_f(f)), where normalize means normalizing the intensity to the interval [0, 1] to ensure that the map clearly shows the sound characteristics during normal elevator operation. The normalization process enables the intensity values of different frequency points to be compared on the same scale, avoiding visual deviations caused by overly large differences in the original intensities. The finally generated sound signature map S_map not only intuitively shows the sound characteristics during normal elevator operation but also facilitates subsequent comparison and analysis with newly collected data, improving the accuracy and reliability of the monitoring system.
[0101] Exemplarily, assume that the intrinsic frequency components F_intrinsic have been extracted from different operating conditions of the elevator. Now, these frequency components will be used to construct the sound signature map:
[0102] Construct the frequency-time matrix: First, according to the determined intrinsic frequency components F_intrinsic, construct the frequency-time matrix M_f_t. For example, if F_intrinsic contains frequencies f1, f2, f3, and the intensities at time points t1, t2, t3 are [0.5, 0.6, 0.7] respectively, then the elements of the constructed matrix are:
[0103] m_ft(t1, f1) = 0.5, m_ft(t2, f1) = 0.6, m_ft(t3, f1) = 0.7
[0104] Calculate the average intensity A_f: Next, calculate the average intensity A_f of each frequency over all time periods. Assume the total time length T_t = 3 (i.e., three time points), then the average intensity A_f(f1) for frequency f1 is calculated as follows:
[0105] A_f(f1) = (0.5 + 0.6 + 0.7) / 3 = 0.6
[0106] Screen the set of core characteristic frequency points P_high: Set the threshold T_A = 0.5 and screen out the set of frequency points P_high with an average intensity higher than T_A. For example, if the average intensity A_f(f1) of frequency f1 is 0.6 and the average intensity A_f(f2) of frequency f2 is 0.4, then P_high contains frequency f1:
[0107] P_high = {f1}
[0108] Draw the sound signature map S_map: Finally, draw the sound signature map S_map to visualize the position and relative intensity of each frequency point. The intensity values after normalization are:
[0109] S_map(f1) = normalize(0.6) = 0.6 / max(A_f) = 0.6 / 0.6 = 1
[0110] Assume that max(A_f) is the maximum average intensity of all frequencies. Then, after normalization, S_map(f1) equals 1. For other frequency points, similar processing is carried out and normalized to the interval [0, 1].
[0111] Through the above steps, the sound identification map S_map reflecting the sound characteristics during normal elevator operation is successfully constructed, providing an important reference basis for subsequent condition monitoring and anomaly detection. This map not only intuitively shows the intensity distribution of key frequency points but also facilitates comparison with newly collected data, improving the accuracy and reliability of the monitoring system.
[0112] Step 6: Compare the newly collected elevator sound with the sound identification map established in Step 5 and record the difference points; it further includes the following sub - processes:
[0113] In the newly collected elevator sound samples, apply the same pre - processing method as in Step 3 to generate a new pre - processed sound segment S_new; the pre - processing ensures that the new sound segment has the same characteristics as the data used to construct the sound identification map before, making the comparison between the two more accurate and meaningful. The pre - processing may include operations such as noise reduction and normalization to improve the quality of subsequent analysis.
[0114] Based on the obtained pre - processed new sound segment S_new, use the short - time Fourier transform STFT to calculate its time - frequency representation T_F_S_new = STFT(S_new), ensuring the same time and frequency resolution as the sound identification map S_map in Step 5; through STFT, the signal in the time domain can be converted into a representation in two dimensions of time and frequency, which helps to understand the signal characteristics more deeply and ensures the consistency of the new sound segment with the time - frequency map S_map in terms of time and frequency, enabling effective comparison.
[0115] According to the obtained time - frequency representation T_F_S_new, construct the frequency - time matrix M_f_t_new of the new sound segment, where m_new(t, f) represents the sound intensity at frequency f at time t, and the formula is: m_new(t, f) = I_new(t, f), where I_new(t, f) is the intensity of the new sample in the time - frequency diagram; constructing the frequency - time matrix M_f_t_new is to create a structured data set that can be directly compared with the sound identification map S_map. In this way, new and old data can be compared in the same framework to identify any significant changes or anomalies.
[0116] Compare the constructed frequency-time matrix \(M_{f_t}^{new}\) with the sound identification atlas \(S_{map}\) established in Step Five element by element to generate visual elements, calculate the difference degree \(D_f\), and the formula is: \(D_f(f)=|m_{new}(t,f)-S_{map}(f)|\). Record all frequency points exceeding the set threshold \(T_D\) as the difference point set \(P_{diff}\). By element-by-element comparison, the difference between the new sound segment and the sound identification atlas in the normal operating state can be quantified. The difference degree \(D_f\) represents the absolute difference between the old and new data at each frequency point, while the set threshold \(T_D\) helps to determine which differences are significant and require further attention.
[0117] Exemplarily, assume that the preliminary work has been completed according to the above steps, and now Step Six will be implemented:
[0118] Preprocess the new sound segment: First, preprocess the newly collected elevator sound samples to generate a new sound segment \(S_{new}\). For example, remove background noise to make the sound segment more suitable for subsequent analysis.
[0119] Calculate the time-frequency representation \(T_{F_S}^{new}\): Then, use the STFT to calculate the time-frequency representation \(T_{F_S}^{new}\) of the preprocessed new sound segment. Assume that \(S_{new}\) contains three time points \([t1, t2, t3]\) and three frequency points \([f1, f2, f3]\), then \(T_{F_S}^{new}\) will reflect the sound intensity changes at these time points and frequency points.
[0120] Construct the frequency-time matrix \(M_{f_t}^{new}\): Next, construct the frequency-time matrix \(M_{f_t}^{new}\) of the new sound segment based on \(T_{F_S}^{new}\). For example, if the intensity of frequency \(f1\) at time point \(t1\) in \(T_{F_S}^{new}\) is 0.8, then the matrix element is:
[0121] \(m_{new}(t1,f1)=0.8\)
[0122] Compare and record the difference points: Finally, compare \(M_{f_t}^{new}\) with the sound identification atlas \(S_{map}\) element by element and calculate the difference degree \(D_f\). For example, if the difference degree of frequency \(f1\) at time point \(t1\) is:
[0123] \(D_f(f1)=|0.8 - S_{map}(f1)|=|0.8 - 1| = 0.2\)
[0124] Assume that the set threshold \(T_D = 0.1\), then since \(D_f(f1)>T_D\), the frequency point \(f1\) is recorded as part of the difference point set \(P_{diff}\).
[0125] Through the above steps, the differences between the newly collected elevator sounds and the sound identification atlas in the normal operating state can be effectively identified, so as to timely detect and locate potential problems. This method not only improves the accuracy of fault diagnosis but also provides important decision-making support for maintenance personnel.
[0126] Step 7: Based on the difference points recorded in Step 6, evaluate whether the current state of the elevator deviates from the normal range; further includes the following sub-processes:
[0127] Based on the visual elements generated in Step 6, extract the color intensity C_v and position coordinates P_v of each element, and construct a color-position matrix M_C_P_v. The formula is: M_C_P_v(p) = C_v(p), where p represents the position coordinates; the color-position matrix M_C_P_v provides an intuitive representation method that combines the position and color intensity of the sound difference points. This step enables a clearer visualization of the distribution of different frequency components in time and space, providing a basis for further analysis.
[0128] Based on the constructed color-position matrix M_C_P_v, assign weights W_v according to the importance of different areas of the elevator to form a weighted color-position matrix M_W_C_P_v. The formula is: M_W_C_P_v(p) = M_C_P_v(p) * W_v(p), ensuring that the visual elements in important areas are emphasized; by introducing the weight W_v, the changes in the sound characteristics of key areas of the elevator can be highlighted. This weighted processing helps to identify abnormal conditions that have a greater impact on the elevator operation status, improving the sensitivity and pertinence of the monitoring system. The principle of this formula is to multiply the color intensity at each position by the corresponding weight through multiplication operations to adjust its importance.
[0129] According to the weighted color-position matrix M_W_C_P_v of the sub, integrate all visual elements to create a multi-layer stack diagram L_stack_v, where each layer represents the state of a specific parameter or area. The formula is: L_stack_v = sum(M_W_C_P_layer_v), where M_W_C_P_layer_v is the weighted matrix corresponding to a specific parameter or area; the multi-layer stack diagram L_stack_v integrates the state information of different parameters or areas into a comprehensive view, facilitating a comprehensive understanding of the elevator operation status. Each layer of the stack diagram reflects the sound characteristics within a specific frequency or time period, enabling the comparison of information in multiple dimensions in the same view and enhancing the interpretability and practicality of the data.
[0130] Using the formed multi-layer stacked graph L_stack_v and combining it with the time axis T_axis_v, a dynamic comprehensive view D_view_v is generated. This view shows the real-time status of each part of the elevator over time. The formula is: D_view_v(t) = L_stack_v + T_axis_v(t), where t represents the time point, ensuring that the view can intuitively reflect the operation of the elevator. By introducing the time axis T_axis_v, the dynamic comprehensive view D_view_v converts the static multi-layer stacked graph into a dynamic image that changes over time, truly reproducing the state evolution of each part during the elevator operation. This method not only improves the real-time monitoring but also provides important decision-making support for fault diagnosis and preventive maintenance.
[0131] Exemplarily, assume that step six has been completed and a series of visual elements of the difference points have been obtained. Now, step seven will be implemented:
[0132] Construct a color-position matrix: First, extract the color intensity C_v and position coordinates P_v of each element from the visual elements generated in step six to construct a color-position matrix M_C_P_v. For example, if a difference point is located at coordinates (x1, y1) and the color intensity is 0.8, then the matrix element is:
[0133] M_C_P_v(x1,y1) = 0.8
[0134] Form a weighted color-position matrix: Next, assign weights W_v according to the importance of different areas of the elevator to form a weighted color-position matrix M_W_C_P_v. For example, if the importance weight of the elevator door area is 2 and that of other areas is 1, then for the difference point (x1, y1) located in the door area, its weighted color intensity is:
[0135] M_W_C_P_v(x1,y1) = 0.8 * 2 = 1.6
[0136] Create a multi-layer stacked graph: Then, according to the formed weighted color-position matrix M_W_C_P_v, integrate all visual elements to create a multi-layer stacked graph L_stack_v. Assume that M_W_C_P_v contains three layers corresponding to different frequencies or time periods, then the calculation of the stacked graph is as follows:
[0137] L_stack_v = sum(M_W_C_P_layer_v)
[0138] Where each layer M_W_C_P_layer_v represents the weighted matrix for a specific frequency or time period.
[0139] Generate a dynamic comprehensive view: Finally, combine the time axis T_axis_v to generate a dynamic comprehensive view D_view_v. Assume that the time axis T_axis_v contains three time points [t1, t2, t3], then the dynamic comprehensive view at each time point t is:
[0140] D_view_v(t) = L_stack_v + T_axis_v(t)
[0141] For example, at time point t1, the dynamic comprehensive view is:
[0142] D_view_v(t1) = L_stack_v + T_axis_v(t1)
[0143] Through the above steps, the current state of the elevator is successfully evaluated based on the recorded difference points to determine whether it deviates from the normal range. The dynamic comprehensive view D_view_v not only intuitively displays the real-time state of each part of the elevator but also provides detailed reference information for the monitoring personnel to help them detect and solve potential problems in a timely manner. This method significantly improves the safety and reliability of elevator operation.
[0144] Step Eight, according to the evaluation result of Step Seven, trigger corresponding maintenance or alarm signals to ensure the safe operation of the elevator; further includes the following sub-processes:
[0145] Receive the dynamic comprehensive view D_view_v(t) generated in Step Seven, parse the color change ΔC_v and position change ΔP_v therein, and form a change vector V_change_v = [ΔC_v, ΔP_v] to represent the changes in the states of each part of the elevator; the change vector V_change_v provides a quantitative index for evaluating the changes in the states of each part of the elevator. By analyzing the changes in color and position, the location and degree of abnormal situations can be accurately located, providing a basis for subsequent decision-making.
[0146] Based on the obtained change vector V_change_v, by comparing the differences with the preset thresholds T_C_v and T_P_v, determine whether there is a state beyond the normal range. Here, the formula is used: Flag_v = (abs(ΔC_v) > T_C_v) OR (abs(ΔP_v) > T_P_v). When Flag_v is true, it indicates that there is an abnormal situation; this formula is used to determine whether there are abnormalities in elevator components. When Flag_v is true, it means that abnormal changes that may affect the safe operation of the elevator have been detected. The principle of this formula is to evaluate two conditions through the logical operator OR. As long as one condition is met (i.e., the color change exceeds the threshold or the position change exceeds the threshold), it is considered that there is an abnormal situation.
[0147] In the case of confirming the existence of an anomaly, calculate the priority index \(I_{priority\_v}\) according to the degree of anomaly. The formula is: \(I_{priority\_v}=\log(|\Delta C_v| + |\Delta P_v|+1)\). This index reflects the urgency that needs to be addressed; the priority index \(I_{priority\_v}\) is used to reflect the urgency that needs to be addressed. By introducing the logarithmic function, the impacts of color change and position change are combined, and it is ensured that a positive value can be obtained even when both are zero (due to adding 1). This step helps the system determine what level of response measures to take, thus improving the response efficiency.
[0148] Based on the calculated priority index \(I_{priority\_v}\), select an appropriate response action \(A_{response\_v}\). If \(I_{priority\_v}\) exceeds the critical value \(L_{critical\_v}\), trigger the alarm signal \(S_{alert\_v}\); otherwise, issue a maintenance notice \(N_{maintenance\_v}\), following the rule: IF \(I_{priority\_v}>L_{critical\_v}\) THEN \(A_{response\_v}=S_{alert\_v}\) ELSE \(A_{response\_v}=N_{maintenance\_v}\). This response mechanism ensures that problems of different severities can be handled promptly and appropriately. For urgent problems, immediately trigger an alarm to remind relevant personnel to take urgent actions; for non-urgent problems, arrange routine maintenance to avoid unnecessary panic and resource waste.
[0149] Exemplarily, assume that the dynamic comprehensive view \(D_{view\_v}(t)\) generated in step seven has been obtained. Now, step eight will be implemented:
[0150] Parse the dynamic comprehensive view: Parse the color change \(\Delta C_v = 0.5\) and position change \(\Delta P_v = 0.3\) at a certain time point \(t\) from \(D_{view\_v}(t)\) to form a change vector:
[0151] \(V_{change\_v}=[0.5, 0.3]\)
[0152] Determine whether there is a state beyond the normal range: Set the color change threshold \(T_{C_v}=0.4\) and the position change threshold \(T_{P_v}=0.2\), and apply the formula:
[0153] \(Flag_v=(|0.5|>0.4) OR (|0.3|>0.2)\)
[0154] The result is \(Flag_v = True\), indicating the existence of an abnormal situation.
[0155] Calculate the priority index: According to the abnormal situation, calculate the priority index:
[0156] I_priority_v = log(abs(0.5)+abs(0.3)+1)
[0157] = log(1.8)
[0158] ≈0.6107
[0159] Select an appropriate response action: Set the threshold value L_critical_v = 0.5, and select a response action based on the value of I_priority_v:
[0160] IF 0.6107 > 0.5 THEN A_response_v = S_alert_v ELSE A_response_v = N_maintenance_v
[0161] Since I_priority_v is greater than L_critical_v, the alarm signal S_alert_v is selected to be triggered.
[0162] Through the above steps, the corresponding maintenance or alarm signal is successfully triggered according to the evaluation results, ensuring the safe operation of the elevator. This method not only improves the speed and accuracy of fault response, but also enhances the automation level of the system and reduces the need for manual intervention.
[0163] On the other hand, the present invention proposes a voiceprint recognition system for elevator operation state monitoring, as Figure 2 shown, including:
[0164] A sound sample collection and reference segment selection module, which is used to collect sound samples generated by the elevator during various working conditions, and select a stable period from the collected sound samples as a reference sound segment;
[0165] A preprocessing and natural frequency component extraction module, which is used to preprocess the reference sound segment, remove noise interference and enhance sound features, transform the processed reference sound segment and extract natural frequency components;
[0166] A sound identification atlas construction and comparison module, which is used to construct a sound identification atlas using the natural frequency components, compare the newly collected elevator sound with the constructed sound identification atlas, and record the difference points;
[0167] A state evaluation and response trigger module, which is used to evaluate whether the current state of the elevator deviates from the normal range based on the difference points, and trigger the corresponding maintenance or alarm signal according to the evaluation results.
[0168] In addition, when the above sound sample collection and reference segment selection module, preprocessing and natural frequency component extraction module, sound identification map construction and comparison module, and status evaluation and response trigger module are executed, they are also used to implement other steps of the above-mentioned voiceprint recognition method for elevator operation status monitoring, which will not be elaborated one by one here.
[0169] In summary, the present invention dynamically adjusts the voiceprint feature model to adapt to the sound changes under different working conditions, effectively removes noise interference and enhances sound features. This method constructs a sound identification map that reflects the sound characteristics during normal elevator operation, compares the newly collected sounds with the recorded difference points in the identification map, and evaluates whether the elevator status deviates from the normal range based on these difference points. Finally, maintenance or alarm signals are triggered according to the evaluation results. This method not only improves the accuracy and reliability of the monitoring system, but also significantly enhances the anti-noise ability and stability of the system, reduces the false alarm rate, and realizes a rapid response to abnormal states, ensuring the safe operation of the elevator and improving the maintenance efficiency and safety.
[0170] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A voiceprint recognition method for elevator operation status monitoring, characterized in that: include: Collecting sound samples generated by the elevator during various working conditions, and selecting a stable period in the collected sound samples as a reference sound segment; Preprocessing the reference sound segment to remove noise interference and enhance sound features, transforming the processed reference sound segment and extracting inherent frequency components; Using the inherent frequency components to construct a sound identification map, comparing the newly collected elevator sound with the constructed sound identification map, and recording the differences; Based on the difference points, it is evaluated whether the current state of the elevator deviates from the normal range, and according to the evaluation result, a corresponding maintenance or alarm signal is triggered.
2. A voiceprint recognition method for elevator operation status monitoring according to claim 1, characterized in that: The sound samples generated by the elevator during various working conditions are collected, including: When the elevator is in different working conditions, the sensor array is used to capture sound signals, and the corresponding environmental variables are recorded, and the average amplitude and frequency distribution of the captured sound signals are calculated; Based on the calculated average amplitude and frequency distribution, the voiceprint vector is constructed for each working condition, and the degree of state change is quantified by comparing the differences between adjacent working conditions.
3. A voiceprint recognition method for elevator operation status monitoring according to claim 2, characterized in that: The step of selecting a stable period from the collected sound samples as a reference sound segment includes: In the acquired sound samples, the stable sections in each sample are identified and marked to ensure that the amplitude fluctuation in the stable sections does not exceed a preset threshold; For each marked stable segment, the energy of the sound signal in the segment is calculated, and the stable segment with the highest energy value is selected as the reference sound segment; The selected reference sound clips are normalized so that the maximum amplitude is normalized to unity.
4. A voiceprint recognition method for elevator operation status monitoring according to claim 3, characterized in that: The preprocessing of the reference sound segment to remove noise interference and enhance sound features includes: For the selected reference sound clip, convert it into the frequency domain, identify and mark the low energy area in the spectrum; A filter function is designed to reduce the influence of noise, and the filter function is applied to the frequency domain representation to achieve denoising and feature enhancement, and then converted back to the time domain to obtain the preprocessed reference sound clip.
5. A voiceprint recognition method for elevator operation status monitoring according to claim 4, characterized in that: The step of transforming the processed reference sound segment and extracting the inherent frequency component comprises: For the preprocessed reference sound clip, a time-frequency representation is generated; Based on the time-frequency representation, identify and record significant frequency points, construct a frequency distribution histogram, and count the number of occurrences of each frequency; The frequency that occurs most frequently is selected as the natural frequency component to ensure that the frequency can reflect the core sound characteristics of the elevator under normal operating conditions.
6. A voiceprint recognition method for elevator operation status monitoring according to claim 5, characterized in that: The step of constructing a sound identification map using the inherent frequency components includes: Based on the determined natural frequency components, a frequency-time matrix is constructed, and the importance of different frequencies is evaluated based on the constructed frequency-time matrix; The core features representing the sound of normal elevator operation are screened out in the frequency-time matrix, and a sound identification map is drawn to ensure that the map clearly shows the sound characteristics of the elevator during normal operation.
7. A voiceprint recognition method for elevator operation status monitoring according to claim 6, characterized in that: The newly collected elevator sound is compared with the constructed sound identification map, and the differences are recorded, including: Preprocessing the newly collected elevator sound samples to generate preprocessed new sound clips; Based on the obtained pre-processed new sound segment, calculate its time-frequency representation to ensure that it has the same time and frequency resolution as the sound identification map; Based on the obtained time-frequency representation, a frequency-time matrix of the new sound clip is constructed, the frequency-time matrix is compared element by element with the sound identification map, visual elements are generated, the difference is calculated, and all frequency points exceeding the set threshold are recorded as difference points.
8. A voiceprint recognition method for elevator operation status monitoring according to claim 7, characterized in that: The evaluation of whether the current state of the elevator deviates from the normal range based on the difference point includes: Based on the generated visual elements, the color intensity and position coordinates of each element are extracted; Based on the extracted information, weights are assigned according to the importance of different areas of the elevator to form a weighted color-position matrix; According to the formed weighted color-position matrix, all visual elements are integrated to create a multi-layer overlay map, and combined with the timeline to generate a dynamic comprehensive view, which shows the real-time status of each part of the elevator over time.
9. A voiceprint recognition method for elevator operation status monitoring according to claim 8, characterized in that: The triggering of corresponding maintenance or alarm signals according to the evaluation results includes: Receiving the dynamic integrated view, parsing the color change and position change therein, and forming a change vector; According to the obtained change vector, by comparing the difference with a preset threshold, determining whether there is a state beyond a normal range; When an abnormality is confirmed, a priority index is calculated according to the degree of abnormality, and a response action is selected based on the calculated priority index; If the priority index exceeds the critical value, an alarm signal is triggered; otherwise, a maintenance notification is issued.
10. A voiceprint recognition system for elevator operation status monitoring for executing the method according to any one of claims 1 to 9, characterized in that: include: A sound sample collection and reference segment selection module is used to collect sound samples generated by the elevator during various working conditions, and select a stable period from the collected sound samples as a reference sound segment; A preprocessing and natural frequency component extraction module, used to preprocess the reference sound segment, remove noise interference and enhance sound features, transform the processed reference sound segment and extract natural frequency components; A sound identification map construction and comparison module, used to construct a sound identification map using the inherent frequency components, compare the newly collected elevator sound with the constructed sound identification map, and record the differences; The status assessment and response trigger module is used to assess whether the current status of the elevator deviates from the normal range based on the difference points, and trigger the corresponding maintenance or alarm signal according to the assessment results.
Citation Information
Patent Citations
Voiceprint recognition and fault diagnosis monitoring alarm system for elevator anomaly
CN110861988A
Wind power cabin monitoring method and system based on sound signal processing
CN117028171A
Speech recognition method and device based on artificial intelligence and robot equipment
CN118865970A
Monitoring of operational condition of an elevator
EP3459889A1
Device and method for identifying speech content
JP2012146116A