A Biometric Elevator Calling Method Based on Wavelet Tight Frames

By using wavelet tight frame and voiceprint feature matching technology, the problems of noise interference and pronunciation differences in the voice call system of elevators have been solved, realizing automatic elevator calling without wake-up words or floor instructions, thus improving the convenience and intelligence of elevators.

CN119774390BActive Publication Date: 2025-10-31HITACHI BUILDING TECH GUANGZHOU CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411904986.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-10-31
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

Existing voice-activated elevator systems suffer from missed or false alarms due to environmental noise interference and differences in user pronunciation, affecting user experience and the level of elevator intelligence.

Method used

We employ wavelet compact frameworks to decompose and denoise speech signals, combine X-Vectors and SpeakerNet models to extract voiceprint features, achieve automatic elevator retrieval through voiceprint matching, and improve elevator retrieval accuracy by combining content recognition models.

Benefits of technology

When the user does not trigger a wake-up word or floor command, the system automatically calls the elevator based on voiceprint characteristics, improving user convenience and the accuracy and reliability of the elevator calling system, reducing mismatches, and improving the level of intelligent elevator services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119774390B_ABST
    Figure CN119774390B_ABST
Patent Text Reader

Abstract

This invention discloses a biometric elevator calling method based on wavelet compact frames. Addressing the poor user experience caused by the frequent missed and false detections of wake words and floor commands in existing voice-based elevator calling systems, this invention first utilizes wavelet compact frames (such as redundant wavelet transform) to denoise and reconstruct the voice signal within the elevator car. Then, it combines X-Vectors and other models to extract voiceprint features and constructs an ID differentiation method based on a large number of voice-ID samples. Voiceprints are monitored within the elevator's single-floor operating range. Based on the internal calling situation, voiceprint and floor matching information is stored in a candidate list. Once a threshold for the frequency of voiceprint occurrence and the corresponding floor proportion is met, it is added to the official list. When a voiceprint is detected again, floor matching is automatically performed for elevator calling. Furthermore, combining wake word and floor command content recognition models can further improve accuracy. This invention improves the convenience and accuracy of elevator calling, integrates multiple technologies to enhance the level of intelligent elevator services, and has significant application value in various elevator scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of elevator control technology, and in particular to a biometric elevator calling method based on wavelet tight frames. Background Technology

[0002] In the field of modern elevator technology, voice-activated elevator calling has emerged as an important means to improve the intelligence level of elevators and user convenience. Its original intention was to allow passengers to quickly and accurately summon the elevator to their designated floor through simple voice commands, without having to manually operate the elevator control panel. This allows passengers to easily use elevator services even when their hands are busy (such as carrying items, pushing a wheelchair, etc.) or when it is inconvenient to touch the elevator buttons.

[0003] However, current voice-activated elevator systems that rely on designated wake words and floor commands have revealed numerous shortcomings in practical applications. On one hand, voice recognition technology is highly susceptible to environmental factors. In the relatively enclosed space of an elevator car with high passenger flow, environmental noise sources are widespread and complex. For example, mechanical noise during elevator operation, airflow noise from the car's ventilation system, noise from passengers entering and exiting the car, and mixed background noise from multiple people talking or chatting simultaneously can all severely interfere with the voice signal. This interference makes it difficult for the voice recognition system to accurately capture wake words and floor commands, leading to frequent missed detections—meaning that valid user commands are not recognized by the system, the elevator cannot respond, and this severely impacts the user's normal elevator usage needs.

[0004] On the other hand, even in relatively quiet environments, voice recognition systems may experience false alarms due to the diversity and ambiguity of voice commands. Different users have different pronunciation habits, speaking speeds, intonations, and accents, causing the same wake-up word or floor command to exhibit a wide range of variations in voice characteristics. For example, regional dialects may cause the pronunciation of the wake-up word to deviate significantly from the system's preset standard pronunciation, leading to it being misjudged as an invalid command. Alternatively, when a user speaks quickly or unclearly, the floor command may be incorrectly identified as a different floor or command, triggering an incorrect elevator call. This not only wastes elevator resources but may also inconvenience other passengers.

[0005] In summary, existing voice-activated elevator calling functions cannot reliably and efficiently meet user needs in actual use due to issues with missed or false detection of wake words and floor commands. This severely restricts the improvement of the quality of intelligent elevator services. An innovative technical solution is urgently needed to overcome these drawbacks and achieve more accurate, intelligent, and convenient elevator calling operations. Summary of the Invention

[0006] The purpose of this invention is to provide a biometric elevator calling method based on wavelet compact frames, which can automatically call elevators based solely on voiceprint even when the user does not trigger a specified wake-up word or floor command. For example, when a user is talking or making a call in the elevator and it is inconvenient to issue a specific voice command, the elevator can still be called based on the corresponding floor matched by the user's voiceprint characteristics, greatly improving the convenience of using the elevator and enhancing the overall user experience.

[0007] To achieve the above objectives, the present invention provides a biometric elevator calling method based on wavelet compact frames, comprising the following steps:

[0008] Speech signal processing

[0009] This invention first employs a wavelet compact frame to decompose the speech signal. The wavelet compact frame possesses excellent time-frequency localization characteristics, enabling the decomposition of the speech signal at different scales and frequencies, thereby effectively extracting the signal's feature information. During the decomposition process, basis functions that conform to compact support and redundancy (such as redundant wavelet transform) are selected, which improves the adaptability to the signal while ensuring signal processing accuracy.

[0010] Thresholding is applied to the wavelet coefficients obtained from the decomposition. By setting an appropriate threshold, wavelet coefficients corresponding to noise are removed, while the main feature coefficients of the speech signal are retained. The signal is then reconstructed, achieving noise reduction and reconstruction of the speech signal. The noise-reduced speech signal can effectively reduce the impact of environmental noise and other interference factors on subsequent speaker extraction and recognition, improving the accuracy and reliability of the entire system.

[0011] Voiceprint feature extraction and ID differentiation

[0012] The reconstructed speech signal is combined with advanced speaker extraction models (such as X-Vectors and SpeakerNet) to extract speaker features. These models can deeply mine the unique features of speaker signals from the speech signal and transform the speech signal into a representative speaker feature vector.

[0013] Based on a large number of different voice-ID samples, this paper constructs an ID differentiation method based on voiceprint features by statistically analyzing the distribution differences of voiceprint features corresponding to different IDs in the feature space. This method does not identify specific user IDs, but focuses on distinguishing whether different audios come from different individuals, thus providing an effective basis for identity differentiation for subsequent voiceprint and floor matching without infringing on user privacy.

[0014] Voiceprint matching with floor level during elevator operation

[0015] The monitoring cycle is defined as the time from when the elevator door opens on a certain floor until the door opens on the next floor. Within this cycle, when a new voiceprint is detected inside the elevator, voiceprint matching is performed using a specific matching mechanism (such as distance-based calculations, including Euclidean distance or cosine similarity methods to calculate the degree of matching between voiceprint features).

[0016] If a new internal call is made during this period, the new floor match and the corresponding ID of the voiceprint will be added to the new entry in the candidate match list; if no new internal calls are made and multiple internal calls have been registered, all registered internal calls and the corresponding ID of the voiceprint will be added to the new entry in the candidate match list. In this way, the association between voiceprint and possible floor call requests is gradually established.

[0017] In the candidate matching list, certain matching rules are set. That is, if a voiceprint appears more than n times and corresponds to a certain floor m times, and if m / n is greater than a preset threshold, then it is considered that there is a strong correlation between the voiceprint and the floor, and the pairing combination of the voiceprint and the floor is added to the official list. The pairing combinations in the official list will serve as the basis for subsequent automatic elevator calling.

[0018] Automatic elevator calling

[0019] When a certain voiceprint is heard again inside the elevator, the system quickly matches the corresponding floor from the official list and automatically registers the internal call, thus enabling automatic elevator calling. Furthermore, when a user triggers a specified wake-up word or floor command, the method of this invention can further improve the accuracy of automatic elevator calling by combining it with a content recognition model. The content recognition model can more accurately parse and judge the user's voice commands, complementing the voiceprint recognition results and further improving the performance and reliability of the entire elevator calling system.

[0020] The present invention has the following beneficial effects:

[0021] 1. Improved elevator calling convenience: The method of this invention can automatically call elevators based solely on voiceprint even when the user does not trigger a specified wake-up word or floor command. For example, when a user is talking or making a call in the elevator and it is inconvenient to issue a specific voice command, the elevator can still be called based on the corresponding floor according to the user's voiceprint characteristics, which greatly improves the convenience of using the elevator and enhances the overall user experience.

[0022] 2. Optimize Voiceprint Processing Accuracy: By decomposing, denoising, and reconstructing the speech signal using a wavelet compact frame, noise interference in the speech signal is effectively reduced, improving the accuracy of subsequent voiceprint extraction. Advanced voiceprint extraction models such as X-Vectors and SpeakerNet further ensure the effectiveness and reliability of voiceprint feature extraction, thus laying a solid foundation for accurate voiceprint matching and floor-based elevator calling.

[0023] 3. Enhanced Reliability of Voiceprint Matching: The voiceprint-based ID differentiation method, constructed using a large number of voice-ID samples, does not involve the storage of specific user ID information and is only used for matching. Without infringing on user privacy, it significantly improves the reliability and accuracy of voiceprint and floor matching by accurately distinguishing and statistically analyzing different voiceprint features, combined with a multi-stage matching mechanism of candidate matching lists and official lists, such as the threshold judgment of the ratio of voiceprint occurrences to corresponding floor occurrences. This reduces the occurrence of false matches and makes elevator calling more intelligent and accurate.

[0024] 4. Integrating multiple technologies to enhance performance: When a user triggers a specified wake-up word or floor command, the accuracy of automatic elevator calling can be further improved by combining content recognition models. The voiceprint recognition technology is organically integrated with traditional voice command recognition technology, giving full play to their respective advantages, and further improving the performance and adaptability of the entire elevator calling system. It can meet the needs of different user scenarios and usage habits, and effectively improve the level of intelligent elevator services. Attached Figure Description

[0025] Other features, objects, and advantages of the present invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings:

[0026] Figure 1 This is a flowchart illustrating a biometric elevator calling method based on wavelet tight frames according to an embodiment of the present invention.

[0027] Figure 2 This is the framework intent of a biometric elevator calling method based on wavelet tight frame according to an embodiment of the present invention. Detailed Implementation

[0028] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. The illustrative embodiments and descriptions of the present invention are used to explain the present invention, but are not intended to limit the present invention.

[0029] It should be noted that all directional indicators (such as up, down, left, right, front, back, upper end, lower end, top, bottom, etc.) in the embodiments of the present invention are only used to explain the relative positional relationship and movement of each component in a certain specific posture (as shown in the figure). If the specific posture changes, the directional indicator will also change accordingly.

[0030] In this invention, unless otherwise explicitly specified and limited, the term "connection" should be interpreted broadly. For example, "connection" can be a fixed connection, a detachable connection, or an integral part; it can be a mechanical connection or an electrical connection; it can be a direct connection or an indirect connection through an intermediate medium; it can be the internal communication of two components or the interaction between two components, unless otherwise explicitly limited. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.

[0031] Furthermore, in this invention, descriptions involving "first," "second," etc., are for descriptive purposes only and should not be construed as indicating or implying their relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined with "first" or "second" may explicitly or implicitly include at least one of that feature. Additionally, the technical solutions of various embodiments can be combined with each other, but only on the basis of being achievable by those skilled in the art. If the combination of technical solutions is contradictory or impossible to implement, such a combination should be considered non-existent and not within the scope of protection claimed by this invention.

[0032] The following is in conjunction with the appendix Figure 1 and attached Figure 2 The present invention will be described in further detail below.

[0033] A biometric elevator calling method based on wavelet compact frames includes the following steps:

[0034] 1. Speech signal acquisition and preprocessing:

[0035] A high-sensitivity microphone array is installed inside the elevator car to collect voice signals. The microphone array can collect sounds from different directions and enhance the strength and clarity of the voice signal through signal processing algorithms.

[0036] The acquired speech signals are digitized and converted into a digital signal format that can be processed by a computer. Preliminary filtering is then performed to remove some obvious high-frequency noise and low-frequency interference signals.

[0037] 2. Speech denoising based on wavelet compact frames:

[0038] The preprocessed speech signal is decomposed using redundant wavelet transform. A suitable wavelet basis function (such as the Daubechies wavelet basis) and the number of decomposition levels (e.g., 3-5 levels) are selected, and the redundancy (e.g., 2-3 times redundancy) is determined based on the characteristics of the speech signal.

[0039] The wavelet coefficients obtained from the decomposition are thresholded. Soft or hard thresholding methods are used, with the threshold value determined based on noise level estimation. For wavelet coefficients with absolute values ​​less than the threshold, the soft thresholding method sets them to 0, while the hard thresholding method discards them directly. The signal is then reconstructed using inverse wavelet transform to obtain the denoised speech signal.

[0040] 3. Voiceprint feature extraction and ID differentiation model training:

[0041] The denoised speech signal is input into X-Vectors or SpeakerNet speaker extraction models. These models are built on deep learning frameworks (such as TensorFlow or PyTorch) and trained on large amounts of speech data.

[0042] A large amount of voice-ID sample data from different individuals is collected. The voice signals are input into a voiceprint extraction model to obtain voiceprint feature vectors. Then, the mean, variance, and other statistics of the voiceprint feature vectors corresponding to different IDs are statistically analyzed in various dimensions to construct an ID differentiation model based on voiceprint features. For example, cluster analysis can be used to cluster the voiceprint feature vectors of different IDs, determine the boundary features between different clusters, and thus distinguish whether different audio belongs to different IDs.

[0043] 4. Matching elevator voiceprints with floor levels during elevator operation:

[0044] When the elevator door opens on a certain floor, the voiceprint and floor matching monitoring program is initiated and continues until the door on the next floor opens. During this period, voice signals inside the elevator are continuously collected and processed.

[0045] For newly acquired speech signals, voiceprint features are extracted and matched with existing voiceprint templates. The distance between voiceprint features is calculated using Euclidean distance or cosine similarity. When the distance is less than a set matching threshold, a new voiceprint is considered to have been matched.

[0046] If a new internal call operation is added during this period, the new floor and the corresponding ID of the matched voiceprint will be added to the new entry in the candidate matching list; if no new internal call is added and multiple internal calls have already been registered, all registered internal calls and the corresponding ID of the voiceprint will be added to the candidate matching list.

[0047] The candidate matching list is analyzed periodically to count the number of times each voiceprint appears (n) and the number of times it appears for each floor. When m / n is greater than a preset threshold (e.g., 0.6), the pairing of the voiceprint and the floor is added to the official list.

[0048] 5. Automatic elevator call execution:

[0049] When a certain voiceprint is detected again inside the elevator, the system searches for the corresponding floor information in the official list. If a match is found, the system automatically registers an internal call operation and controls the elevator to go to the corresponding floor.

[0050] When a user triggers a specified wake word or floor command, the content recognition model parses the voice command, extracting the wake word and floor information. Simultaneously, the voiceprint recognition system continues to work, matching the voiceprint corresponding to the command with existing voiceprint templates to confirm the user's identity. If both match successfully, the corresponding elevator call operation is executed, further improving the accuracy and reliability of elevator calling.

[0051] Example:

[0052] This paper applies a wavelet-based biometric elevator calling method to the elevator system of a high-rise office building.

[0053] First, four high-sensitivity microphones are evenly distributed on the top of each elevator car, forming a microphone array. The sampling frequency of this microphone array is set to 16kHz, which can effectively collect voice signals from all directions inside the car.

[0054] After the voice signal is acquired, it enters the preprocessing stage. For example, during an elevator operation, the acquired raw voice signal contains obvious mechanical humming noise from the elevator and background noise from passengers talking slightly inside the car. After digital conversion, a high-pass filter with a cutoff frequency of 50Hz is used to remove the low-frequency mechanical humming noise, and a low-pass filter with a cutoff frequency of 8kHz is used to filter out weak high-frequency interference signals, resulting in a preliminarily purified voice signal.

[0055] Next, speech denoising based on a wavelet compact frame was performed. The Daubechies wavelet basis was selected, with a decomposition level of 4 layers and a redundancy of 2.5 times. Soft thresholding was applied to the decomposed wavelet coefficients, with the threshold set to 0.3 based on the statistical characteristics of the elevator noise environment in the office building. After reconstruction, noise in the speech signal was significantly suppressed, and previously unclear speech details became clearly discernible.

[0056] Then, voiceprint feature extraction and ID differentiation model training were performed. Voice-ID sample data from 500 different employees in the office building were collected, and voiceprint features were extracted using the X-Vectors model built on the TensorFlow framework. After training on a large amount of data, the constructed ID differentiation model can effectively distinguish the voiceprint features of different employees. For example, when two employees are talking in the elevator car, the model can accurately determine that their voiceprints come from different individuals, even without knowing the specific employee's identity information.

[0057] During the voiceprint and floor matching phase of the elevator operation, when the elevator opens from the 1st floor, three passengers enter the car. The voice of one of these passengers during their conversation is captured by the system. After processing, a new voiceprint is matched. As the elevator ascends, at the 3rd floor, a passenger presses the internal call button for the 10th floor. At this point, the system adds the ID corresponding to the 10th floor and the new voiceprint to the candidate matching list. The elevator continues to run, and two more passengers enter the car and converse. During this time, the voiceprint is detected multiple times again, and no new internal call operations occur. However, internal calls for the 15th and 20th floors have already been registered, so these two floors and their corresponding IDs are also added to the candidate matching list. When the voiceprint appears a total of 8 times, corresponding to the 10th floor 3 times, the 15th floor 2 times, and the 20th floor 3 times, since 3 / 8 > 0.6 (the preset threshold), the pairings of the 10th, 15th, and 20th floors with this voiceprint are added to the official list.

[0058] Finally, during the automatic elevator call execution phase, when the passenger rides the elevator again and speaks inside the car, the system quickly matches the corresponding floor in the official list. If the passenger also triggers the specified wake-up word and the floor command for the 25th floor, the content recognition model accurately parses the command and, combined with voiceprint recognition to confirm the passenger's identity, the elevator will not only go to the previously matched floor but also add the 25th floor to the elevator call sequence. This further improves the accuracy and flexibility of elevator calls, providing office building employees with an efficient and convenient elevator service experience.

[0059] This implementation case demonstrates that the elevator calling method of the present invention can operate effectively in complex elevator usage scenarios, overcoming many problems of traditional voice calling and significantly improving the intelligence level and user experience of elevator calling.

[0060] In summary, compared with the prior art, the present invention has the following beneficial effects:

[0061] 1. Improved elevator calling convenience: The method of this invention can automatically call elevators based solely on voiceprint even when the user does not trigger a specified wake-up word or floor command. For example, when a user is talking or making a call in the elevator and it is inconvenient to issue a specific voice command, the elevator can still be called based on the corresponding floor according to the user's voiceprint characteristics, which greatly improves the convenience of using the elevator and enhances the overall user experience.

[0062] 2. Optimize Voiceprint Processing Accuracy: By decomposing, denoising, and reconstructing the speech signal using a wavelet compact frame, noise interference in the speech signal is effectively reduced, improving the accuracy of subsequent voiceprint extraction. Advanced voiceprint extraction models such as X-Vectors and SpeakerNet further ensure the effectiveness and reliability of voiceprint feature extraction, thus laying a solid foundation for accurate voiceprint matching and floor-based elevator calling.

[0063] 3. Enhanced Reliability of Voiceprint Matching: The voiceprint-based ID differentiation method, constructed using a large number of voice-ID samples, does not involve the storage of specific user ID information and is only used for matching. Without infringing on user privacy, it significantly improves the reliability and accuracy of voiceprint and floor matching by accurately distinguishing and statistically analyzing different voiceprint features, combined with a multi-stage matching mechanism of candidate matching lists and official lists, such as the threshold judgment of the ratio of voiceprint occurrences to corresponding floor occurrences. This reduces the occurrence of false matches and makes elevator calling more intelligent and accurate.

[0064] 4. Integrating multiple technologies to enhance performance: When a user triggers a specified wake-up word or floor command, the accuracy of automatic elevator calling can be further improved by combining content recognition models. The voiceprint recognition technology is organically integrated with traditional voice command recognition technology, giving full play to their respective advantages, and further improving the performance and adaptability of the entire elevator calling system. It can meet the needs of different user scenarios and usage habits, and effectively improve the level of intelligent elevator services.

[0065] The technical solutions provided by the embodiments of the present invention have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of the embodiments of the present invention. The descriptions of the embodiments above are only for helping to understand the principles of the embodiments of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the embodiments of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A biometric elevator calling method based on wavelet compact frames, characterized in that, Includes the following steps: Step S1: Decompose the speech signal based on the wavelet compact frame, transform it using basis functions that conform to compact support and redundancy, and reconstruct the signal after thresholding the wavelet coefficients to achieve noise reduction and reconstruction of the speech signal. Step S2: Extract voiceprint features using the reconstructed speech signal combined with the voiceprint extraction model; Step S3: Based on a large number of different voice-ID samples, statistically analyze the differences in voiceprints of different IDs, and construct an ID differentiation method based on voiceprint features; Step S4: Starting from the moment the elevator opens its door on a certain floor, and ending before the door opens on the next floor, if a new voiceprint is matched inside the elevator during this period, a match is made using distance-based calculation or a matching mechanism: If a new internal call is added during this period, the ID corresponding to the new floor match and the voiceprint is added to the new entry in the candidate match list; if no new internal call is added and multiple internal calls have been registered, all registered internal calls and the ID corresponding to the voiceprint are added to the new entry in the candidate match list. Step S5: In the candidate matching list, when a certain voiceprint appears more than n times and corresponds to a certain floor m times, if m / n is greater than the preset threshold, the pairing combination of the voiceprint and the floor is added to the formal list. Step S6: When a certain voiceprint is detected in the elevator, match the corresponding floor from the official list and automatically register the internal call.

2. The biometric elevator calling method based on wavelet compact frames according to claim 1, characterized in that: The voiceprint extraction models include, but are not limited to, the X-Vectors model and the SpeakerNet model.

3. The biometric elevator calling method based on wavelet compact frames according to claim 1, characterized in that: The ID differentiation method based on voiceprint features only distinguishes whether the audio belongs to different IDs, and does not identify specific IDs.

4. The biometric elevator calling method based on wavelet compact frames according to claim 1, characterized in that: The wavelet compact frame employs redundant wavelet transform for speech signal decomposition and reconstruction.

5. The biometric elevator calling method based on wavelet compact frames according to claim 1, characterized in that: The ID differentiation method based on voiceprint features constructs a differentiation model by statistically analyzing the distribution differences of voiceprint feature vectors across different dimensions.

6. The biometric elevator calling method based on wavelet compact frames according to claim 1, characterized in that: The distance calculation in the matching mechanism uses Euclidean distance or cosine similarity to calculate the degree of matching between voiceprint features.

7. The biometric elevator calling method based on wavelet compact frames according to claim 1, characterized in that: When a user triggers a specified wake word or floor command, a content recognition model is used to improve the accuracy of automatic elevator calling.

Citation Information

Patent Citations

  • Elevator calling method and system based on voice technology

    CN115557339A

  • Voiceprint recognition exhales terraced control system

    CN207483097U