Voice Print and Household Number Combination-Based Common Entrance Access Control System
A contactless access control system for multi-family housing uses voice fingerprint and household number recognition to address hygiene, security, and operational issues, offering secure and convenient entry by integrating voice recognition and fingerprint verification, with noise and pronunciation normalization for improved accuracy.
Patent Information
- Authority / Receiving Office
- KR · KR
- Patent Type
- Applications
- Current Assignee / Owner
- INTELLECTURE FUTURE IP MANAGEMENT CO LTD
- Filing Date
- 2026-06-30
- Publication Date
- 2026-07-21
AI Technical Summary
Conventional access control systems for multi-family housing complexes face hygiene issues, security vulnerabilities, and operational inconveniences due to physical contact and reliance on media like keypads, cards, and smartphones, and existing voice recognition technologies are not designed to handle the unique environmental and linguistic challenges of such settings.
A contactless access control system that combines voice fingerprint and household number recognition, allowing entry through a single spoken voice utterance, integrating voice recognition and fingerprint verification to generate a dynamic pass key, and supporting multiple registrations per household with noise and reverberation compensation.
Provides secure, hygienic, and convenient access by eliminating physical contact, enhancing recognition accuracy through noise and pronunciation normalization, and preventing unauthorized entry, while supporting flexible access policies for multi-resident households.
Smart Images

Figure PAT00001_ABST
Abstract
Description
Technology Field
[0001] The present invention relates to a common entrance access control system for a multi-family housing complex, and more specifically, to a common entrance access control system based on the combination of voice fingerprint and household number, wherein a household number information extracted from the spoken voice and the speaker's unique voice fingerprint information are combined to form a dynamic pass key through a single action in which a user attempting to enter verbally speaks the household number of their resident unit, thereby enabling the common entrance door to be opened contactless without any incidental actions such as separate keypad input, card tagging, or smartphone operation. Background Technology
[0002] In multi-unit dwellings such as apartments, villas, and officetels, the exterior and interior are separated by a common entrance door installed at the common entrance of the complex or building, and a common entrance access control system is operated to prevent unauthorized entry by non-residents. Conventional common entrance access control systems for multi-unit dwellings generally include keypad input methods, card or RFID tag methods, facial recognition methods, and mobile terminal integration methods.
[0003] The keypad input method is the most widely adopted method in which the entrance door opens when a resident enters their unit number and a pre-set password into the keypad of the lobby phone installed at the common entrance. However, this keypad input method has limitations: firstly, there are hygiene issues as physical contact with an input device used jointly by multiple residents is unavoidable; secondly, there is a constant risk that the password entry process may be exposed to passersby or installed video equipment (so-called "over-the-shoulder peeping attacks"); thirdly, the number of digits and patterns of the password can be traced back from fingerprints or wear accumulated on the keypad surface; and fourthly, the act of accurately entering the unit number and password causes significant inconvenience to children, the elderly, or residents carrying luggage in both hands.
[0004] The card or RFID tag method involves opening the door by tagging an access card issued to the resident to a reader near the main entrance. While this method partially resolves hygiene issues and password exposure concerns associated with keypad input, it imposes the burden on residents to carry their access cards at all times. Furthermore, it presents security vulnerabilities where a person obtaining a lost or stolen card can impersonate a resident to gain entry, and the risk of unauthorized access via cloned cards is increasing due to advancements in card duplication technology.
[0005] Facial recognition is a method of controlling access by having cameras installed at common entrances capture the faces of approaching users and compare them to pre-registered facial images of residents. However, this method has limitations, as recognition rates are significantly affected by lighting conditions, whether a mask is worn, whether a hat or glasses is worn, facial angle, changes in makeup, and seasonal variations in appearance. Furthermore, security issues have been reported, such as vulnerability to facial forgery attacks using printed photos or images displayed on screens. Additionally, facial images are sensitive information that directly exposes a user's identity and appearance, posing a significant burden in terms of personal information protection.
[0006] Meanwhile, although voice recognition and speaker recognition technologies have made significant advancements in fields such as telecommunications, finance, and mobile device unlocking, these conventional voice-based authentication technologies differ fundamentally from the field of access control for common entrances of multi-family housing to which this invention belongs, in terms of their application environment, user behavior patterns, technical problems to be solved, and operational effects. For example, voice identity authentication in telephone networks utilizes voice as an additional auxiliary means in an environment where the caller's phone number already functions as a primary indicator of the caller's identity; voice authentication in financial transactions is structured to auxiliaryly verify a claim via voice after the user has made an identity claim by inserting a card or separately entering an account number; and voice unlocking of mobile devices is a single-user authentication for a privately owned device, which differs fundamentally from the technical environment of a common entrance in multi-family housing shared by multiple households.
[0007] Furthermore, while there have been some reported cases in the field of conventional common entrance access control where voice recognition was merely listed as one of the possible authentication methods, these merely refer to voice as a general biometric authentication method equivalent to fingerprint or facial recognition. The mechanism in which a user verbally pronounces their apartment number—whereby apartment number information and voice fingerprint information are simultaneously extracted and combined from a single utterance to form a dynamic pass-through key, the spoken apartment number functions as a claim of the speaker's residency status, and the voice fingerprint of the same utterance verifies the authenticity of that claim—has never been disclosed or suggested in any conventional common entrance access control technology.
[0008] Furthermore, the common entrance of a multi-family dwelling has the following unique technical environment, which is fundamentally different from a personal terminal or 1:1 transaction channel occupied by a single user. First, a single common entrance input device is shared by tens to thousands of households; second, since multiple residents (couples, children, cohabitants, etc.) usually reside in each household, there are multiple legitimate entrants per household; third, as many passersby are located around the common entrance at the same time, there is a high possibility that voices other than the speaker's are captured simultaneously; fourth, environmental noise such as vehicle traffic, elevator operation sounds, children crying, rain, and wind is always present; and fifth, there is a special characteristic that there is a high possibility of voice recognition errors because there are various variations in the pronunciation of Korean house numbers, such as Sino-Korean pronunciation, pronunciation with omitted digits, pronunciation with mixed English numbers, and pronunciation with separated digits, and pronunciations between adjacent house numbers are very similar. Conventional voice authentication technologies in other fields are not designed to take into account the unique environmental and linguistic conditions of multi-family housing as described above; therefore, simply applying these conventional technologies to common entrances does not resolve the aforementioned inherent problems.
[0009] Therefore, there is an urgent technical need for a new type of access control system that can overcome the limitations of conventional access control methods while simultaneously solving the technical challenges unique to the common entrance environment of multi-family housing. The problem to be solved
[0010] The present invention has been devised to solve the problems of the prior art described above, and aims to provide a completely contactless and media-free access control system for the access control of the common entrance of a multi-family housing complex, which enables entry solely through a single voice act of uttering one's unit number once, without the user making any physical contact with input devices that multiple people jointly touch, such as keypads, card readers, or fingerprint scanners, and without the need to carry or operate physical media such as separate access cards or smartphones.
[0011] More specifically, the present invention aims to provide a novel access control mechanism in which the utterance itself functions as an identity claim for resident eligibility and the voice fingerprint verifies the authenticity of the claim by simultaneously extracting household number information specifying the household number to be entered from a single spoken voice of a user and voice fingerprint information representing the unique physical voice characteristics of the speaker, generating a dynamic pass key by combining the two, and opening the door only when the voice fingerprint of the resident registered for the spoken household number matches the real-time voice fingerprint.
[0012] In addition, another objective of the present invention is to provide an operational flow specialized for multi-family environments, such as supporting multiple voice fingerprint registrations per household in accordance with the environmental characteristics of multi-family housing where multiple residents live in a single household, branching access policies by resident type, and automatically switching to a call to the wall pad of the corresponding household rather than simply refusing access when a visitor who is not a legitimate resident calls out the apartment number.
[0013] Furthermore, another objective of the present invention is to comprehensively solve technical challenges arising from the common entrance environment of multi-family housing itself, such as normalizing various variants of Korean lake pronunciation, correcting speech recognition errors between adjacent lakes through backfeedback of voice fingerprint matching results, actively processing noise and reverberation unique to the common entrance environment, accurately separating and recognizing only the utterance of the main speaker even in a multi-passenger environment, and ensuring security by detecting recording and playback attacks. means of solving the problem
[0014] To achieve the above objective, the common entrance access control system based on the combination of voice fingerprint and household number according to the present invention comprises: a voice input unit that receives a voice of a user including the household number of the destination to be entered; a database unit in which unique voice fingerprint data of a legitimate user is matched and stored for each household number; and a control unit that is linked to the voice input unit and the database unit to simultaneously extract household number information and real-time voice fingerprint information from the voice and generate a dynamic pass key by combining them to determine whether entry is permitted.
[0015] In addition, the system according to the present invention may further include at least one of the following: a reverse-binding mechanism in which the user's speech itself functions as a claim for access rights to the corresponding household and the authenticity of the claim is verified by whether the voice fingerprint matches; a mechanism for registering multiple resident voice fingerprints per household and policy branching by resident type; an automatic intercom call switching mechanism for a speaker who is not a legitimate resident; a mechanism for normalizing Korean apartment pronunciation variations and correcting misrecognition of adjacent apartments; a mechanism for modeling common entrance environment noise and reverberation compensation; a mechanism for separating multiple speakers and directional beamforming; and a recording playback attack detection mechanism. Effects of the invention
[0016] The present invention provides the following significant effects compared to conventional technology in the field of access control for common entrances of multi-family housing.
[0017] First, the present invention is recognized for its inventive step in that it is a technology specialized in the specific field of physical access control, namely access control for common entrances of multi-family housing. Conventional examples where voice recognition technology and speaker recognition technology have been combined are generally fields that differ essentially from the field of access control for common entrances of multi-family housing to which the present invention belongs. These fields include remote transaction authentication in telephone networks, identity authentication in financial institution call centers, auxiliary authentication in automated deposit and withdrawal terminals, and unlocking of personal mobile devices. These technologies in dissimilar fields have been developed to solve identity identification problems within their respective fields and were not designed to solve the unique technical challenges arising in the physical space access control environment of common entrances in multi-family housing—namely, sharing by multiple households, multiple residents per household, noise specific to common entrances, variations in Korean apartment number pronunciation, misrecognition of adjacent apartments, and multiple passersby. Therefore, it is difficult for a person skilled in the art to clearly derive the motivation to combine voice authentication technology from a dissimilar field with the application of it to the common entrance of a multi-family housing complex, and even if a simple application is attempted, the aforementioned inherent problems are not resolved. In this respect, the present invention has pioneered a new technical domain that cannot be reached through the simple application of conventional technology.
[0018] Second, the present invention provides a significant improvement in user convenience by integrating the multi-step actions of conventional access control methods (such as keypad number input + password input, card tagging + PIN input, or number input + separate biometric authentication) into a single action, in that it simultaneously extracts and combines information designating the target household for entry and voice fingerprint information verifying the speaker's residency eligibility from a single voice act in which a user utters their household number. Since entry is completed simply by uttering one's number once, even if the user is carrying luggage in both hands, holding a child, or holding an umbrella, the precise finger manipulation and tagging actions at accurate locations required by conventional keypad or card systems are all eliminated.
[0019] Third, the generation number uttered by the user in the present invention provides a novel functional effect that distinguishes it from conventional technology in that it possesses a semantic nature essentially different from simple identification markers such as passwords, commands, and account numbers in conventional voice authentication technology. While the spoken text in conventional voice authentication technology was merely a marker to indicate the speaker's identity or identify a transaction, the utterance of the generation number in the present invention functions as an active assertion of the right to reside that the speaker is living in that space, and simultaneously as an act that specifies the target of entry into that space. This semantic combination is realized as a reverse-binding effect stating that "entry is prohibited for voices that are not of the resident registered for the uttered unit number," thereby enabling space-speaker combined authority verification where, if a resident of unit A utters unit B, entry is denied because there is no authority to reside in unit B, even if the speaker's voice fingerprint is registered in the system. This represents a new authority verification paradigm distinct from conventional simple 1:1 speaker verification or 1:N speaker identification.
[0020] Fourth, the present invention provides the effect of fundamentally avoiding the processing burden and accuracy degradation issues associated with 1:N speaker identification in multi-family housing environments where multiple households share a single common entrance. Conventional 1:N speaker identification methods require comparing the voice fingerprints of all registered residents with the input voice one by one, which results in longer processing times and higher misrecognition rates as the number of households increases. However, since the search space is immediately reduced to 1:1 or 1:minor based on the apartment number spoken by the user, the present invention ensures consistent processing speed and high accuracy regardless of the size of the complex. Furthermore, by supporting the registration of multiple resident voice fingerprints per household, it enables the separate identification of each family member and allows for policy branching, such as entry / exit times, based on resident type, thereby flexibly responding to various residential forms in multi-family environments.
[0021] Fifth, the present invention provides effects optimized for the Korean environment by actively resolving the linguistic challenges unique to Korean multi-family housing, namely the variability of Korean number pronunciation and the similarity of adjacent number pronunciations. A pronunciation variation normalization module pre-emptively absorbs the variability of Korean number pronunciation, where the same "301" is pronounced in various ways depending on the speaker, such as "Sambaeilho," "Samgongilho," "Three Zero One," or "Sam-Yeong-ilho." Furthermore, voice recognition errors between adjacent numbers with very similar pronunciations—such as "301" and "3001ho," or "201" and "210ho"—are corrected using a bidirectional combination method that determines the final number by comparing the similarity between the registered voice fingerprint of each candidate number and the input voice fingerprint. Consequently, the accuracy of number pronunciation recognition in a Korean environment is significantly improved. This Korean-specific processing is something that was not addressed at all in conventional foreign technologies designed for English voice authentication environments.
[0022] Sixth, the present invention provides the effect of fundamentally blocking security incidents caused by the loss, theft, misappropriation, or duplication of a medium by completely eliminating dependency on the medium. Since the user does not need to carry a separate access card or operate a smartphone, the inherent vulnerabilities of conventional medium-based access control methods, such as unauthorized entry due to card loss, impersonation entry due to card duplication, and the theft of authority due to smartphone theft, do not occur. Voice fingerprints are biometric characteristics derived from the unique vocal organ structure of the user's body, making forgery or duplication very difficult; furthermore, the present invention further enhances security by defending against recording playback attacks through an additional liveness detection module.
[0023] Seventh, the present invention realizes completely contactless access in which users make no physical contact whatsoever with input devices that are jointly touched by multiple people, such as keypads, fingerprint scanners, and card readers, thereby providing significant advantages over conventional contact-based access control methods in terms of hygiene, such as preventing the spread of infectious diseases. This is a particularly important effect during public health crises like pandemics and is a unique incidental effect of the present invention that conventional contact-based access control methods lack.
[0024] Eighth, the present invention dramatically improves recognition reliability in actual common entrance environments compared to conventional simple speech recognition applications by integrating and applying technical means such as microphone array beamforming, environmental noise modeling, de-reverb processing, speaker separation, and distance normalization to actively handle acoustic environmental challenges unique to common entrances, such as multi-passenger environments, noise environments specific to common entrances, and speech environments with variable distance and direction.
[0025] Ninth, the present invention enables integrated linkage with the complex infrastructure, such as transmitting a notification containing the time of entry, identified family members, and video data to the wall pad of the household upon permission of entry, and automatically calling an elevator, thereby providing additional effects such as family safety monitoring and improved convenience of movement beyond simple access control. Brief explanation of the drawing
[0026] FIG. 1 is a system configuration diagram schematically illustrating the overall configuration of a common entrance access control system based on voice fingerprint and generation number combination according to one embodiment of the present invention. FIG. 2 is a detailed block diagram illustrating the internal configuration of a voice input unit according to one embodiment of the present invention. FIG. 3 is a block diagram illustrating the internal configuration and signal processing path of a control unit according to an embodiment of the present invention. FIG. 4 is a structural diagram illustrating the data structure of a database unit according to one embodiment of the present invention. FIG. 5 is a flowchart illustrating the overall operation flow of an access control system according to one embodiment of the present invention. FIG. 6 is a conceptual diagram illustrating a processing flow in which generation number information and voice fingerprint information are extracted in parallel from a single spoken voice according to an embodiment of the present invention. FIG. 7 is a conceptual diagram illustrating the process of generating a dynamic pass key and comparing it with a registered voice fingerprint according to one embodiment of the present invention. FIG. 8 is a flowchart illustrating a reverse-binding denial scenario for space-speaker combined authority verification according to one embodiment of the present invention. FIG. 9 is a data structure diagram illustrating a voice fingerprint registration structure for multiple residents per household according to one embodiment of the present invention. FIG. 10 is a flowchart illustrating the branching flow of an access policy by resident type according to one embodiment of the present invention. FIG. 11 is a sequence diagram illustrating the flow of automatic visitor intercom switching in case of a mismatch in the voice fingerprint of a party resident according to one embodiment of the present invention. FIG. 12 is a conceptual diagram illustrating the normalization processing of Korean lake pronunciation variations according to one embodiment of the present invention. FIG. 13 is a flowchart illustrating a bidirectional correction flow based on voice fingerprint matching for adjacent lake misrecognition according to one embodiment of the present invention. FIG. 14 is a flowchart illustrating the immediate rejection flow for unregistered lakes within a complex according to one embodiment of the present invention. FIG. 15 is a block diagram illustrating the application structure of a common entrance environment noise model according to one embodiment of the present invention. FIG. 16 is a plan view illustrating a spatial sound source separation area by beamforming of a microphone array according to one embodiment of the present invention. FIG. 17 is a graph illustrating the effect of voice fingerprint energy normalization on the change in firing distance according to one embodiment of the present invention. FIG. 18 is a conceptual diagram illustrating speaker separation processing in a multi-speaker environment according to one embodiment of the present invention. FIG. 19 is a flowchart illustrating the operation flow of a liveness detection module according to one embodiment of the present invention. FIG. 20 is a graph illustrating a comparison between a natural variation pattern of lake ignition and a suspected recording playback pattern according to one embodiment of the present invention. FIG. 21 is a sequence diagram illustrating the operation flow of an arbitrary additional word challenge sequence according to an embodiment of the present invention. FIG. 22 is a timeline illustrating the temporal change of voice fingerprint adaptation update according to one embodiment of the present invention. FIG. 23 is a flowchart illustrating a progressive voice fingerprint registration flow according to one embodiment of the present invention. FIG. 24 is a sequence diagram of an automatic elevator call interlock that occurs simultaneously with entry permission according to one embodiment of the present invention. FIG. 25 is a sequence diagram of access notification linkage to a wall pad and a mobile terminal according to an embodiment of the present invention. FIG. 26 is a comparative diagram illustrating a comparison between the present invention and conventional keypad and card methods. FIG. 27 is a perspective view illustrating the appearance of a voice input unit and related components installed near a common entrance door according to one embodiment of the present invention. Specific details for implementing the invention
[0027] A common entrance access control system based on the combination of voice fingerprint and household number according to a preferred embodiment of the present invention will be described in detail below with reference to the attached drawings. In describing the present invention, detailed descriptions regarding related known technologies or configurations are omitted if it is determined that such detailed descriptions may obscure the essence of the present invention. Furthermore, the terms used in this specification are used to appropriately describe embodiments of the present invention and are not intended to limit the present invention.
[0028] Referring to FIG. 1, a voice fingerprint and household number combination-based common entrance access control system (10) according to one embodiment of the present invention is installed near the common entrance of a complex or building of a multi-family housing complex and is largely composed of a voice input unit (100) that receives a user's spoken voice, a control unit (200) that processes the received spoken voice to determine whether to allow entry, a database unit (300) that stores registered voice fingerprint data and related information for each household number, and an output / interlocking unit (400) that operates according to the entry permission decision of the control unit (200). In addition, the access control system (10) according to the present invention can be connected to communicate with external systems such as a complex integrated server (510), a wall pad (520) installed in each household, a resident's mobile terminal (530), and a complex elevator control system (540). A common entrance door (610) on which the access control system (10) according to the present invention is installed is equipped with an access door locking device (410), and a voice input unit (100), a guide display (620), and a video camera (630) are arranged in the vicinity thereof.
[0029] The voice input unit (100) is configured to receive voices spoken by a user attempting to enter or exit, installed near the common entrance door (610), and is configured to include a microphone array (110), a beamforming module (120), a noise preprocessing unit (130), a dereverb module (140), an analog-to-digital converter (150), and a distance normalization unit (160), as shown in FIG. 2. The microphone array (110) consists of a plurality of microphones (111, 112, 113, 114) arranged at a certain spatial interval. In this embodiment, an example is described in which four microphones are arranged in a square or linear arrangement, but the number of microphones and the arrangement form are not limited to this. The beamforming module (120) utilizes the phase difference of signals received from each microphone (111 to 114) of the microphone array (110) to strengthen the signal for sound sources within a certain distance range in front of the common entrance door (610), that is, within a distance range of about 0.3 meters to 1.5 meters where a normal user is expected to speak a number, and spatially suppress sound sources outside of that range. As a result, the voice of a user speaking from the front of the common entrance door (610) is strongly collected, while the voice of other pedestrians passing by the side or rear, or the noise of vehicles passing through the adjacent lane, is largely shielded.
[0030] The noise preprocessing unit (130) is configured to additionally remove environmental noise from the voice signal that has passed through the beamforming module (120). It selectively separates and removes the corresponding environmental noise component from the input signal by referring to an environmental noise model that has been pre-learned and stored in the environmental noise model DB (340) within the database unit (300)—e.g., a vehicle traffic noise model, an elevator operation sound model, a child crying sound model, a rain and wind noise model, a companion chatter model, etc. The de-reverb module (140) is configured to remove reverberation components generated by hard surfaces such as walls, ceilings, and floors of the common entrance. By performing inverse filtering based on the reverberation impulse response measured in advance according to the structure of the common entrance, it restores the signal to a form close to the user's direct speech. This reverberation removal processing provides a significant effect on improving recognition accuracy, especially in the common entrances of hotel-type officetels or luxury apartment complexes where the ceiling is high and the walls are smooth. The analog-to-digital converter (150) converts the processed analog voice signal into digital voice data and transmits it to the control unit (200), and the distance normalization unit (160) estimates the distance between the speaker and the microphone array (110) using the signal strength and arrival time difference measured at the microphone array (110), and then normalizes the volume and frequency characteristics that change according to the distance.
[0031] The control unit (200) is a core component of the present invention that processes digital voice data transmitted from the voice input unit (100) to determine whether entry is permitted, and is configured to include a voice recognition engine (210), a voice fingerprint extraction engine (220), a dynamic pass key generation unit (230), a matching comparison unit (240), a speaker separation module (250), a liveness detection module (260), a progressive registration module (270), and an adaptive update module (280), as illustrated in FIG. 3. The voice recognition engine (210) is configured to extract text spoken by a user, particularly generation number information, from input digital voice data, and includes a pronunciation variation normalization module (211) and an STT candidate generation unit (212). The voice fingerprint extraction engine (220) is configured to extract real-time voice fingerprint information representing the unique physical voice characteristics of the speaker from the same digital voice data.
[0032] One of the key features of the present invention is that, as illustrated in FIG. 6, a single digital voice data stream transmitted from the voice input unit (100) is simultaneously input in parallel to the voice recognition engine (210) and the voice fingerprint extraction engine (220). That is, when a user utters "No. 301" or "No. 301" once, the digital data of the utterance is branched, on the one hand, converted into generation number information in the text "301" by the voice recognition engine (210), and on the other hand, converted into a feature vector representing the unique voice characteristics of the speaker (e.g., a 39-dimensional feature vector such as MFCC, delta, delta-delta coefficients, etc.) by the voice fingerprint extraction engine (220). These two processes proceed simultaneously in time, and the results of the processing are all transmitted to the dynamic pass key generation unit (230).
[0033] As illustrated in FIG. 7, the dynamic pass key generation unit (230) is configured to generate a dynamic pass key by combining the generation number information extracted from the voice recognition engine (210) and the real-time voice fingerprint information extracted from the voice fingerprint extraction engine (220). This dynamic pass key is newly generated for every utterance, and since the voice fingerprint changes slightly depending on the voice micro-characteristics at the time of utterance even if the same resident utters the same apartment number, the dynamic pass key also changes slightly each time. Therefore, unlike conventional fixed passwords or fixed RFID card codes, the dynamic pass key of the present invention has the nature of a one-time key that is updated for every utterance, thereby further enhancing security. The matching comparison unit (240) searches for and retrieves the registered voice fingerprint data of the generation corresponding to the generation number information extracted from the database unit (300), and then compares and determines whether the real-time voice fingerprint within the dynamic pass key matches the registered voice fingerprint within a preset threshold.
[0034] As illustrated in FIG. 4, the database section (300) includes a household number-voice fingerprint matching table (310) that stores a household number and a voice fingerprint of a legitimate resident of the corresponding household by matching them, a resident type / time zone information table (320) that stores information on the type of each resident (household head, accompanying resident, minor, short-term registrant, etc.) and the time zone for entry and exit, an auxiliary vocabulary pool (330) that stores a list of auxiliary vocabulary for each user that can be used as a challenge word when detecting liveness, and an environment noise model DB (340) that stores an environment noise model of the common entrance. As illustrated in FIG. 9, the household number-voice fingerprint matching table (310) is configured so that multiple resident voice fingerprints can be registered for a single household number, for example, for household 301, the voice fingerprint of the father who is the household head, the voice fingerprint of the mother who is the spouse, the voice fingerprint of child 1, and the voice fingerprint of child 2 can all be registered, and resident type information is matched and stored for each voice fingerprint.
[0035] The output / interlocking unit (400) is configured to perform actual door opening and auxiliary interlocking operations according to the entry permission decision of the control unit (200), and includes an entry door locking device control unit that unlocks the locking device (410) of the common entrance door (610), an elevator calling unit (420) that transmits a signal to the elevator control system (540) to automatically call the elevator to the floor of the unit where entry is permitted, and a wall pad notification unit (430) that transmits the entry fact, identified resident information, and video data to the wall pad (520) and registered mobile terminal (530) of the unit.
[0036] Referring to FIG. 5, the overall operation flow of an access control system (10) according to an embodiment of the present invention is described as follows. When a user attempting to enter approaches the vicinity of the common entrance door (610) and speaks the number of their household, such as "301" or "301," the microphone array (110) of the voice input unit (100) receives this. The beamforming module (120) strengthens the signal in the direction of the speaker and suppresses lateral noise sources, the noise preprocessing unit (130) removes environmental noise components such as vehicle traffic sounds, elevator sounds, and children crying sounds by referring to the noise model of the environmental noise model DB (340), and the de-reverb module (140) removes reverberation components caused by the structure of the common entrance. The processed voice signal is converted into digital data by the analog-to-digital converter (150), and the signal strength and frequency characteristics according to the speaking distance are normalized by the distance normalization unit (160).
[0037] Normalized digital voice data is transmitted to the control unit (200), and within the control unit (200), the same data stream is simultaneously input in parallel to the voice recognition engine (210) and the voice fingerprint extraction engine (220). The voice recognition engine (210) converts various pronunciation variations such as "301," "301," "301," and "3-0-1" into standardized generation number text "301" through the pronunciation variation normalization module (211), and the STT candidate generation unit (212) can generate multiple number candidates based on the reliability of the voice recognition result. The voice fingerprint extraction engine (220) generates real-time voice fingerprint information by extracting acoustic feature vectors derived from the shape of the speaker's vocal organs and vocal habits from the same voice data.
[0038] The dynamic pass key generation unit (230) generates a dynamic pass key by combining the extracted household number information "301" with real-time voice fingerprint information, and the matching comparison unit (240) retrieves all resident voice fingerprints registered for household number "301" from the household number-voice fingerprint matching table (310). The matching comparison unit (240) calculates the similarity between the real-time voice fingerprint and each of the registered voice fingerprints to determine whether there is a registered voice fingerprint that shows a similarity greater than or equal to a preset threshold, and if a matching registered voice fingerprint exists, it further checks the resident's type information and allowed time zone information in the resident type / time zone information table (320). If the time zone policy is not violated, the matching comparison unit (240) outputs an entry permission signal to the output / linkage unit (400), and as the door lock (410) is released and the common entrance door (610) is opened, the elevator call unit (420) transmits an automatic elevator call signal to the floor of the corresponding household, and the wall pad notification unit (430) transmits an entry notification including identified resident information, entry time, and video data captured by the video camera (630) to the wall pad (520) and mobile terminal (530) of the corresponding household.
[0039] The space-speaker combined authority verification, i.e., the reverse-binding mechanism, which is a core technical feature of the present invention, will be explained in detail with reference to FIG. 8. In the first scenario, it is assumed that the legitimate resident of Room 301 speaks "Room 301" at the common entrance. In this case, the voice recognition engine (210) extracts the unit number "301," the voice fingerprint extraction engine (220) extracts the voice fingerprint of the resident, and the matching comparison unit (240) finds a match among the voice fingerprints registered for "301" in the unit number-voice fingerprint matching table (310) and allows entry. In the second scenario, it is assumed that the legitimate resident of Room 301 speaks "Room 501" either accidentally or intentionally. In this case, the voice recognition engine (210) extracts the generation number "501", the voice fingerprint extraction engine (220) extracts the speaker's voice fingerprint (which is the fingerprint registered in Room 301), and the matching comparison unit (240) searches for voice fingerprints registered in "501" in the generation number-voice fingerprint matching table (310), but since none of them match the speaker's real-time voice fingerprint, entry is denied. Even though the speaker's voice fingerprint is registered throughout the system, entry is not permitted because there is no resident authority for the spoken room number, and this is the core of the reverse-binding effect of the present invention. As a third scenario, assume a case where an unregistered external visitor speaks an arbitrary room number, such as "Room 301". In this case, the matching comparison unit (240) compares the voice fingerprints registered in the spoken "301" with the speaker's real-time voice fingerprint, but since there is no match, entry is denied and the system switches to an automatic intercom call mode as described below.
[0040] An embodiment regarding the processing of multiple residents per household is described with reference to FIGS. 9 and 10. The household number-voice fingerprint matching table (310) of the database unit (300) according to the present invention is configured so that multiple resident voice fingerprints can be registered for a single household number. For example, if a four-person family consisting of a father, mother, child 1 (high school student), and child 2 (elementary school student) resides in unit 301, four voice fingerprints are registered for the household number "301", and information from the resident type / time zone information table (320) is matched to each voice fingerprint. For example, the father and mother are allowed free entry and exit 24 hours a day as heads of household or accompanying residents, child 1 is a minor but a high school student and is allowed free entry and exit before 23:00, and child 2 is a minor and is allowed free entry and exit before 21:00, but a policy may be set so that entry and exit are restricted or an accompanying person is required during other times. The matching comparison unit (240) identifies which resident has a matching real-time voice fingerprint, and then applies an entry policy based on the current time by referring to the resident type / time zone information table (320). For example, if Child 2 says "Room 301" at 22:00, the voice fingerprint matches, but since it violates the time zone policy, the system does not immediately refuse entry but instead notifies the wall pad (520) of the relevant unit of the child's attempt to enter and asks for remote approval from the guardian. This branching of policies by resident type clearly demonstrates the effectiveness of the present invention in flexibly responding to various residential forms in multi-family housing.
[0041] An embodiment regarding automatic intercom switching for visitors is described with reference to FIG. 11. When an external visitor, such as a delivery driver, a courier, or a relative, arrives at the main entrance and speaks the number of the unit they wish to visit, such as "Room 301," the matching comparison unit (240) compares the voice fingerprints registered for the spoken "301" with the visitor's real-time voice fingerprints, but does not find a match. At this time, the conventional access control system simply outputs a rejection response, but the system (10) according to the present invention first checks whether the unit number "301" extracted from the voice recognition engine (210) is a valid unit registered in the database unit (300), and if it is a valid unit, immediately sends an automatic call to the wall pad (520) of Room 301 via the complex integrated server (510). The video of the visitor captured by the video camera (630) is displayed on the wall pad (520) of Room 301, and at the same time, a visit notification can be sent to the mobile terminal (530) of the resident of Room 301. The resident can check the visitor through the wall pad (520) or the mobile terminal (530) and input remote approval or rejection, and if remote approval is input, the door lock (410) is released and the visitor is allowed to enter. Thus, the visitor can also proceed with the remote approval process by the resident naturally without having to press a separate intercom call button or operate a keypad, simply by speaking the room number.
[0042] Referring to Fig. 12, an example of normalization of phonetic variations for Korean house numbers is described. In Korean, the same generation number "Ho 301" can be pronounced in various forms depending on the speaker's age, region of origin, and speech habits. For example, a speaker using standard Sino-Korean pronunciation may pronounce it as "Sambaeilho," a speaker conscious of the digits may pronounce it as "Samgongilho" or "Samyongilho," a young generation speaker who mixes English numbers may pronounce it as "Threezeroone" or "Threeowon," and a speaker who clearly separates the digits may pronounce it as "Sam-gong-il-ho" or "Sam-yeong-il-ho." Additionally, "Ho 200" may be pronounced as "Ibaekho," "Igonggongho," or "Ibaegongho." The pronunciation variation normalization module (211) within the speech recognition engine (210) holds a dictionary of variations of these Korean number pronunciations in advance and maps all of the above pronunciation variations to the standardized generation number text "301" or "200". The pronunciation variation normalization module (211) is configured to enable reasonable normalization even for new pronunciation variations that are not predefined in the dictionary by utilizing a statistical model learned from various pronunciation cases along with rule-based mapping based on Korean number notation rules. This Korean-specific processing is not handled in conventional foreign speech authentication technologies designed based on English.
[0043] Referring to FIG. 13, an embodiment regarding bidirectional correction based on voice fingerprint matching for misrecognition of adjacent numbers is described. In Korean, some adjacent numbers have very similar pronunciations, so there are cases where it is difficult for the voice recognition engine (210) to determine the correct number on its own. For example, "301" and "3001," "201" and "210," and "1101" and "1110" correspond to such adjacent pronunciation numbers. In the present invention, in such cases, the STT candidate generation unit (212) is configured to generate a plurality of number candidates whose reliability exceeds a threshold, and for example, for the utterance "301," a candidate "301" with a reliability of 0.55 and a candidate "3001" with a reliability of 0.45 can be output simultaneously. The bidirectional combined re-ranking unit (241) within the matching comparison unit (240) queries the registered voice fingerprints in the database unit (300) for each of these multiple lake number candidates and calculates the similarity with the real-time voice fingerprint extracted by the voice fingerprint extraction engine (220) for each candidate. That is, the highest similarity with the registered voice fingerprints of "301" and the highest similarity with the registered voice fingerprints of "3001" are calculated, respectively, and the candidate lake number with significantly higher voice fingerprint similarity is confirmed as the final lake number. This is a new method for confirming the lake number by combining two independent information sources—the reliability of the voice recognition result and the voice fingerprint matching similarity—bidirectionally, and the recognition accuracy is significantly improved compared to the conventional unidirectional ASR + speaker verification combined method.
[0044] Referring to FIG. 14, an example of immediate refusal for unregistered units within the complex is described. When the voice recognition engine (210) extracts a unit number from a user's speech, and the unit number is a unit that does not actually exist within the complex (e.g., in a case where "3501" is spoken even though the complex only has up to 30 floors), or is an unregistered vacant unit, the matching comparison unit (240) does not perform any voice fingerprint matching processing and immediately outputs an entry refusal response, and displays a notice through the guidance display (620) stating, "The unit in question is an unregistered unit." This immediately blocks the burden of voice fingerprint matching processing from invalid unit inputs, thereby saving system resources, and at the same time allows for monitoring potential intrusion attempts by separately logging unregistered unit attempt patterns.
[0045] An embodiment regarding noise processing in a common entrance environment and microphone array beamforming is described with reference to FIGS. 15 and 16. Unlike indoor environments where general voice recognition is used, the common entrance environment of a multi-family housing complex is exposed to a wide variety of strong noise sources. Common noises in the common entrance environment include vehicle traffic sounds from adjacent parking lots or complex entrance roads, elevator operation sounds within the complex, crying or shouting of children, raindrop sounds and wind sounds during rain, chatter from companions or subsequent passersby, and complex broadcasts heard from a distance. In the present invention, an environmental noise model for these noise sources unique to the common entrance is pre-trained and stored in an environmental noise model DB (340), and a noise preprocessing unit (130) refers to this to selectively separate and remove the corresponding noise components from the input signal. The environmental noise model can be further trained and adapted to suit the environmental characteristics of each complex where the access control system (10) according to the present invention is installed. Meanwhile, as illustrated in FIG. 16, the beamforming module (120) of the microphone array (110) forms a beam area approximately ±30 degrees to the left and right and a distance of approximately 0.3 to 1.5 meters centered on the front of the common entrance door (610), and the voice of a speaker located within this beam area is selectively enhanced, while sound sources outside of it are spatially suppressed. Thus, when a user speaks their apartment number from the front of the common entrance door, the sound of other residents talking on the side or vehicle noise from the rear lane is naturally shielded.
[0046] Referring to FIG. 17, an embodiment regarding voice fingerprint normalization according to changes in speech distance is described. Even if the same user speaks the same number, the received volume and frequency characteristics change depending on the distance from the microphone array (110). For example, there is a significant difference in the received volume when the user speaks from a close distance of about 0.3 meters compared to when the user speaks from a far distance of 1.5 meters, and also, the degree of attenuation of high-frequency components varies due to the air absorption effect according to distance. The distance normalization unit (160) estimates the speech distance from the signal strength of the microphone array (110) and the difference in arrival time, and then corrects the volume and frequency variations according to distance to normalize the signal characteristics into standardized signals. As a result, the accuracy of voice fingerprint matching can be stably maintained even if the user does not speak from the same location every time, and the reliability of comparison between the voice fingerprint measured at the registration stage and the voice fingerprint measured at the authentication stage is improved.
[0047] Referring to FIG. 18, an embodiment regarding speaker separation in a multi-speaker environment is described. In the common entrance of a multi-unit apartment building, situations frequently occur where the voices of a companion, an adjacent passerby, or another resident talking on the phone nearby are collected together while one user speaks their apartment number. For example, this occurs when a parent speaks their apartment number while returning home with their child, or when the voice of a passerby talking on the phone in the adjacent building is captured together. In such cases, the speaker separation module (250) within the control unit (200) operates to separate the voices of multiple speakers who speak simultaneously by applying a speaker diarization algorithm to the input digital voice data. Among the separated voices of multiple speakers, the speaker separation module (250) selects only the voice of the main speaker who spoke the text corresponding to the apartment number and transmits it to the voice recognition engine (210) and the voice fingerprint extraction engine (220). This prevents the voice of a companion or a passerby from affecting the matching.
[0048] An embodiment regarding liveness detection is described with reference to FIGS. 19 and 20. A replay attack, which involves pre-recording and playing back the voice of a legitimate resident, is known as a representative attack method against a voice authentication system. The liveness detection module (260) within the control unit (200) according to the present invention is configured to detect and defend against such replay attacks and includes a speech natural variation verification unit (261) and a challenge vocabulary selection unit (262). The speech natural variation verification unit (261) takes into account that even when the same resident speaks the same apartment number, there are natural variations in human speech—such as minute differences in speech speed, slight fluctuations in pitch, and differences in pause times between phonemes within a word—and compares the acoustic characteristics of the current speech with the past speech history of the same resident. If all acoustic features of the current utterance match abnormally accurately with a specific utterance in the past—that is, if identity is observed to an extent that exceeds the range of natural human variation—the liveness detection module (260) determines this as a recording playback attack and denies entry regardless of whether the voice fingerprint matches. Additionally, the challenge vocabulary selection unit (262) requests an additional utterance from the user by selecting a random word from the auxiliary vocabulary pool (330) within the database unit (300) when a suspected liveness situation occurs or according to a pre-set security policy. The auxiliary vocabulary pool (330) stores words that the user has previously registered, and by verifying the possibility of responding to the randomly selected challenge vocabulary and the voice fingerprint of the additional utterance together, the defense against recording playback attacks is further strengthened.
[0049] An embodiment regarding progressive voice fingerprint registration and adaptive updating is described with reference to FIGS. 22 and 23. Conventional speaker recognition systems generally collect multiple speech samples in bulk through a separate registration procedure when registering a new user; however, the progressive registration module (270) according to the present invention accumulates and collects the voice of the room number spoken by a new resident each time they enter and exit, thereby progressively refining the voice fingerprint model. That is, basic registration is completed by providing only a single, simple voice sample during the initial registration of a new resident, and the voice fingerprint is naturally reinforced during the subsequent daily entry and exit process. This significantly reduces the burden on the user during the registration stage, and at the same time, the diversity of speech in the actual usage environment is naturally reflected in the voice fingerprint model, thereby improving recognition accuracy. Additionally, the adaptive updating module (280) accumulates speech voice data for which entry and exit are permitted based on voice fingerprint matching, reflecting the fact that the user's voice changes over time, and periodically adaptively updates the registered voice fingerprint data for each user. This absorbs changes in vocal range due to puberty in growing minors, a downward shift in vocal range due to aging, and voice variations caused by temporary changes in condition such as a cold, ensuring that recognition accuracy for the same user is maintained over time.
[0050] An embodiment regarding automatic elevator calling and wall pad entry notification is described with reference to FIGS. 24 and 25. In the access control system (10) according to the present invention, when the matching comparison unit (240) makes a decision to allow entry, the elevator calling unit (420) transmits a floor call signal for the corresponding unit to the elevator control system (540) in parallel with the output of a door locking device (410) release signal. For example, when the resident of unit 301 is authenticated and entry is allowed, the elevator on the 1st floor is already automatically called to the 3rd floor at the time the door opens, so the elevator is waiting for the resident by the time they arrive in front of the elevator. Convenience is provided so that the resident, carrying luggage in both hands, does not need to operate a separate elevator call button. At the same time, the wall pad notification unit (430) transmits an entry notification to the wall pad (520) and mobile terminal (530) of the corresponding unit, including a video of the resident captured by the video camera (630), identified family member information (e.g., "Child 1 has entered"), and the time of entry. This allows other family members to check a family member's return home or outing in real time, making it particularly useful for monitoring child safety in households with minor children.
[0051] The voice fingerprint and generation number combination-based common entrance access control system (10) according to the present invention described above may be implemented with various modifications within the scope that some components or functions are included in the technical concept of the present invention. For example, the number and arrangement of microphones in the microphone array (110), specific algorithms of the voice recognition engine (210) and the voice fingerprint extraction engine (220), the data structure of the database unit (300), and the learning method of the environmental noise model DB (340) may be selected in various ways depending on the application environment. In addition, the present invention may be extended to fields such as access control of guest room areas in hotels, access control by floor in office buildings, and access control by zone in industrial facilities, in addition to the common entrance of multi-family housing, where the same technical concept can be applied. Each component of the system according to the present invention may be implemented as hardware, software, or a combination thereof, and each module within the control unit (200) may be executed integrally by a single processor or distributed and executed by multiple processors. The database section (300) may be implemented locally within the access control system (10), or it may be implemented remotely on the integrated server (510) side so that the access control system (10) accesses it through a communication network.
[0052] As a specific application example of the access control system (10) according to the present invention, consider a case where the system is applied to a large apartment complex with a scale of 1,000 households. There are multiple buildings within the complex, and a common entrance door (610) is installed in each building. In each common entrance, a voice input unit (100), a video camera (630), and a guide display (620) according to the present invention are installed. A database unit (300) in which the voice fingerprints of residents of all 1,000 households are registered is operated in the complex integrated server (510), and the access control system (10) of each building is connected to the complex integrated server (510) through a communication network. When a resident arrives at the common entrance of their building and speaks their apartment number, the registered voice fingerprint of the household corresponding to the spoken number is retrieved from among the households registered in that building. Since an immediate 1:minority matching based on the apartment number is performed rather than a 1:N matching, the authentication processing time is maintained constant even if the complex is large. Meanwhile, if a resident of another building who has come to the wrong building calls out their own apartment number—for example, if a resident of building 102 accidentally comes to the common entrance of building 103 and calls out their apartment number—matching fails because the speaker is not registered in the database of that building, and the system displays a notice through the guidance display (620) saying "This is not an apartment number registered in this building" and then switches to automatic intercom call mode, or if it is confirmed that the speaker is a resident of another building by querying the integrated server (510), it can be processed in guest mode.
[0053] As another application example of the access control system (10) according to the present invention, consider a case where it is utilized for the safety management of family members. For instance, when a child of a dual-income couple returns home from school and passes through the common entrance by speaking their apartment number, the system (10) according to the present invention identifies the child using a voice fingerprint and immediately transmits the time of entry and video to the parent's mobile terminal (530). The parent can confirm the child's safe return even from their workplace and can respond immediately if the child returns home later than usual or if no entry notification is received. Furthermore, if a minor child's nighttime outing falls outside the permitted time zone, the system immediately transmits an approval request to the parent's mobile terminal (530), and the system can be operated so that entry is permitted only upon remote approval from the parent. This family safety monitoring function is an additional effect of the present invention that is not provided by conventional simple access control systems.
[0054] In addition, the access control system (10) according to the present invention provides an effect in terms of personal information protection in that it can ensure access safety without collecting information on the movement path of the resident. The voice fingerprint itself is an abstract acoustic feature vector that does not directly expose the user's identity or appearance, and unlike conventional facial recognition systems that store the resident's face image, the user's visual identity is not stored in the system. The voice fingerprint model is appropriately encrypted and stored in the database unit (300), and since it is very difficult to restore the user's original voice from the model itself, the risk of exposure of the user's identity is significantly lower compared to facial images even if data leakage occurs.
[0055] In addition, the access control system (10) according to the present invention can significantly reduce the rate of security incidents compared to conventional keypad or card methods, in that the visible movement path of a resident is not tracked and even if a key card or mobile terminal is lost and acquired by an unrelated third party, entry to the common entrance is impossible with only that. Security incidents that frequently occurred in conventional access control systems, such as password exposure, loss of access card, and duplication of access card, are completely prevented from occurring in the present invention. This enhanced security effect is particularly significant in an environment where a large number of residents live, such as a multi-family housing complex.
[0056] Although preferred embodiments of the present invention have been described above, the present invention is not limited to the above embodiments, and it is evident that a person skilled in the art can derive various modifications and equivalent other embodiments within the scope of the technical concept of the present invention. Accordingly, the true scope of protection of the present invention should be determined by the appended claims. Explanation of the symbols
[0057] 10: Common entrance access control system based on voice, fingerprint, and household number combination 100: Voice input section 110: Microphone array 111, 112, 113, 114: Microphone 120: Beamforming Module 130: Noise Preprocessing Unit 140: De-reverb Module 150: Analog-to-Digital Converter 160: Distance Normalization 200: Control unit 210: Voice recognition engine 211: Pronunciation Variation Normalization Module 212: STT Candidate Generation Section 220: Voice fingerprint extraction engine 230: Dynamic pass-through key generation unit 240: Matching Comparison Section 241: Bidirectional combined reranking section 250: Speaker Separation Module 260: Liveness Detection Module 261: Ignition Natural Variation Verification Unit 262: Challenge Vocabulary Selection Section 270: Incremental Enrollment Module 280: Adaptive Update Module 300: Database Department 310: Household Number-Voice Fingerprint Matching Table 320: Resident Type / Time Zone Information Table 330: Auxiliary Vocabulary Pool 340: Environmental Noise Model DB 400: Output / Interlocking Unit 410: Door lock 420: Elevator Call Center 430: Wallpad Notification Unit 510: Complex Integrated Server 520: Wallpad 530: Mobile terminal 540: Elevator Control System 610: Common entrance door 620: Information display 630: Video camera
Claims
Claim 1 A system for controlling access to a common entrance of a multi-family housing complex using only the voice of a user, comprising: a voice input unit for receiving the voice of a user including the household number of the destination to be entered; a database unit in which unique voice fingerprint data of a legitimate user is matched and stored for each household number; and a control unit linked to the voice input unit and the database unit, wherein the control unit extracts household number information through voice recognition from the voice data of the user received through the voice input unit, extracts real-time voice fingerprint information by analyzing the frequency characteristics of the voice data, searches for unique voice fingerprint data of a registered user corresponding to the extracted household number information in the database unit, generates a dynamic pass key by mutually combining the household number information obtained by voice recognition and the analyzed real-time voice fingerprint information, and compares the generated dynamic pass key with the unique voice fingerprint data of the registered user searched in the database unit, and outputs a control signal to open the common entrance door if they match. Claim 2 A voice fingerprint and generation number combination-based common entrance access control system according to claim 1, wherein the control unit receives a single voice signal stream spoken by a user and simultaneously inputs it in parallel to a voice recognition algorithm that analyzes text content and a voice fingerprint extraction algorithm that analyzes individual biometric features to generate the dynamic pass key in real time. Claim 3 A common entrance access control system based on the combination of voice fingerprint and household number, wherein, in claim 1, the voice input unit includes a directional microphone module that removes external environmental noise, and the control unit extracts a household number and a voice fingerprint after undergoing a preprocessing process that filters ambient noise from the received voice data. Claim 4 A common entrance access control system based on the combination of voice fingerprint and household number, wherein, in claim 1, the voice of the user is a voice referring to the household number of the household the user wishes to enter, and the voice itself functions as an identity claim asserting access rights to the household, and the control unit verifies the authenticity of the identity claim by whether the voice fingerprint matches. Claim 5 A voice fingerprint and generation number combination-based common entrance access control system according to claim 1, wherein the control unit refuses to open the door even if the extracted real-time voice fingerprint information matches the registered voice fingerprint of another generation in the database unit, but does not match the registered voice fingerprint of the generation corresponding to the extracted generation number information. Claim 6 A voice fingerprint and household number combination-based common entrance access control system according to claim 1, wherein the database unit registers and stores unique voice fingerprint data for each of a plurality of residents living in a household corresponding to a single household number, and the control unit opens the door when the extracted real-time voice fingerprint information matches any one of the plurality of voice fingerprints registered for the extracted household number. Claim 7 A common entrance access control system based on the combination of voice fingerprints and household numbers, characterized in that, in claim 6, the database unit stores resident type information and access permission time zone information by matching them for each of the plurality of resident voice fingerprints, and the control unit determines whether to open the door according to the resident type and permission time zone matched to the matched voice fingerprint. Claim 8 A common entrance access control system based on the combination of voice fingerprint and household number, wherein, in claim 1, the control unit, when the extracted real-time voice fingerprint information does not match any voice fingerprint registered in the extracted household number information, sends an automatic call to the wall pad or communication terminal of the household corresponding to the extracted household number and switches to a remote access approval procedure by the resident of the household. Claim 9 A voice fingerprint and household number combination-based common entrance access control system according to claim 1, wherein the control unit, in extracting household number information from the spoken voice, includes a pronunciation variation normalization module that normalizes variations of Korean number pronunciation, namely Sino-Korean pronunciation, Sino-Korean digit omission type, English number mixed type, and digit-separated pronunciation, and matches them to the same household number. Claim 10 A voice fingerprint and generation number combination-based common entrance access control system according to claim 1, wherein the control unit, when a plurality of generation number candidates are derived by the voice recognition, compares the extracted real-time voice fingerprint information with the registered voice fingerprint of each candidate generation and determines the generation number candidate showing the highest voice fingerprint similarity as the final generation number. Claim 11 A voice fingerprint and household number combination-based common entrance access control system according to claim 1, wherein the control unit outputs a rejection response immediately without performing a voice fingerprint matching process when the extracted household number is a household number not registered in the database unit. Claim 12 In paragraph 3, the noise filtering preprocessing process of the control unit is characterized by including an environment-adaptive noise model that selectively separates the corresponding environmental noise from the user's spoken voice by pre-learning acoustic models of noise unique to the environment of a common entrance of a multi-unit dwelling, namely, vehicle traffic noise, elevator operation sound, child crying sound, conversation sound of a companion, and rain and wind noise. Claim 13 A voice fingerprint and household number combination-based common entrance access control system, wherein, in paragraph 3, the voice input unit further includes a de-reverb processing module for compensating for reverberation caused by the wall and ceiling structures of the common entrance, and the voice fingerprint extraction is performed based on reverberation-compensated voice data. Claim 14 A voice fingerprint and generation number combination-based common entrance access control system according to claim 1, wherein the control unit extracts a voice fingerprint after normalizing the energy level of the spoken voice in order to compensate for changes in volume and frequency characteristics due to changes in the distance between the voice input unit and the speaker. Claim 15 A common entrance access control system based on the combination of voice fingerprints and household numbers, wherein, in the first paragraph, the control unit applies a speaker separation algorithm to separate only the voice of the main speaker uttering the household number when multiple speaker voices are simultaneously received through the voice input unit, and extracts a household number and a voice fingerprint from the separated main speaker's voice. Claim 16 A voice fingerprint and household number combination-based common entrance access control system according to claim 3, wherein the voice input unit comprises a microphone array composed of a plurality of microphones, and the microphone array performs beamforming on sound sources within a certain distance range in front of the common entrance door to spatially separate and suppress sound sources outside that range. Claim 17 A voice fingerprint and generation number combination-based common entrance access control system according to claim 1, wherein the control unit further includes a liveness detection module that performs recording playback attack detection on the received spoken voice, and is characterized by refusing to open the door regardless of whether the voice fingerprint matches when recording playback is suspected as a result of the liveness detection. Claim 18 A voice fingerprint and generation number combination-based common entrance access control system according to claim 17, wherein the liveness detection module compares the user's past speech history for the same generation number with the speech rate, pitch, and pause pattern of the current speech, and determines that it is a recording playback when the degree of agreement is high enough to exceed the range of natural variation of human speech. Claim 19 A common entrance access control system based on the combination of voice fingerprints and generation numbers, wherein, in claim 17, the control unit requests an additional utterance from the user by selecting a random word from a user-specific auxiliary vocabulary pool registered in the database unit when a pre-set security policy or liveness suspicion situation occurs, and verifies the text and voice fingerprints extracted from the additional utterance together. Claim 20 A voice fingerprint and generation number combination-based common entrance access control system according to claim 1, wherein the control unit accumulates and collects speech voice data for which entry is permitted by voice fingerprint matching, and periodically adaptively updates the registered voice fingerprint data to reflect temporal changes in the user's voice. Claim 21 A voice fingerprint and household number combination-based common entrance access control system according to claim 1, further comprising a progressive registration module that, when registering a voice fingerprint of a new resident in the database, accumulates and collects the voice of the household number spoken by the new resident each time they enter and exit to progressively build a voice fingerprint model. Claim 22 A voice fingerprint and household number combination-based common entrance access control system according to claim 1, wherein the control unit transmits a control signal to the elevator control system to automatically call the elevator to the floor corresponding to the extracted household number simultaneously with the output of a door opening control signal. Claim 23 A voice fingerprint and household number combination-based common entrance access control system according to claim 1, wherein the control unit transmits an access notification including the time of entry, resident information identified by speech, and video data to a wall pad or registered mobile terminal of a household corresponding to the extracted household number simultaneously with the opening of the entrance door. Claim 24 A method for controlling access to a common entrance of a multi-family housing complex using only the voice of a user, comprising: (a) receiving a voice of voice including the unit number of the unit the user wishes to enter through a voice input unit; (b) extracting unit number information from the received voice of voice through voice recognition; (c) extracting unique voice fingerprint information of the speaker from the same voice of voice in parallel with step (b); (d) searching for a registered voice fingerprint of the corresponding unit from a database of registered voice fingerprints by unit number using the extracted unit number information as an index key; (e) comparing the searched registered voice fingerprint with the extracted real-time voice fingerprint; and (f) opening the common entrance door only when the comparison result matches; a method for controlling access to a common entrance based on the combination of voice fingerprint and unit number. Claim 25 A computer-readable recording medium having a program for executing the method of paragraph 24 on a computer.