Safety box unlocking method based on facial recognition
By analyzing the dynamic facial muscle features of users reading random sentences, and combining static facial features with dynamic pronunciation features to form a multi-dimensional authentication system, the problem of safe facial recognition being vulnerable to static forgery attacks has been solved, thereby improving security and convenience, and supporting remote unlocking and local verification.
Patent Information
- Application Number
- CN202511385592.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-09-26
AI Technical Summary
Existing facial recognition technology for safes is easily fooled by photos, videos, or 3D masks, posing a security risk and making it difficult to effectively prevent unauthorized unlocking.
By analyzing the dynamic facial muscle features of users when reading random sentences, a multi-dimensional authentication system is formed by combining static facial features with dynamic pronunciation features. A clustering algorithm is used to construct a pronunciation set and generate a semantically fluent pronunciation template. Combined with optical flow method to capture dynamic features, a sentence password is generated in real time for authentication.
It effectively resists static deception attacks using photos, videos, or 3D masks, increasing the difficulty of cracking, reducing false rejection rates, ensuring security and convenience, conforming to modern human-computer interaction trends, and supporting remote unlocking and local verification.
Smart Images

Figure CN120977038A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of safe unlocking, in particular to a safe unlocking method based on facial recognition. BACKGROUND
[0002] In today's society, with the continuous development of technology, as an important safety protection device, the unlocking method of the safe is also evolving. The unlocking method of the safe in the prior art is commonly to use facial recognition technology.
[0003] The reason why facial recognition technology is widely used is because of its remarkable convenience. Users do not need to carry additional keys or cards, and only need to stand in front of the safe and perform facial scanning through the camera. The system can quickly identify and complete the unlocking operation, greatly improving the use efficiency, especially in some scenarios where the safe needs to be frequently opened. The convenience is more prominent.
[0004] However, although facial recognition technology performs excellently in terms of convenience, it has a security risk that cannot be ignored. In actual application scenarios, facial recognition can be deceived by photos, videos or 3D masks. For example, a criminal can obtain a user's photo, process it using technical means to make it pass the verification of the facial recognition system, or make a highly simulated video or 3D mask to simulate the user's facial features, thereby bypassing the security detection and illegally obtaining the items in the safe. This security risk poses a serious threat to the safe storage of important items in the safe, and also makes the safe unlocking method relying solely on facial recognition technology face many challenges in actual application, and a more secure and reliable unlocking technology is needed to make up for its shortcomings. SUMMARY
[0005] The purpose of the present application is to provide a safe unlocking method based on facial recognition, which solves the following technical problems.
[0006] The purpose of the present application can be achieved by the following technical solutions: A safe unlocking method based on facial recognition, comprising the following steps: Step S1: setting a plurality of pronunciation combinations, obtaining facial muscle features of each pronunciation combination; clustering each pronunciation combination according to the facial muscle features to obtain a plurality of pronunciation sets, and determining standard facial features of each pronunciation set, obtaining pronunciation combinations corresponding to the standard facial features, denoted as typical pronunciation combinations; and obtaining the association features between each pronunciation combination in the pronunciation set and the typical pronunciation combination; Step S2: setting a pronunciation template according to each typical pronunciation combination, inputting a template pronunciation video of the user reading the pronunciation template, and obtaining facial pronunciation template features of the user according to the template pronunciation video; Step S3: when the user unlocks the safe, the safe randomly generates a sentence password, acquires a face video of the user reading the sentence password, denoted as a current face video, and obtains a current face feature sequence according to the current face video; the safe generates a current template feature sequence of the sentence password according to the face pronunciation template feature; Compare the current template feature sequence and the current face feature sequence to determine whether the authentication is passed, and if the authentication is passed, the safe is unlocked.
[0007] As a further scheme of the present application, the obtaining process of the face muscle features of the pronunciation combination comprises: Randomly combine the initial consonants and final consonants two by two to obtain a plurality of pronunciation combinations, and select a plurality of experimental personnel to acquire face videos of the experimental personnel reading each pronunciation combination, and extract face muscle features in each face video.
[0008] As a further scheme of the present application, the extracting process of the face muscle features of the pronunciation combination further comprises: Divide the face video of the pronunciation combination into a plurality of video frames, locate the face range of the experimental personnel in the video frames based on a face detection algorithm, and mark core feature points on the face range through a face key point detection model, the core feature points including an upper lip edge, a lower lip edge, a mouth corner, a lip bead, a lower jaw line and a chin tip; based on all the core feature points, a sub-region is segmented through a pixel mask technology, the sub-region including a lip, a lower jaw, a cheek, a forehead and an ear; all the sub-regions are divided into a pronunciation-related region and a pronunciation-unrelated region, and all the pronunciation-unrelated regions are excluded, the pronunciation-related region including the lip, the lower jaw and the cheek, and the pronunciation-unrelated region including the forehead and the ear; Mark the core feature points in the pronunciation-related region as key points, and obtain dynamic features of the key points in the face video by sequentially obtaining pixel positions of the key points in each video frame through an optical flow method, the dynamic features including a moving trajectory, a moving acceleration and a displacement peak value; Filter key frames from all the video frames of the face video, and obtain static features based on the pixel positions of the key points in the key frames, the static features including an opening degree, a round-lip degree, a mouth corner angle, an angle between a lower jaw line and a horizontal line, and a convexity value of a cheekbone region; Convert the dynamic features and the static features into numerical vectors of a uniform dimension, and mark the numerical vectors as the face muscle features of the pronunciation combination.
[0009] As a further scheme of the present application, the filtering process of the key frames comprises: In all video frames of the face video, the first video frame is recorded as a start frame, and the remaining video frames are recorded as remaining frames; the pixel positions of the key points in the start frame are obtained and recorded as start positions, and the pixel positions of the key points in the remaining frames are sequentially obtained and recorded as remaining positions; the displacement amounts of the key points in the remaining frames compared with the key points in the start frame are obtained according to the start positions and the remaining positions, the displacement total amount of the remaining frames is obtained according to the displacement amounts of the key points in the remaining frames; the remaining frame with the maximum value of the displacement total amount is selected and recorded as a key frame of the face video.
[0010] As a further scheme of the present application, the process of clustering each pronunciation combination comprises: presetting k cluster centers, obtaining facial muscle features of each pronunciation combination, regarding each facial muscle feature as an independent cluster, sequentially merging the two independent clusters closest to each other into the same cluster until the number of the final clusters is k; recording the pronunciation combinations corresponding to all the facial muscle features belonging to the same cluster as a pronunciation set.
[0011] As a further scheme of the present application, the process of determining the standard facial feature of the pronunciation set comprises: for any facial muscle feature in the pronunciation set, recording the facial muscle feature as a to-be-tested feature and recording the remaining facial muscle features in the pronunciation set as remaining features; obtaining the similarity mean value of the to-be-tested feature and each remaining feature wherein n is the total number of the remaining features, Of i represents the i-th remaining feature, i∈[1, n] and i is a positive integer, and Ct represents the to-be-tested feature; selecting the facial muscle feature with the maximum similarity mean value as the standard facial feature of the pronunciation set.
[0012] As a further scheme of the present application, the process of obtaining the association feature between the pronunciation combination and the typical pronunciation combination comprises: recording both the dynamic feature and the static feature as a sub-feature of the facial muscle feature, obtaining a standard facial feature of the typical pronunciation combination, presetting the weights of each sub-feature of the standard facial feature, and constantly adjusting the values of the weights so that the similarity between the standard facial feature and the facial muscle feature corresponding to the pronunciation combination reaches a preset similarity threshold; when the similarity between the standard facial feature and the facial muscle feature corresponding to the pronunciation combination reaches the preset similarity threshold, recording that the standard facial feature is consistent with the facial muscle feature, obtaining the values of the weights of each sub-feature at this time, integrating to obtain a weight sequence, and recording the weight sequence as the association feature of the values of the standard facial feature and the facial muscle feature, that is, the association feature between the pronunciation combination and the typical pronunciation combination.
[0013] As a further scheme of the present application, the setting process of the pronunciation template comprises: A Chinese character library is established, and all Chinese characters corresponding to the typical pronunciation combination in the Chinese character library are obtained, denoted as a typical Chinese character set, and a plurality of typical Chinese character sets are obtained according to the typical Chinese character set of each typical pronunciation combination; based on a natural language processing tool, a plurality of Chinese characters are selected in each typical Chinese character set to form a sentence, and the sentence is denoted as a pronunciation template.
[0014] The present application has the following beneficial effects: The present application effectively resists static deception attacks of photos, videos or 3D masks by analyzing dynamic facial muscle features of a user reading a random sentence; a multi-dimensional authentication system is formed by combining facial static features and dynamic pronunciation features, and the cracking difficulty is greatly increased; different sentence passwords are generated each time to force the user to read in real time, avoiding pre-recording attacks; the user individual differences (such as dialects and pronunciation habits) are adapted through the association feature model (such as weight sequence), reducing the false rejection rate while ensuring security; the pronunciation related areas are accurately segmented, the dynamic features are captured by combining the optical flow method, and the representative static features are extracted by key frame screening, improving the effectiveness of the features; the pronunciation set is constructed based on the clustering algorithm, and the pronunciation template with smooth semantics is generated by combining the natural language processing, considering the biological feature collection and user experience; in the present application, the user only needs to read the sentence generated by the system to unlock, without the need for a physical key or complex operation, which meets the modern human-computer interaction trend; the dynamic feature sequence is aligned by time stamp, ensuring the real-time and continuity of the authentication process; in summary, the present application solves the security risks of traditional facial recognition vulnerable to static counterfeit attacks by dynamic facial muscle feature analysis + random sentence live detection, improves the safe deposit box anti-deception ability, and takes into account the convenience of user use and system adaptability. BRIEF DESCRIPTION OF DRAWINGS
[0015] The present application will be further described below with reference to the drawings.
[0016] Figure 1 is a flow diagram of a safe unlocking method based on facial recognition according to the present application; Figure 2 is a flow diagram of step S1 of a safe unlocking method based on facial recognition according to the present application. DETAILED DESCRIPTION
[0017] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of the present application.
[0018] Referring to Figure 1 The application is a safe unlocking method based on facial recognition, comprising the following steps: Step S1: randomly combine the initial consonants and final consonants two by two to obtain a plurality of pronunciation combinations, and select a plurality of experimental personnel to obtain the facial videos of the experimental personnel reading each pronunciation combination, and extract the facial muscle features in each facial video; According to the facial muscle features, each pronunciation combination is clustered to obtain a plurality of pronunciation sets, and the standard facial features of each pronunciation set are determined, and the pronunciation combination corresponding to the standard facial features is recorded as the typical pronunciation combination; and the correlation features between each pronunciation combination in the pronunciation set and the typical pronunciation combination are obtained; It is worth noting that the initial consonants include special initial consonants y and w, and the pronunciation combination does not need to have a tone, only the combination of initial consonants and final consonants; As a preferred embodiment of the application, the extraction process of the facial muscle features of the pronunciation combination comprises: Divide the facial video of the pronunciation combination into a plurality of video frames, locate the facial range of the experimental personnel in the video frame based on a face detection algorithm, and mark the core feature points on the facial range through a facial key point detection model, the core feature points including the upper lip edge, the lower lip edge, the mouth corner, the lip pearl, the lower jaw line and the chin tip; based on all core feature points, a sub-region is segmented through a pixel mask technology, the sub-region including the lip, the lower jaw, the cheek, the forehead and the ear; all sub-regions are divided into pronunciation-related regions and pronunciation-unrelated regions, and all pronunciation-unrelated regions are excluded, the pronunciation-related regions including the lip, the lower jaw and the cheek, and the pronunciation-unrelated regions including the forehead and the ear; Mark the core feature points in the pronunciation-related regions as key points, and obtain the dynamic features of the key points in the facial video by sequentially obtaining the pixel positions of the key points in each video frame through the optical flow method, the dynamic features including the moving track, the moving acceleration and the displacement peak value; Filter out the key frames in all video frames of the facial video, and obtain the static features based on the pixel positions of each key point in the key frames, the static features including the opening degree, the round-lip degree, the mouth corner angle, the angle between the lower jaw line and the horizontal line, and the convexity value of the cheekbone region; Convert the dynamic features and the static features into numerical vectors of uniform dimensions, and record the numerical vectors as the facial muscle features of the pronunciation combination; The filtering process of the key frames comprises: In all video frames of the face video, the first video frame is recorded as a starting frame, and the remaining video frames are recorded as remaining frames; the pixel positions of the key points in the starting frame are obtained and recorded as starting positions, and the pixel positions of the key points in each of the remaining frames are sequentially obtained and recorded as remaining positions; the displacement amounts of the key points in the remaining frames compared to the key points in the starting frame are obtained according to the starting positions and the remaining positions, the total displacement amount of the remaining frames is obtained according to the displacement amounts of the key points in the remaining frames, and the remaining frame with the maximum value of the total displacement amount is selected as the key frame of the face video; Specifically, the process of obtaining the facial muscle features of the experimental personnel is to capture the facial muscle movement patterns related to pronunciation from dynamic videos and convert them into quantifiable and analyzable feature data; first, the face video is preprocessed, and the face video is image enhanced and standardized; to solve the problem of uneven light, the brightness and contrast are adjusted through histogram equalization and other methods to ensure that the facial details are clear and visible; The pronunciation action is mainly driven by specific muscle groups of the face, such as the lip muscles, the jaw muscles, and the cheek muscles, and these key areas need to be located and segmented first; specifically, the overall range of the face in the video frame is located using a face detection algorithm, such as MTCNN and YOLO, and then the core feature points related to pronunciation are marked using a face key point detection model, such as Dlib and FaceMesh, for example: 20-30 core feature points of the lip area, such as the upper lip edge, the lower lip edge, the corner of the mouth, and the lip pearl; the lower jaw line and chin tip of the jaw and chin area; the muscle attachment points near the zygomatic arch and alar region; based on the above core feature points, the lip, jaw, and cheek muscle activity sub-regions are segmented through polygon fitting or pixel mask technology, and the forehead, ears, and other areas unrelated to pronunciation are excluded to reduce redundant information interference; Further, the movement of facial muscles during pronunciation is the core of the feature, such as the opening and closing of the lips, the stretching of the corners of the mouth, and the lifting of the jaw, and quantitative indicators need to be extracted from both dynamic changes and static forms; the pixel displacement direction and distance of the key regions in consecutive frames are calculated by the optical flow method to quantify the speed and acceleration of muscle movement, with the speed representing the average speed of the lips from closing to opening; the Euclidean distance change of the feature points in adjacent frames is calculated to reflect the amplitude of the mouth corner distance increasing or decreasing with pronunciation, reflecting the strength of muscle contraction or relaxation, such as the mouth corner distance when pronouncing "a" being greater than when pronouncing "i"; the dynamic features of a single pronunciation video are arranged along the time axis, and the statistical properties of the time series are extracted, for example, the duration, which is the duration of a muscle movement state, such as the duration of the lip round state when pronouncing the long vowel "u"; the displacement peak, which is the maximum displacement of muscle movement, such as the displacement peak corresponding to the maximum force of the instantaneous closing of the lips when pronouncing the explosive sound "p"; In the key frame of the pronunciation process, the key frame is the moment when the pronunciation is the clearest, the morphological features of the muscle activity area are extracted, the typical facial state of the pronunciation is reflected, including the opening degree, the round lip degree, the mouth corner angle, the angle between the lower jaw line and the horizontal line, and the convexity value of the malar region; wherein, the opening degree is the ratio of the upper and lower lip distance to the face width, the round lip degree is the ellipse fitting degree of the lip contour, the mouth corner angle is the angle between the mouth corner line and the horizontal line, the angle between the lower jaw line and the horizontal line reflects the degree of opening, and the convexity value of the malar region is the area of the cheek bulge, such as the area of the slight cheek bulge when pronouncing 'o'; and through local binary pattern (LBP), gray level co-occurrence matrix (GLCM) and other algorithms, the skin texture changes of the muscle activity area are extracted, such as the skin wrinkle texture density caused by muscle contraction, to assist in distinguishing the subtle muscle movement differences, such as the subtle lip movements of's' and'sh'; Further, the above extracted dynamic and static features are converted into numerical vectors of uniform dimensions, the features are numerically coded, the statistical quantities such as peak value, duration and slope of the time series in the dynamic features are converted into specific numerical values, the geometric parameters (such as angle, distance and area) in the static features are directly taken as numerical values, and the texture features are converted into numerical sequences through histogram statistics; through principal component analysis method, the feature dimension is reduced, the redundant features are eliminated, and the calculation efficiency is improved; The facial muscle features of each experimental personnel reading the pronunciation combination are obtained in sequence, and the average feature of the facial muscle features of each experimental personnel is obtained as the facial muscle feature of the pronunciation combination, so that the obtained result has universality; It is worth noting that the duration of all the obtained facial videos is the same by default; As a preferred embodiment of the present application, the process of clustering each pronunciation combination includes: A preset k cluster centers are obtained, the facial muscle features of each pronunciation combination are obtained, each facial muscle feature is regarded as an independent cluster, the two nearest independent clusters are merged into the same cluster in sequence until the number of the final cluster is k; the pronunciation combinations corresponding to all the facial muscle features belonging to the same cluster are recorded as a pronunciation set; As a preferred embodiment of the present application, the process of determining the standard facial feature of the pronunciation set includes: For any facial muscle feature in the pronunciation set, the facial muscle feature is recorded as a to-be-tested feature, and the remaining facial muscle features in the pronunciation set are recorded as remaining features; the average similarity of the to-be-tested feature and each remaining feature is obtained Wherein n is the total number of the remaining features, Of i The i-th remaining feature is represented as i∈[1, n] and i is a positive integer, and Ct represents the to-be-tested feature; the facial muscle feature with the maximum average similarity is selected and recorded as the standard facial feature of the pronunciation set; Specifically, the direction similarity is measured by calculating the cosine value of the angle between two vectors, and the value range of the similarity average is [-1, 1], the closer to 1 indicates the more similar the directions, the closer to -1 indicates the opposite directions, and 0 indicates the orthogonal directions; As a preferred embodiment of the present application, the process of obtaining the association feature between the pronunciation combination and the typical pronunciation combination comprises: The dynamic feature and the static feature are recorded as sub-features of the facial muscle feature, the standard facial feature of the typical pronunciation combination is obtained, the weights of each sub-feature of the standard facial feature are preset, and the values of each weight are adjusted constantly, so that the similarity between the standard facial feature and the facial muscle feature corresponding to the pronunciation combination reaches a preset similarity threshold; When the similarity between the standard facial feature and the facial muscle feature corresponding to the pronunciation combination reaches the preset similarity threshold, it is recorded that the standard facial feature is consistent with the facial muscle feature, and the values of each weight at this time are obtained, and the weight sequence is integrated to obtain the association feature between the standard facial feature and the facial muscle feature value, that is, the association feature between the pronunciation combination and the typical pronunciation combination; It can be understood that in obtaining the association feature between the pronunciation combination and the typical pronunciation combination, it is first determined that the dynamic feature and the static feature are both sub-features of the facial muscle feature. Then, the standard facial feature of the typical pronunciation combination is obtained, and initial weights of each sub-feature of the standard feature are preset; Then, by constantly adjusting the weights of each sub-feature, the similarity between the standard facial feature and the facial muscle feature corresponding to the pronunciation combination to be analyzed gradually approaches and finally reaches a preset similarity threshold; this adjustment process needs to combine the feature matching algorithm to repeatedly iterate and optimize the weight distribution; When the similarity between the two satisfies the preset threshold, it is considered that the standard facial feature is consistent with the facial muscle feature of the pronunciation combination, and the weight values of each sub-feature are extracted at this time, and these weights are integrated to form an ordered weight sequence; the weight sequence can accurately reflect the association degree between the standard facial feature and each sub-feature of the facial muscle feature of the pronunciation combination to be analyzed, that is, the association feature between the pronunciation combination and the typical pronunciation combination; Step S2: According to each typical pronunciation combination, a pronunciation template is set, a template pronunciation video of the user reading the pronunciation template is input, and according to the template pronunciation video, the facial pronunciation template feature of the user is obtained; As a preferred embodiment of the present application, the setting process of the pronunciation template comprises: A Chinese character library is established, and all Chinese characters corresponding to the typical pronunciation combination in the Chinese character library are obtained, denoted as a typical Chinese character set, and a plurality of typical Chinese character sets are obtained according to the typical Chinese character set of each typical pronunciation combination; a sentence is formed by selecting a plurality of Chinese characters in each typical Chinese character set based on a natural language processing tool, and the sentence is denoted as a pronunciation template. Specifically, when the pronunciation template is set, a Chinese character library covering a large number of Chinese characters is first established to ensure that the Chinese characters corresponding to all types of typical pronunciation combinations can be covered; then, for each typical pronunciation combination, all Chinese characters containing the pronunciation combination are selected from the Chinese character library, and these Chinese characters are classified and integrated to form a typical Chinese character set corresponding to each typical pronunciation combination, thereby obtaining a plurality of independent typical Chinese character sets; Then, the typical Chinese character sets are processed by means of a natural language processing tool, and a proper amount of Chinese characters are selected in each typical Chinese character set according to the frequency of use of the Chinese characters, semantic relevance, and matching degree with the pronunciation combination; and the selected Chinese characters are combined and optimized by means of the natural language processing tool to form a sentence that is smooth in semantics and complies with grammatical rules, and such a sentence is determined as a pronunciation template. For example, if a typical pronunciation combination corresponds to "ang", and the typical Chinese character set thereof contains "ang", "wang", and "fang", appropriate Chinese characters are selected from the set by means of a natural language processing tool to form a pronunciation template such as "angshou wangyuanfang". As a preferred embodiment of the present application, the process of obtaining the facial pronunciation template features of the user includes: The template pronunciation video of the user reading the pronunciation template is image-enhanced and standardized, the pronunciation area is segmented by means of face detection and core feature point positioning technology, and the key points of the pronunciation area are obtained, the dynamic features and static features of each key point are captured in the template pronunciation video, and finally the facial pronunciation template features of the user are obtained; It can be understood that the process of obtaining the facial pronunciation template features of the user refers to the process of obtaining the facial muscle features of the experimenter in step S1, which is not described herein again; Step S3: When the user unlocks the safe, the safe randomly generates a sentence password, obtains a facial video of the user reading the sentence password, denoted as a current facial video, obtains a current facial feature sequence according to the current facial video, and generates a current template feature sequence of the sentence password according to the facial pronunciation template features; The current template feature sequence and the current facial feature sequence are compared to determine whether the authentication is passed, and if the authentication is passed, the safe is unlocked; As a preferred embodiment of the present application, the process of generating the sentence password includes: A preset minimum number threshold m, based on natural language processing technology in the Chinese character library randomly selects N Chinese characters, wherein N≥m, a new sentence is formed by N Chinese characters, and the new sentence is recorded as a sentence password; Specifically, in the generation of the sentence password, a minimum number threshold m of Chinese characters is preset to ensure that the password has a certain length and complexity; relying on natural language processing technology, random selection is performed in the established Chinese character library, and the number N of selected Chinese characters needs to meet the condition of N≥m; In the selection process, the natural language processing technology will screen the Chinese characters to ensure that the selected Chinese characters have a certain relevance in semantics, so that the formed sentence can have basic fluency; finally, the N Chinese characters are combined into a new sentence according to a reasonable sentence order, and the sentence is determined as a sentence password; for example, if m is preset to 5, 6 Chinese characters may be randomly selected to form a sentence password such as “spring breeze willow swallow return”; As a preferred embodiment of the present application, the obtaining process of the current face feature sequence comprises: Obtain the play time sequence of the current face video, divide the play time sequence into a plurality of play time stamps, obtain the video frame corresponding to each play time stamp, and record it as a current video frame; obtain the facial muscle features in each video frame through face detection and core feature point positioning technology, and record it as a current facial feature; associate each current facial feature with the play time stamp corresponding to the current video frame to obtain the current facial feature sequence; As a preferred embodiment of the present application, the generation process of the current template feature video of the sentence password comprises: Obtain all pronunciation combinations in the sentence password, record them as password pronunciation combinations, and obtain the pronunciation set to which the password pronunciation combination belongs, record it as a password pronunciation set; obtain the typical pronunciation combination corresponding to the standard face feature in the password pronunciation set, record it as a password typical pronunciation; and in each face pronunciation template feature of the user, screen out the face pronunciation template feature corresponding to the password typical pronunciation, record it as a password template feature; obtain the association feature between the password typical pronunciation and the password pronunciation combination, and correct the password template feature according to the association feature to obtain the current template feature of the password pronunciation combination; Based on audio recognition technology, obtain the play time stamp corresponding to each password pronunciation combination, and associate the current template feature of each password pronunciation combination with the play time stamp to obtain the current template feature sequence; The process of obtaining the play time stamp corresponding to each password pronunciation combination based on audio recognition technology is: The audio of the user reading the sentence password is analyzed first to identify the start and end time of each password pronunciation combination in the audio; then the time is corresponded with the playing time sequence of the face video to obtain the playing time stamp of each password pronunciation combination in the video; As a preferred embodiment of the present application, the process of determining whether the authentication is passed includes: In the current template feature sequence and the current face feature sequence, the current template features and the current face features at each playing time stamp are compared in sequence, and if the current template features and the current face features at each time stamp are consistent, the authentication is passed; In the process of determining whether the authentication is passed, the features of the corresponding time points in the current template feature sequence and the current face feature sequence are compared one by one based on the playing time stamp; specifically, starting from the first playing time stamp, the current template features and the current face features at the time stamp are checked in sequence, including the consistency of dynamic features and static features; if the features at all time stamps are completely matched without any difference, it is determined that the authentication is passed; as long as the features of any time stamp are not matched, the authentication fails; this process ensures that the user's face pronunciation features and the expected features are completely matched through the feature comparison of the whole time sequence, thereby improving the accuracy and security of identity authentication.
[0019] The safe box of the present application supports remote unlocking in the unlocking process, constructs a secure transmission system through the SM2 encryption algorithm, and ensures the authenticity of the instruction through the public and private key signature verification of the vault system (Client) and the safe box system (Server). The data is transmitted in JSON format through the Socket short link (port 9000), wherein the srvId and box_id parameters of the single unlocking instruction interface are designed; and the safe box supports entity network connection to realize remote communication, and remote unlocking needs to specify the device through IP positioning and verify the box_id; The safe box local authentication process based on facial muscle features of the present application covers core links such as facial feature extraction, template setting, password generation, feature comparison and authentication judgment, and finally realizes the local verification logic of feature matching consistency unlocking. Combined with the remote instruction interaction of the safe box, a complete closed loop of local authentication and remote instruction execution is formed.
[0020] The above describes one embodiment of the present application in detail, but the content described is only a preferred embodiment of the present application and cannot be considered as limiting the scope of the present application. Any equivalent changes and improvements made within the scope of the present application should still belong to the inventive scope of the present application.
Claims
1. A method for unlocking a safe based on facial recognition, characterized in that, Includes the following steps: Step S1: Set several pronunciation combinations and obtain the facial muscle features of each pronunciation combination; based on the facial muscle features, cluster each pronunciation combination to obtain several pronunciation sets, and determine the standard facial features of each pronunciation set, obtain the pronunciation combination corresponding to the standard facial features, and record it as a typical pronunciation combination; and obtain the correlation features between each pronunciation combination in the pronunciation set and the typical pronunciation combination. Step S2: Set a pronunciation template according to each typical pronunciation combination, record a template pronunciation video of the user reading the pronunciation template, and obtain the user's facial pronunciation template features based on the template pronunciation video; Step S3: When the user unlocks the safe, the safe randomly generates a password, and the user reads the password aloud. This is recorded as the current facial video. Based on the current facial video, the current facial feature sequence is obtained. The safe generates the current template feature sequence of the statement password based on the facial pronunciation template features; Compare the current template feature sequence with the current facial feature sequence to determine whether authentication is successful. If authentication is successful, the safe is unlocked.
2. The method for unlocking a safe based on facial recognition according to claim 1, characterized in that, In step S1, the process of obtaining the facial muscle features of the pronunciation combination includes: Initials and finals were randomly paired to obtain several pronunciation combinations. Several experimenters were selected, and facial videos of the experimenters reading each pronunciation combination were obtained. Facial muscle features were extracted from each facial video.
3. The method for unlocking a safe based on facial recognition according to claim 2, characterized in that, In step S1, the extraction process of facial muscle features of the pronunciation combination further includes: The facial video of the pronunciation combination is divided into several video frames. The facial range of the experimenter in the video frames is located based on a face detection algorithm, and core feature points are marked on the facial range using a facial key point detection model. The core feature points include the upper lip edge, lower lip edge, corners of the mouth, philtrum, jawline, and chin tip. Based on all core feature points, sub-regions are segmented using pixel masking technology. The sub-regions include the lips, jaw, cheeks, forehead, and ears. All sub-regions are divided into pronunciation-related regions and pronunciation-independent regions, and all pronunciation-independent regions are excluded. The pronunciation-related regions include the lips, jaw, and cheeks, and the pronunciation-independent regions include the forehead and ears. The core feature points in the pronunciation-related region are recorded as key points. The pixel positions of the key points in each video frame are obtained sequentially by optical flow method to obtain the dynamic features of the key points in the facial video. The dynamic features include movement trajectory, movement acceleration and displacement peak. Keyframes are selected from all video frames of the facial video. Static features are obtained based on the pixel positions of each key point in the keyframes. The static features include the degree of opening and closing, the degree of lip roundness, the angle of the corner of the mouth, the angle between the jawline and the horizontal line, and the convexity value of the cheekbone area. The dynamic and static features are transformed into numerical vectors of a unified dimension, and the numerical vectors are denoted as the facial muscle features of the pronunciation combination.
4. The method for unlocking a safe based on facial recognition according to claim 3, characterized in that, In step S1, the keyframe filtering process includes: In all video frames of the facial video, the first video frame is designated as the start frame, and the remaining video frames are designated as the rest frames. The pixel positions of each key point in the start frame are obtained and designated as the start position. The pixel positions of each key point in each of the rest frames are obtained sequentially and designated as the rest positions. Based on the start position and the rest positions, the displacement of each key point in the rest frames relative to the key points in the start frame is obtained. Based on the displacement of each key point in the rest frames, the total displacement of the rest frames is obtained. The rest frame with the largest total displacement value is selected and designated as the key frame of the facial video.
5. A method for unlocking a safe based on facial recognition according to claim 1, characterized in that, In step S1, the process of clustering each pronunciation combination includes: Predetermine k cluster centers, obtain facial muscle features for each pronunciation combination, treat each facial muscle feature as an independent cluster, and merge the two closest independent clusters into the same cluster until the final number of clusters is k; record the pronunciation combinations corresponding to all facial muscle features belonging to the same cluster as a pronunciation set.
6. The method for unlocking a safe based on facial recognition according to claim 1, characterized in that, In step S1, the process of determining the standard facial features for the pronunciation set includes: For any facial muscle feature within the pronunciation set, this facial muscle feature is designated as the test feature, and the remaining facial muscle features within the pronunciation set are designated as other features; the average similarity between the test feature and each of the other features is then obtained. Where n is the total number of the remaining features, Of i Let i represent the i-th remaining feature, i∈[1,n] and i is a positive integer, and Ct represent the feature to be tested; select the facial muscle feature with the largest mean similarity and denote it as the standard facial feature of the pronunciation set.
7. A method for unlocking a safe based on facial recognition according to claim 3, characterized in that, In step S1, the process of obtaining the association features between the pronunciation combination and the typical pronunciation combination includes: Both the dynamic and static features are recorded as sub-features of the facial muscle features. The standard facial features of the typical pronunciation combination are obtained. The weights of each sub-feature of the standard facial features are preset. The values of each weight are continuously adjusted so that the similarity between the standard facial features and the facial muscle features corresponding to the pronunciation combination reaches a preset similarity threshold. When the similarity between the standard facial feature and the facial muscle feature corresponding to the pronunciation combination reaches a preset similarity threshold, the standard facial feature and the facial muscle feature are considered to be consistent. The weight values of each sub-feature are then obtained and integrated to obtain a weight sequence, which is recorded as the association feature between the standard facial feature and the facial muscle feature value, i.e., the association feature between the pronunciation combination and the typical pronunciation combination.
8. A method for unlocking a safe based on facial recognition according to claim 1, characterized in that, In step S2, the process of setting the pronunciation template includes: A Chinese character database is established, and all Chinese characters corresponding to typical pronunciation combinations in the database are obtained and denoted as typical Chinese character sets. Based on the typical Chinese character sets of each typical pronunciation combination, several typical Chinese character sets are obtained. Based on natural language processing tools, several Chinese characters are selected from each typical Chinese character set to form a sentence, and the sentence is denoted as a pronunciation template.
Citation Information
Patent Citations
Speech recognition-based unlocking method, and intelligent door lock system thereof
CN106920303A
Intelligent lock with face recognition function for locker and using method
CN113445829A
Autonomous language learning system based on big data speech recognition
CN120236589A
Intelligent conference management method and system
WO2019148583A1