User release method and device, computer device and storage medium

By using a multimodal biometric verification method, the quality scores and weights of fingerprint images, facial images, and voice signals are obtained and processed, solving the problem of easy forgery of single biometric features in traditional IoT devices and achieving higher accuracy and security in identity verification.

CN121747231BActive Publication Date: 2026-05-08SHENZHEN JOOAN TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN JOOAN TECH CO LTD
Filing Date
2026-02-28
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In traditional IoT devices, single biometric verification is easily forged, resulting in low accuracy of identity verification.

Method used

A multimodal biometric verification method is adopted to acquire the user's fingerprint image, facial image and voice signal, determine their quality scores and assign weights, and determine the release result by similarity calculation.

Benefits of technology

This significantly improves the accuracy of identity verification, ensuring the accuracy and security of release results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121747231B_ABST
    Figure CN121747231B_ABST
Patent Text Reader

Abstract

The application relates to a user release method and device, computer equipment and a storage medium. The method comprises the following steps: acquiring a first fingerprint image, a first face image and a first voice signal of a first user; determining the quality scores of the first fingerprint image, the first face image and the first voice signal respectively, so as to determine the first weight of the first fingerprint image, the second weight of the first face image and the third weight of the first voice signal respectively; determining the first similarity between the first fingerprint image and a second fingerprint image of a second user, the second similarity between the first face image and a second face image of the second user, and the third similarity between the first voice signal and a second voice signal of the second user; and determining the release result for the first user based on the first weight, the first similarity, the second weight, the second similarity, the third weight and the third similarity. The method can improve the identity verification accuracy for the first user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to a user release method, apparatus, computer equipment, and storage medium. Background Technology

[0002] As people's living standards rapidly improve, access control security has gradually gained attention. Traditional IoT devices typically use a single fingerprint image or a single facial information to verify visitor identity. If the verification information is singular, it can be easily forged, leading to incorrect access. Summary of the Invention

[0003] Therefore, it is necessary to provide a user access method, apparatus, computer device, computer-readable storage medium, and computer program product that can improve the accuracy of authentication for the first user in order to address the above-mentioned technical problems.

[0004] Firstly, this application provides a user permission method, including:

[0005] Acquire the first fingerprint image, the first facial image, and the first voice signal of the first user;

[0006] The quality scores of the first fingerprint image, the first face image, and the first voice signal are determined respectively.

[0007] Based on the quality scores of the first fingerprint image, the first face image, and the first voice signal, a first weight of the first fingerprint image, a second weight of the first face image, and a third weight of the first voice signal are determined respectively.

[0008] The first similarity between the first fingerprint image and the second fingerprint image of the second user, the second similarity between the first face image and the second face image of the second user, and the third similarity between the first voice signal and the second voice signal of the second user are determined respectively, wherein the second user is any user stored.

[0009] Based on the first weight, first similarity, second weight, second similarity, third weight, and third similarity, the release result for the first user is determined.

[0010] Secondly, this application also provides a user release device, comprising:

[0011] The acquisition module is used to acquire the first fingerprint image, the first facial image, and the first voice signal of the first user;

[0012] The first determining module is used to determine the quality scores of the first fingerprint image, the first face image, and the first voice signal, respectively.

[0013] The second determining module is used to determine the first weight of the first fingerprint image, the second weight of the first face image, and the third weight of the first voice signal based on the quality scores of the first fingerprint image, the first face image, and the first voice signal, respectively.

[0014] The third determining module is used to determine the first similarity between the first fingerprint image and the second fingerprint image of the second user, the second similarity between the first face image and the second face image of the second user, and the third similarity between the first voice signal and the second voice signal of the second user, wherein the second user is any user stored.

[0015] The fourth determination module is used to determine the release result for the first user based on the first weight, the first similarity, the second weight, the second similarity, the third weight, and the third similarity.

[0016] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement some or all of the steps described in any method of the first aspect of the embodiments of this application.

[0017] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements some or all of the steps described in any method of the first aspect of the embodiments of this application.

[0018] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements some or all of the steps described in any method of the first aspect of the embodiments of this application.

[0019] The aforementioned user release method, apparatus, computer equipment, computer-readable storage medium, and computer program product acquire a first fingerprint image, a first facial image, and a first voice signal of a first user, and determine the quality scores of the first fingerprint image, the first facial image, and the first voice signal respectively. Based on these quality scores, a first weight for the first fingerprint image, a second weight for the first facial image, and a third weight for the first voice signal are determined. Simultaneously, a first similarity between the first fingerprint image and a second fingerprint image of a second user, a second similarity between the first facial image and a second facial image of a second user, and a third similarity between the first voice signal and a second voice signal of a second user are determined. Furthermore, based on the first weight, first similarity, second weight, second similarity, third weight, and third similarity, a release result for the first user is determined. It can be seen that the user release method provided in this application can accurately determine the release result for the first user through multimodal biometrics, thereby significantly improving the accuracy of identity verification for the first user. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a diagram illustrating the application environment of the user release method in one embodiment;

[0022] Figure 2 This is a flowchart illustrating a user release method in one embodiment;

[0023] Figure 3 This is a structural block diagram of a user release device in one embodiment;

[0024] Figure 4 This is an internal structural diagram of a computer device in one embodiment;

[0025] Figure 5 This is a diagram of the internal structure of a computer device in another embodiment. Detailed Implementation

[0026] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0027] The user release method provided in this application embodiment can be applied to, for example, Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network and is connected to fingerprint sensor 106, camera 108, and microphone 110. A data storage system can store the data that server 104 needs to process. The data storage system can be integrated onto server 104 or located on the cloud or other network servers. Terminal 102 can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and IoT devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, smart vehicle devices, projection devices, etc. Server 104 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing cloud computing services.

[0028] In one exemplary embodiment, such as Figure 2 As shown, a user permission method is provided, which is applied to... Figure 1 Taking the terminal in the example, the explanation includes the following steps 202 to 210. Wherein:

[0029] Step 202: Obtain the first fingerprint image, the first facial image, and the first voice signal of the first user.

[0030] The first user refers to the user who needs to be authenticated.

[0031] The first fingerprint image refers to the fingerprint image of the first user captured by the fingerprint sensor.

[0032] The first facial image refers to the facial image of the first user captured by the camera.

[0033] The first voice signal refers to the voice signal of the first user collected by the microphone.

[0034] The fingerprint sensor, camera, and microphone are each connected to the terminal.

[0035] Optionally, the fingerprint sensor can be a capacitive sensor or an optical fingerprint sensor; the camera can be an infrared network surveillance camera; and the microphone can be an array microphone or a microelectromechanical system microphone.

[0036] In an exemplary embodiment, the acquisition of the first user's first fingerprint image, first facial image, and first voice signal includes: acquiring the first user's first fingerprint image through a fingerprint sensor, acquiring the first user's first facial image through a camera, and acquiring the first user's first voice signal through a microphone.

[0037] Optionally, the terminal may acquire the first user's first fingerprint image, first facial image, and first voice signal when it determines that the camera's monitoring screen has changed, i.e., when it determines that the first user has appeared in the camera's monitoring screen.

[0038] Optionally, the terminal can be a terminal in an access control management system. Thus, by executing the user access release method provided in this application embodiment, the terminal can improve the security of the access control management system by increasing the accuracy of identity verification for the first user. Optionally, the access control management system can be used in residential communities, corporate parks, and other monitored areas requiring access control management.

[0039] Step 204: Determine the quality scores of the first fingerprint image, the first face image, and the first voice signal, respectively.

[0040] The quality score of the first fingerprint image is used to quantify the quality of the first fingerprint features extracted from it. A higher quality score indicates that the first fingerprint image can extract higher quality first fingerprint features. The same applies to the quality scores of the first face image and the first speech signal, which will not be elaborated upon further below.

[0041] The quality score of the first face image is used to quantify the quality of the first face features extracted from the first face image.

[0042] The quality score of the first speech signal is used to quantify the quality of the first speech features extracted based on the first speech signal.

[0043] Optionally, the first fingerprint feature, the first facial feature, and the first voice feature can be presented as feature vectors, and the feature vectors of different biometric modalities have a unified vector format.

[0044] Optionally, the quality scores of the first fingerprint image, the first face image, and the first voice signal can be represented by 0 to 1 or by 0 to 100%.

[0045] Step 206: Based on the quality scores of the first fingerprint image, the first face image, and the first voice signal, determine the first weight of the first fingerprint image, the second weight of the first face image, and the third weight of the first voice signal, respectively.

[0046] The first weight of the first fingerprint image is used to quantify the importance of the first fingerprint image in the release result for the first user.

[0047] The second weight of the first face image is used to quantify the proportion of importance of the first face image in the release result for the first user.

[0048] The third weight of the first voice signal is used to quantify the importance of the first voice signal in the release result for the first user.

[0049] The first weight is positively correlated with the quality score of the first fingerprint image, the second weight is positively correlated with the quality score of the first face image, and the third weight is positively correlated with the quality score of the first speech signal.

[0050] Optionally, the sum of the first weight, the second weight, and the third weight can be 1.

[0051] Step 208: Determine the first similarity between the first fingerprint image and the second fingerprint image of the second user, the second similarity between the first face image and the second face image of the second user, and the third similarity between the first voice signal and the second voice signal of the second user, wherein the second user is any of the stored users.

[0052] The first similarity between the first fingerprint image and the second fingerprint image of the second user can be determined by a fingerprint matching algorithm. Optionally, the fingerprint matching algorithm can be the Minutia Cylinder-Code (MCC) algorithm.

[0053] The second similarity between the first facial image and the second user's second facial image can be determined by a face recognition algorithm. Optionally, the face recognition algorithm can be a multi-task cascaded convolutional network (MTCNN) algorithm.

[0054] The third similarity between the first speech signal and the second user's second speech signal can be determined by a voiceprint recognition algorithm. Optionally, the voiceprint recognition algorithm can be a Gaussian Mixture Model (GMM) algorithm.

[0055] A second user refers to any user pre-stored in the authorized user database on the terminal. This authorized user database includes multiple authorized users who have pre-registered their second fingerprint image, second facial image, and second voice signal with the terminal.

[0056] In an exemplary embodiment, the above-described determination of the first similarity between the first fingerprint image and the second fingerprint image of the second user, the second similarity between the first face image and the second face image of the second user, and the third similarity between the first voice signal and the second voice signal of the second user includes: determining the first similarity between the first fingerprint feature corresponding to the first fingerprint image and the second fingerprint feature corresponding to the second fingerprint image of the second user, the second similarity between the first face feature corresponding to the first face image and the second face feature corresponding to the second face image of the second user, and the third similarity between the first voice feature corresponding to the first voice signal and the second voice feature corresponding to the second voice signal of the second user.

[0057] Step 210: Based on the first weight, first similarity, second weight, second similarity, third weight, and third similarity, determine the release result for the first user.

[0058] The decision to allow access to the first user can be based on a weighted sum of the first weight, first similarity, second weight, second similarity, third weight, and third similarity.

[0059] The outcome for the first user includes whether the first user is allowed to proceed or denied.

[0060] Optionally, the terminal can count the number of times the access control system has triggered a denial of access to the first user, and generate an alarm signal when the number of triggers is greater than or equal to a threshold. The alarm signal is then sent to a first target terminal, which can be a terminal used by the administrator of the access control system. Sending the alarm signal to the first target terminal can serve as an alarm reminder to the administrator of the access control system regarding the first user.

[0061] Optionally, after determining the release result for the first user, the terminal may also generate a log event based on the release result for the first user and the first similarity threshold, and upload the log event to the server or send it to the first target terminal.

[0062] In the aforementioned user release method, a first fingerprint image, a first facial image, and a first voice signal of a first user are acquired, and quality scores for each of these images are determined. Based on these quality scores, a first weight for the first fingerprint image, a second weight for the first facial image, and a third weight for the first voice signal are determined. Simultaneously, a first similarity between the first fingerprint image and the second fingerprint image of a second user, a second similarity between the first facial image and the second facial image of a second user, and a third similarity between the first voice signal and the second voice signal of a second user are determined. Finally, based on these weights, the release result for the first user is determined. It can be seen that the user release method provided in this application can accurately determine the release result for the first user through multimodal biometrics, thereby significantly improving the accuracy of identity verification for the first user.

[0063] In an exemplary embodiment, determining to allow access for the first user when the fusion similarity is greater than or equal to a first similarity threshold includes:

[0064] If the fusion similarity is greater than or equal to the first similarity threshold, the user type of the first user is determined; the user type of the first user is either a formal user or a temporary access user.

[0065] If the first user's user type is a formal user, then grant access to the first user.

[0066] If the first user's user type is a temporary access user, a one-time verification code is generated and sent to the second target terminal; the second target terminal is the terminal corresponding to the first user's access destination; after receiving the confirmation instruction from the second target terminal regarding the one-time verification code, it is determined that the first user can be allowed access.

[0067] If the first user is a temporary user and has not received a confirmation instruction from the second target terminal for the one-time verification code, then the first user will not be allowed to pass.

[0068] In one exemplary embodiment, after determining the release result for the first user, a log event can be generated for the release result for the first user.

[0069] In an exemplary embodiment, determining the quality scores of the first fingerprint image, the first face image, and the first voice signal respectively includes:

[0070] The quality score of the first fingerprint image is determined based on the first image parameters of the first fingerprint image.

[0071] The quality score of the first face image is determined based on the second image parameters of the first face image.

[0072] The quality score of the first speech signal is determined based on the speech parameters of the first speech signal.

[0073] The first image parameter refers to a parameter that characterizes the quality of the first fingerprint features extracted from the first fingerprint image. The first image parameter reflects the imaging quality of the first fingerprint image.

[0074] The second image parameter refers to a parameter that characterizes the quality of the first facial features extracted from the first facial image. The second image parameter reflects the imaging quality of the first facial image.

[0075] Speech parameters are parameters that characterize the quality of the first speech features extracted from the first speech signal. Speech parameters reflect the acquisition quality of the first speech signal.

[0076] In this embodiment, based on the first image parameters of the first fingerprint image, the second image parameters of the first face image, and the voice parameters of the first voice signal, the quality scores of the first fingerprint image, the first face image, and the first voice signal can be determined respectively. Thus, by accurately determining the quality scores of different modal biometrics, the high accuracy of the subsequently obtained first weight, second weight, and third weight is ensured. Consequently, the release result for the first user can be accurately determined, thereby significantly improving the accuracy of identity verification for the first user.

[0077] In an exemplary embodiment, the first image parameters of the first fingerprint image include the clarity, ridge integrity, and effective area ratio of the first fingerprint image; the second image parameters of the first face image include the illumination uniformity, occlusion rate, and key point confidence of the first face image; and the voice parameters of the first voice signal include the signal-to-noise ratio and effective duration of the voice signal.

[0078] The above-mentioned determination of the quality score of the first fingerprint image based on the first image parameters of the first fingerprint image includes:

[0079] The quality score of the first fingerprint image is determined based on its clarity, ridge integrity, and effective area ratio. The quality score of the first fingerprint image is positively correlated with its clarity, ridge integrity, and effective area ratio.

[0080] The above-mentioned determination of the quality score of the first face image based on the second image parameters of the first face image includes:

[0081] Based on the illumination evenness, occlusion rate, and key point confidence of the first face image, the quality score of the first face image is determined. The quality score of the first face image is positively correlated with the illumination evenness and key point confidence of the first face image, respectively, and negatively correlated with the occlusion rate.

[0082] The above-mentioned determination of the quality score of the first speech signal based on the speech parameters of the first speech signal includes:

[0083] The quality score of the first speech signal is determined based on the signal-to-noise ratio and effective duration of the speech signal. The quality score of the first speech signal is positively correlated with the signal-to-noise ratio and effective duration of the speech signal, respectively.

[0084] The sharpness of the first fingerprint image is used to quantify the edge sharpness between ridges and valleys in the first fingerprint image (which can also be understood as the contrast between ridges and valleys). The sharpness of the first fingerprint image reflects its ability to resolve fingerprint image details.

[0085] The ridge integrity of the first fingerprint image is used to quantify the connectivity of the ridges as continuous curves in the first fingerprint image. The ridge integrity of the first fingerprint image reflects the degree to which the first fingerprint image is not affected by image noise.

[0086] The effective area ratio of the first fingerprint image is used to quantify the ratio between the area of ​​the region in the first fingerprint image that can be used to extract the first fingerprint feature and the total area of ​​the first fingerprint image.

[0087] The illumination uniformity of the first facial image is used to quantify the evenness of lighting at different locations within the first facial image. The illumination uniformity of the first facial image reflects whether there are overexposure, underexposure, or other uneven lighting phenomena in the first facial image.

[0088] The occlusion rate of the first face image is used to quantify the proportion of the first face image that is obscured by non-face objects. The occlusion rate of the first face image reflects the integrity of the effective face region of the first face image.

[0089] The keypoint confidence score of the first face image is used to quantify the accuracy of keypoint localization in the first face image. The keypoint confidence score of the first face image reflects the accuracy of extracting the corresponding first facial features.

[0090] The signal-to-noise ratio (SNR) of the first speech signal is used to quantify the energy ratio between the effective speech portion and the ambient noise portion of the first speech signal. The signal-to-noise ratio (SNR) of the first speech signal reflects the purity of the first speech signal relative to ambient noise.

[0091] The effective duration of the first speech signal is used to quantify the duration ratio between the duration of the effective speech portion of the first speech signal and the total duration of the first speech signal.

[0092] In this embodiment, the quality score of the first fingerprint image is determined based on its clarity, ridge integrity, and effective area ratio. Similarly, the quality score of the first face image is determined based on its illumination uniformity, occlusion rate, and key point confidence. Furthermore, the quality score of the first speech signal is determined based on its signal-to-noise ratio and effective speech duration. This multi-dimensional information allows for the comprehensive determination of quality scores for different biometric modalities, improving the accuracy of these scores and ensuring the high accuracy of the subsequently obtained first, second, and third weights. Consequently, the release result for the first user can be accurately determined, significantly improving the accuracy of user authentication.

[0093] In one exemplary embodiment, the method further includes:

[0094] If the quality score of the first fingerprint image is less than a first score threshold, the quality score of the first face image is less than a second score threshold, or the quality score of the first voice signal is less than a third score threshold, the steps of acquiring the first fingerprint image, the first face image, and the first voice signal of the first user, and determining the quality scores of the first fingerprint image, the first face image, and the first voice signal respectively are executed.

[0095] If the quality score of the first fingerprint image is less than the first score threshold, the quality score of the first face image is less than the second score threshold, or the quality score of the first voice signal is less than the third score threshold, it indicates that the acquisition quality of at least one modality of biometrics is low. This situation will lead to a decrease in the accuracy of the subsequent release result for the first user.

[0096] The step of acquiring the first fingerprint image, the first face image, and the first voice signal of the first user, and determining the quality scores of the first fingerprint image, the first face image, and the first voice signal respectively, refers to reacquiring the first fingerprint image, the first face image, and the first voice signal of the first user, and re-determining the quality scores of the reacquiring first fingerprint image, the reacquiring first face image, and the reacquiring first voice signal respectively.

[0097] In this embodiment, if the acquisition quality of at least one of the multiple biometric modalities is low, the terminal re-executes the steps of acquiring the first fingerprint image, the first facial image, and the first voice signal of the first user, and re-executes the steps of determining the quality scores of the first fingerprint image, the first facial image, and the first voice signal respectively. That is to say, the terminal can ensure that the final release result for the first user will only be determined when the quality of the multiple biometric features is qualified. In other words, it ensures that the final release result for the first user can be accurately determined. Based on this, the accuracy of identity verification for the first user can be significantly improved.

[0098] In an exemplary embodiment, determining the quality score of the first fingerprint image based on the first image parameters of the first fingerprint image includes:

[0099] Based on the first image parameters of the first fingerprint image and the contamination level of the fingerprint sensor used to acquire the first fingerprint image, the quality score of the first fingerprint image is determined; there is a negative correlation between the quality score of the first fingerprint image and the contamination level of the fingerprint sensor.

[0100] The above-mentioned determination of the quality score of the first face image based on the second image parameters of the first face image includes:

[0101] The quality score of the first face image is determined based on the second image parameters of the first face image and the brightness value of the first face image; there is a positive correlation between the quality score of the first face image and the brightness value of the first face image.

[0102] The above-mentioned determination of the quality score of the first speech signal based on the speech parameters of the first speech signal includes:

[0103] The quality score of the first speech signal is determined based on the speech parameters and the environmental noise index of the first speech signal; the quality score of the first speech signal is negatively correlated with the environmental noise index of the first speech signal.

[0104] The contamination level of a fingerprint sensor refers to the degree to which the surface of the fingerprint sensor is covered by contaminants. Based on this, the contamination level of a fingerprint sensor can be determined by the ratio of the area covered by contaminants to the surface area of ​​the fingerprint sensor.

[0105] The brightness value of the first face image can be the average brightness of all pixels in the first face image.

[0106] The environmental noise index of the first speech signal refers to the energy level of the background environmental noise during the initial period when the first speech signal is just beginning to be recorded and the first user has not yet spoken.

[0107] Optionally, the terminal may store image parameters of the fingerprint image, a first correspondence between the fingerprint sensor and the contamination level of the fingerprint image, and so that, based on the first correspondence, the quality score of the first fingerprint image can be determined based on the first image parameters of the first fingerprint image and the contamination level of the fingerprint sensor used to acquire the first fingerprint image.

[0108] Optionally, the terminal may store a second correspondence between the image parameters of the face image, the brightness value of the face image, and the quality score of the face image. Based on this second correspondence, the quality score of the first face image can be determined based on the second image parameters of the first face image and the brightness value of the first face image.

[0109] Optionally, the terminal may store a third correspondence between the speech parameters of the speech signal, the environmental noise index of the speech signal, and the quality score of the speech signal. Based on this third correspondence, the quality score of the first speech signal can be determined based on the speech parameters of the first speech signal and the environmental noise index of the first speech signal.

[0110] In an exemplary embodiment, the method further includes: determining a reference blank image of the fingerprint sensor; the reference blank image refers to an image acquired by the fingerprint sensor in an initial state where there is no contaminant coverage; the method further includes: determining a real-time blank image of the fingerprint sensor before acquiring a first fingerprint image of a first user through the fingerprint sensor; the real-time blank image refers to an image acquired by the fingerprint sensor in a real-time state where there is no finger pressing; matching the reference blank image with the real-time blank image, and using the contaminant coverage area in the real-time blank image as the area of ​​the fingerprint sensor surface covered by contaminants; and determining the contamination level of the fingerprint sensor based on the area of ​​the fingerprint sensor surface covered by contaminants and the fingerprint sensor surface area.

[0111] The matching of the reference blank image and the real-time blank image can be achieved by using pixel difference calculation to identify pixels in the real-time blank image that have grayscale anomalies relative to the reference blank image. These pixels with grayscale anomalies are then used as contaminant-covered pixels. By counting the number of contaminant-covered pixels, the contaminant-covered area in the real-time blank image can be obtained, and this area can be used as the area of ​​the fingerprint sensor surface covered by contaminants.

[0112] Optionally, the real-time blank image can be the blank image corresponding to the time point before the first fingerprint image is acquired.

[0113] Optionally, the contamination level of the fingerprint sensor = the area of ​​the fingerprint sensor surface covered by contaminants / the surface area of ​​the fingerprint sensor.

[0114] In this embodiment, the quality score of the first fingerprint image is determined based on the first image parameters of the first fingerprint image and the contamination level of the fingerprint sensor used to acquire the first fingerprint image. Similarly, the quality score of the first face image is determined based on the second image parameters of the first face image and the brightness value of the first face image. Furthermore, the quality score of the first voice signal is determined based on the voice parameters of the first voice signal and the environmental noise index of the first voice signal. Thus, the quality scores of different modalities of biometrics can be fully determined through multi-dimensional information, improving the accuracy of the quality scores of different modalities of biometrics. This ensures that the subsequently obtained first weight, second weight, and third weight have high accuracy, thereby accurately determining the release result for the first user. Based on this, the accuracy of identity verification for the first user can be significantly improved.

[0115] In an exemplary embodiment, determining the quality score of the first fingerprint image based on the first image parameters of the first fingerprint image includes:

[0116] The quality score of the first fingerprint image is determined based on the clarity, ridge integrity, effective area ratio, and contamination level of the fingerprint sensor used to acquire the first fingerprint image.

[0117] The above-mentioned determination of the quality score of the first face image based on the second image parameters of the first face image includes:

[0118] The quality score of the first face image is determined based on the illumination equalization, occlusion rate, key point confidence, and brightness value of the first face image.

[0119] The above-mentioned determination of the quality score of the first speech signal based on the speech parameters of the first speech signal includes:

[0120] The quality score of the first speech signal is determined based on the signal-to-noise ratio, effective duration of the speech, and environmental noise index of the first speech signal.

[0121] In an exemplary embodiment, determining the quality score of the first fingerprint image based on the clarity, ridge integrity, effective area ratio, and contamination level of the fingerprint sensor used to acquire the first fingerprint image includes:

[0122] The quality score of the first fingerprint image is determined based on the clarity of the first fingerprint image, the first fingerprint weight, the ridge integrity, the second fingerprint weight, the effective area ratio, the third fingerprint weight, the contamination level of the fingerprint sensor used to acquire the first fingerprint image, and the fourth fingerprint weight.

[0123] Among them, the first fingerprint weight corresponds to the clarity of the first fingerprint image, the second fingerprint weight corresponds to the integrity of the ridge, the third fingerprint weight corresponds to the effective area ratio, and the fourth fingerprint weight corresponds to the contamination level of the fingerprint sensor. Furthermore, the first fingerprint weight is greater than the second fingerprint weight, the second fingerprint weight is greater than the third fingerprint weight, and the third fingerprint weight is greater than the fourth fingerprint weight.

[0124] For example, the quality score of the first fingerprint image = sharpness × first fingerprint weight + ridge integrity × second fingerprint weight + effective area ratio × third fingerprint weight - contamination degree of the fingerprint sensor × fourth fingerprint weight.

[0125] In a straightforward manner, the clarity of the first fingerprint image, the weight of the first fingerprint, the integrity of the ridge, the weight of the second fingerprint, the effective area ratio, the weight of the third fingerprint, the contamination level of the fingerprint sensor, and the weight of the fourth fingerprint can all be normalized values.

[0126] In this embodiment, clarity is a crucial foundation for first fingerprint feature extraction. Low clarity of the first fingerprint image makes it difficult to extract reliable first fingerprint features. Similarly, incomplete ridges and small effective area ratios in the first fingerprint image can also lead to varying degrees of difficulty in first fingerprint feature extraction. Therefore, in determining the quality score of the first fingerprint image, the clarity of the first fingerprint image should be the primary consideration, followed by the different effects of ridge integrity, effective area ratio, and fingerprint sensor contamination. Based on this, the accuracy of determining the quality score of the first fingerprint image can be significantly improved.

[0127] In an exemplary embodiment, determining the quality score of the first face image based on the illumination equalization, occlusion rate, keypoint confidence, and brightness value of the first face image includes:

[0128] The quality score of the first face image is determined based on the illumination equalization, first face weight, occlusion rate, second face weight, key point confidence, third face weight, brightness value of the first face image, and fourth face weight of the first face image.

[0129] Among them, the first face weight corresponds to the illumination balance of the first face image, the second face weight corresponds to the occlusion rate, the third face weight corresponds to the key point confidence, and the fourth face weight corresponds to the brightness value of the first face image. Furthermore, the third face weight is greater than the second face weight, the second face weight is greater than the first face weight, and the first face weight is greater than the fourth face weight.

[0130] For example, the quality score of the first face image = illumination equalization × first face weight - occlusion rate × second face weight + key point confidence × third face weight + brightness value of the first face image × fourth face weight.

[0131] In a straightforward manner, the illumination equalization of the first face image, the weight of the first face, the occlusion rate, the weight of the second face, the confidence of the key points, the weight of the third face, the brightness value of the first face image, and the weight of the fourth face can all be normalized values.

[0132] In this embodiment, since the keypoint confidence score is the alignment anchor for the first face image, a low keypoint confidence score indicates poor quality, which will negatively impact the accuracy of the first user's access result. Similarly, occlusion or uneven lighting in the first face image will also make the first face feature extraction more difficult. Therefore, in determining the quality score of the first face image, the keypoint confidence score should be the primary consideration, followed by the effects of occlusion rate, lighting balance, and brightness value. This approach significantly improves the accuracy of determining the quality score of the first face image.

[0133] In an exemplary embodiment, determining the quality score of the first speech signal based on the signal-to-noise ratio, effective duration of the speech, and environmental noise index of the first speech signal includes:

[0134] The quality score of the first speech signal is determined based on the signal-to-noise ratio, first speech weight, effective speech duration, second speech weight, environmental noise index of the first speech signal, and third speech weight.

[0135] The first speech weight corresponds to the signal-to-noise ratio of the first speech signal, the second speech weight corresponds to the effective duration of the speech, and the third speech weight corresponds to the environmental noise index of the first speech signal. Furthermore, the first speech weight is greater than the third speech weight, and the third speech weight is greater than the second speech weight.

[0136] For example, the quality score of the first speech signal = signal-to-noise ratio × first speech weight + effective speech duration × second speech weight - environmental noise index of the first speech signal × third speech weight.

[0137] In a straightforward manner, the signal-to-noise ratio of the first speech signal, the first speech weight, the effective duration of the speech, the second speech weight, the environmental noise index of the first speech signal, and the third speech weight can all be normalized values.

[0138] In this embodiment, since the signal-to-noise ratio (SNR) directly measures the ratio between effective speech information and interference noise in the first speech signal, a low SNR indicates poor quality, which will negatively impact the accuracy of the first user's release result. Similarly, the effective duration of the first speech signal and the environmental noise index will also lead to varying degrees of difficulty in extracting the first speech features. Therefore, in determining the quality score of the first speech signal, the SNR should be considered first, followed by the different effects of the effective duration and environmental noise index. Based on this, the accuracy of determining the quality score of the first speech signal can be significantly improved.

[0139] In an exemplary embodiment, the determination of the release result for the first user based on the first weight, first similarity, second weight, second similarity, third weight, and third similarity includes:

[0140] Based on the first weight, the second weight, and the third weight, the first similarity, the second similarity, and the third similarity are weighted and summed to obtain the fused similarity.

[0141] If the fusion similarity is greater than or equal to the first similarity threshold, then the first user is allowed to proceed.

[0142] The fusion similarity refers to the weighted sum of the first, second, and third similarities, calculated using the first, second, and third weights.

[0143] Since the first weight and the first similarity correspond to the first fingerprint image, the second weight and the second similarity correspond to the first face image, and the third weight and the third similarity correspond to the first speech signal, the fusion similarity = first weight × first similarity + second weight × second similarity + third weight × third similarity.

[0144] In a straightforward manner, if the fusion similarity is less than the first similarity threshold, then the first user will not be allowed to proceed.

[0145] In this embodiment, access to the first user is only granted when the weighted sum of the first similarity, second similarity, and third similarity based on the first weight, second weight, and third weight, i.e., the fused similarity, is greater than or equal to the first similarity threshold. Since the fused similarity can fully reflect the comprehensive similarity between the first user's fingerprint, face, voice, and registered biometric information, the accuracy of identity verification for the first user can be significantly improved based on the fused similarity.

[0146] In one exemplary embodiment, the method further includes:

[0147] A first similarity threshold is determined based on a preset baseline value, the abnormality of the first user's passage behavior, and the security level of the monitored area. The first similarity threshold, with the preset baseline value remaining unchanged, is positively correlated with the abnormality of the first user's passage behavior and the security level of the monitored area, respectively. The monitored area is the area where the first facial image is acquired.

[0148] The preset benchmark value refers to the benchmark value set by the user on the terminal in advance. Simply put, the preset benchmark value reflects the standard value of the first similarity threshold. For example, if a user has higher accuracy requirements for the release results for different users, the preset benchmark value can be set to a larger value.

[0149] In a straightforward manner, the preset baseline value remains unchanged during the determination of the first similarity threshold. That is, the preset baseline value is a fixed value during the determination of the first similarity threshold, while the abnormality of the first user's passage behavior and the security level of the monitored area can be variable values.

[0150] The abnormality of the first user's passage behavior can reflect the degree of deviation between the first user's current passage behavior and historical passage behavior.

[0151] Optionally, the degree of abnormality of the first user's passage behavior can be determined based on the time abnormality of the first user's passage behavior and the degree of abnormality of the first user's terminal device identifier.

[0152] The monitored area can also be understood as the location where the cameras are installed.

[0153] The security level of a monitored area can be determined by the importance or security sensitivity of the people and / or items in the monitored area.

[0154] Optionally, the security level of a monitored area can be determined based on the area type of the monitored area or based on the value of the items stored in the monitored area.

[0155] Optionally, the first similarity threshold = preset baseline value + the abnormality of the first user's passage behavior × the security level of the monitored area.

[0156] In this embodiment, the first similarity threshold is determined based on a preset benchmark value, the abnormality of the first user's access behavior, and the security level of the monitored area. Thus, in access control management scenarios where the first user's access behavior is abnormal and / or the security level of the monitored area is high, the strictness of the access result for the first user can be further improved, thereby significantly improving the accuracy of identity verification for the first user.

[0157] In one exemplary embodiment, the method further includes:

[0158] If the fusion similarity is greater than or equal to the second similarity threshold, obtain the real-time access time and historical passage time of the first user;

[0159] The above-mentioned determination of the first similarity threshold based on a preset baseline value, the anomaly degree of the first user's passage behavior, and the security level of the monitored area includes:

[0160] The real-time access time of the first user is matched with the historical access time to obtain the degree of time deviation between the real-time access time and the historical access time;

[0161] The first similarity threshold is determined based on the preset baseline value, the degree of time deviation, and the security level of the monitored area.

[0162] The second similarity threshold is less than a preset benchmark value. For example, if the preset benchmark value is 80%, the second similarity threshold can be 70%.

[0163] The first user's real-time access time can be understood as the current time point.

[0164] The historical passage time of the first user can be determined by the average of the passage times corresponding to each passage record in the multiple passage records of the first user within the historical time period.

[0165] In this embodiment of the application, the first similarity threshold is defined as: a preset baseline value + the degree of time deviation × the security level of the monitored area.

[0166] In a straightforward manner, if the fusion similarity is greater than or equal to the second similarity threshold, then the first user is considered a candidate user with high credibility but who has not yet been officially approved.

[0167] In a straightforward manner, the time units for the first user's historical access time and real-time access time should be consistent. That is to say, if it is desired that the degree of time deviation only reflects the degree of deviation within 24 hours of a day, then the time units for both historical access time and real-time access time should be specified to within 24 hours.

[0168] For example, suppose the access control system is applied in a corporate park, and suppose the first user is an employee of the corporate park, and the working hours of the corporate park are from 9:00 to 18:00. Then the first user's historical access time is mostly concentrated around 9:00. Therefore, when the first user's real-time access time is 22:00, that is, the time deviation between the first user's real-time access time and the historical access time is large, it indicates that the first user's access behavior is not as expected. Therefore, it is necessary to increase the first similarity threshold used to determine whether the first user can be allowed to pass.

[0169] Optionally, when the terminal is connected to cameras in multiple different monitoring areas, the abnormality of the first user's passage behavior can also be determined by the degree of position deviation of the first user. For example, the first user's real-time access location can be matched with historical passage locations to obtain the degree of position deviation between the real-time access location and the historical passage location. The real-time access location is the monitoring area where the first user is currently located, and the historical passage location can be determined by the number of times each passage record corresponds to a passage location in multiple passage records within a historical time period. Based on this, a first similarity threshold can be determined based on a preset baseline value, the degree of time deviation, the degree of position deviation, and the security level of the monitoring area. In this case, the first similarity threshold = preset baseline value + degree of time deviation × degree of position deviation × security level of the monitoring area.

[0170] In this embodiment, the real-time access time of the first user is matched with the historical access time to obtain the time deviation between the real-time access time and the historical access time. Based on a preset benchmark value, the first similarity threshold is adjusted according to the time deviation. This makes the determination process of the first similarity threshold more flexible, better able to cope with different user access behaviors, and significantly improves the accuracy of identity verification for the first user and the security of the access control management system.

[0171] In an exemplary embodiment, the acquisition of the first user's first fingerprint image, first facial image, and first voice signal includes:

[0172] After receiving the first user's permission request, the system acquires the first user's first fingerprint image, first facial image, first voice signal, and terminal device identifier; the terminal device identifier is the terminal device identifier corresponding to the terminal used by the first user to send the permission request.

[0173] The above methods also include:

[0174] If the fusion similarity is greater than or equal to the second similarity threshold, the preset terminal device identifier of the first user is obtained;

[0175] The above-mentioned determination of the first similarity threshold based on a preset baseline value, the anomaly degree of the first user's passage behavior, and the security level of the monitored area includes:

[0176] Determine the identifier matching result between the terminal device identifier and the preset terminal device identifier;

[0177] Based on the preset benchmark value, the identifier matching result, and the security level of the monitored area, the first similarity threshold is determined.

[0178] Among them, the terminal device identifier refers to the coded identifier that can uniquely identify the terminal used by the first user to send the release request.

[0179] Optionally, if the matching result is marked as a successful match, it will not affect the determination process of the first similarity threshold; if the matching result is marked as a failed match, it will increase the determination process of the first similarity threshold.

[0180] In this embodiment of the application, the first similarity threshold is equal to a preset baseline value plus the identifier matching result multiplied by the security level of the monitored area. When the identifier matching result is a successful match, the identifier matching result is 1; when the identifier matching result is a failed match, the identifier matching result is greater than 1. The specific value can be set by the user in the terminal based on actual needs.

[0181] In this embodiment, the terminal device identifier is matched with a preset terminal device identifier to obtain an identifier matching result. Based on a preset benchmark value, the first similarity threshold is adjusted according to the identifier matching result. This makes the determination process of the first similarity threshold more flexible, avoids the adverse situation of incorrect access due to biometric impersonation, and significantly improves the accuracy of identity verification for the first user and the security of the access control management system.

[0182] In one exemplary embodiment, the method further includes:

[0183] Based on the first facial image, determine the age group of the first user;

[0184] If the age group of the first user is either a child or an elderly person, a first weight adjustment coefficient, a second weight adjustment coefficient, and a third weight adjustment coefficient are generated based on the age group of the first user.

[0185] The above-mentioned decision to grant access to the first user, based on the first weight, first similarity, second weight, second similarity, third weight, and third similarity, includes:

[0186] Based on the first weight adjustment coefficient, the first weight of the first fingerprint image is adjusted to obtain the adjusted first weight;

[0187] Based on the second weight adjustment coefficient, the second weight of the first face image is adjusted to obtain the adjusted second weight;

[0188] Based on the third weight adjustment coefficient, the third weight of the first speech signal is adjusted to obtain the adjusted third weight;

[0189] Based on the adjusted first weight, first similarity, adjusted second weight, second similarity, adjusted third weight, and third similarity, the release result for the first user is determined;

[0190] Specifically, when the first user's age group is the child age group, the adjusted first weight is greater than the unadjusted first weight, the adjusted second weight is less than the unadjusted second weight, and the adjusted third weight is less than the unadjusted third weight; when the first user's age group is the elderly age group, the adjusted first weight is less than the unadjusted first weight, the adjusted second weight is greater than the unadjusted second weight, and the adjusted third weight is greater than the unadjusted third weight.

[0191] Optionally, based on the first face image, the age range of the first user can be determined using an age recognition algorithm. For example, the age range of the first user can be determined by extracting facial features such as facial texture features from the first face image using an age recognition algorithm. For instance, the age range of the first user can be determined by the smoothness of the facial texture, as represented by the facial texture features of the first face image.

[0192] Facial texture features refer to the image features in the first facial image used to characterize the distribution and variation patterns of wrinkles, fine lines, and folds on the surface of the first user's facial skin. Facial texture smoothness is used to quantify the smoothness of the first user's facial skin. It is easy to understand that facial texture smoothness is negatively correlated with the number and depth of wrinkles on the human skin surface.

[0193] Optionally, facial texture smoothness can be determined by factors such as the direction of gradient change, the degree of gradient change, and the contrast of facial texture features.

[0194] For example, if the facial texture smoothness of the first user's face, as represented by the facial texture features of the first face image, is greater than or equal to a first smoothness threshold, then the first user's age group can be considered to be the child age group; if the facial texture smoothness of the first user's face, as represented by the facial texture features of the first face image, is less than or equal to a second smoothness threshold, then the first user's age group can be considered to be the elderly age group, where the first smoothness threshold is greater than the second smoothness threshold. Meanwhile, considering that between the child and elderly age groups, there exists a middle-aged and young adult age group where the facial texture smoothness is greater than the second smoothness threshold but less than the first smoothness threshold, and no weight adjustment is needed, the first smoothness threshold and the second smoothness threshold are not equal and there is a certain difference between them. The difference between the first smoothness threshold and the second smoothness threshold can be greater than or equal to a difference threshold.

[0195] Further optionally, the age range corresponding to the children's age group can be defined based on the magnitude of the first smoothness threshold, and the age range corresponding to the elderly age group can be defined based on the second smoothness threshold. This application does not impose any limitations on this.

[0196] For example, an age recognition algorithm can be an age recognition algorithm based on a convolutional neural network.

[0197] In this embodiment, by determining the age group of the first user, and adjusting the weights of different modalities of biometric features for the first user in the child age group and the first user in the elderly age group respectively, specifically considering that the face and voice of the first user in the child age group are still in a period of rapid growth and change, while fingerprint features are the most stable biometric features of the first user in the child age group, and considering that the thinning of the stratum corneum in the elderly age group leads to reduced sebum secretion, the adjustment direction of the first weight of the first fingerprint image is different for the first user in different age groups. Thus, the user release method provided in this embodiment can have higher user targeting and flexibility, and can more accurately determine the release result for the first user. Based on this, the accuracy of identity verification for the first user can be significantly improved.

[0198] In one exemplary embodiment, the user release method provided in this application includes the following steps:

[0199] Acquire the first fingerprint image, the first facial image, and the first voice signal of the first user;

[0200] The quality score of the first fingerprint image is determined based on the clarity of the first fingerprint image, the first fingerprint weight, the ridge integrity, the second fingerprint weight, the effective area ratio, the third fingerprint weight, the contamination level of the fingerprint sensor used to acquire the first fingerprint image, and the fourth fingerprint weight. The first fingerprint weight corresponds to the clarity of the first fingerprint image, the second fingerprint weight corresponds to the ridge integrity, the third fingerprint weight corresponds to the effective area ratio, and the fourth fingerprint weight corresponds to the contamination level of the fingerprint sensor. Furthermore, the first fingerprint weight is greater than the second fingerprint weight, the second fingerprint weight is greater than the third fingerprint weight, and the third fingerprint weight is greater than the fourth fingerprint weight.

[0201] The quality score of the first face image is determined based on the illumination equalization, first face weight, occlusion rate, second face weight, keypoint confidence, third face weight, brightness value of the first face image, and fourth face weight of the first face image. Among them, the first face weight corresponds to the illumination equalization of the first face image, the second face weight corresponds to the occlusion rate, the third face weight corresponds to the keypoint confidence, and the fourth face weight corresponds to the brightness value of the first face image. Furthermore, the third face weight is greater than the second face weight, the second face weight is greater than the first face weight, and the first face weight is greater than the fourth face weight.

[0202] The quality score of the first speech signal is determined based on the signal-to-noise ratio, first speech weight, effective speech duration, second speech weight, environmental noise index of the first speech signal, and third speech weight of the first speech signal. The first speech weight corresponds to the signal-to-noise ratio of the first speech signal, the second speech weight corresponds to the effective speech duration, and the third speech weight corresponds to the environmental noise index of the first speech signal. Furthermore, the first speech weight is greater than the third speech weight, and the third speech weight is greater than the second speech weight.

[0203] Based on the first facial image, determine the age group of the first user;

[0204] Based on the quality scores of the first fingerprint image, the first face image, and the first voice signal, a first weight of the first fingerprint image, a second weight of the first face image, and a third weight of the first voice signal are determined respectively.

[0205] If the age group of the first user is either a child or an elderly person, a first weight adjustment coefficient, a second weight adjustment coefficient, and a third weight adjustment coefficient are generated based on the age group of the first user.

[0206] Based on the first weight adjustment coefficient, the first weight of the first fingerprint image is adjusted to obtain the adjusted first weight;

[0207] Based on the second weight adjustment coefficient, the second weight of the first face image is adjusted to obtain the adjusted second weight;

[0208] Based on the third weight adjustment coefficient, the third weight of the first speech signal is adjusted to obtain the adjusted third weight; wherein, when the age group of the first user is a child, the adjusted first weight is greater than the original first weight, the adjusted second weight is less than the original second weight, and the adjusted third weight is less than the original third weight; when the age group of the first user is an elderly person, the adjusted first weight is less than the original first weight, the adjusted second weight is greater than the original second weight, and the adjusted third weight is greater than the original third weight.

[0209] The first similarity between the first fingerprint image and the second fingerprint image of the second user, the second similarity between the first face image and the second face image of the second user, and the third similarity between the first voice signal and the second voice signal of the second user are determined respectively, wherein the second user is any user stored.

[0210] The fusion similarity is obtained by weighting and summing the adjusted first weight, first similarity, adjusted second weight, second similarity, adjusted third weight, and third similarity.

[0211] If the fusion similarity is greater than or equal to the similarity threshold, then the first user is allowed access.

[0212] In this embodiment of the application, the fusion similarity = adjusted first weight × first similarity + adjusted second weight × second similarity + adjusted third weight × third similarity.

[0213] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0214] Based on the same inventive concept, this application also provides a user release device for implementing the user release method described above. The solution provided by this device is similar to the implementation described in the above method; therefore, the specific limitations in one or more user release device embodiments provided below can be found in the limitations of the user release method described above, and will not be repeated here.

[0215] In one exemplary embodiment, such as Figure 3 As shown, a user release device is provided, including: an acquisition module 302, a first determination module 304, a second determination module 306, a third determination module 308, and a fourth determination module 310, wherein:

[0216] The acquisition module 302 is used to acquire the first fingerprint image, the first face image, and the first voice signal of the first user.

[0217] The first determining module 304 is used to determine the quality scores of the first fingerprint image, the first face image, and the first voice signal, respectively.

[0218] The second determining module 306 is used to determine the first weight of the first fingerprint image, the second weight of the first face image, and the third weight of the first voice signal based on the quality scores of the first fingerprint image, the first face image, and the first voice signal, respectively.

[0219] The third determining module 308 is used to determine the first similarity between the first fingerprint image and the second fingerprint image of the second user, the second similarity between the first face image and the second face image of the second user, and the third similarity between the first voice signal and the second voice signal of the second user, wherein the second user is any user stored.

[0220] The fourth determining module 310 is used to determine the release result for the first user based on the first weight, the first similarity, the second weight, the second similarity, the third weight, and the third similarity.

[0221] In an exemplary embodiment, in determining the quality scores of the first fingerprint image, the first face image, and the first voice signal, the first determining module 304 is specifically configured to determine the quality score of the first fingerprint image based on first image parameters of the first fingerprint image; determine the quality score of the first face image based on second image parameters of the first face image; and determine the quality score of the first voice signal based on voice parameters of the first voice signal.

[0222] In an exemplary embodiment, the first image parameters of the first fingerprint image include the sharpness, ridge integrity, and effective area ratio of the first fingerprint image; the second image parameters of the first face image include the illumination uniformity, occlusion rate, and key point confidence of the first face image; the voice parameters of the first voice signal include the signal-to-noise ratio and effective duration of the voice signal; regarding determining the quality score of the first fingerprint image based on the first image parameters of the first fingerprint image, the first determining module 304 is specifically used to determine the quality score of the first fingerprint image based on the sharpness, ridge integrity, and effective area ratio of the first fingerprint image, wherein the quality score of the first fingerprint image is positively correlated with the sharpness, ridge integrity, and effective area ratio of the first fingerprint image, respectively; regarding determining the quality score of the first fingerprint image based on the first face ... respectively; regarding determining the quality score of the first fingerprint image based on the first face image, the first determining module 304 is specifically used to determine the quality score of the first fingerprint image based on the sharpness, ridge integrity, and effective area ratio of the first fingerprint image, respectively. Regarding the determination of the quality score of the first face image based on the second image parameters, the first determining module 304 is specifically used to determine the quality score of the first face image based on the illumination equalization, occlusion rate, and key point confidence of the first face image. The quality score of the first face image is positively correlated with the illumination equalization and key point confidence of the first face image, and negatively correlated with the occlusion rate. Regarding the determination of the quality score of the first speech signal based on the speech parameters of the first speech signal, the first determining module 304 is specifically used to determine the quality score of the first speech signal based on the signal-to-noise ratio and effective speech duration of the first speech signal. The quality score of the first speech signal is positively correlated with the signal-to-noise ratio and effective speech duration of the first speech signal, respectively.

[0223] In an exemplary embodiment, if the quality score of the first fingerprint image is less than a first score threshold, the quality score of the first face image is less than a second score threshold, or the quality score of the first voice signal is less than a third score threshold, the acquisition module 302 re-executes the acquisition of the first user's first fingerprint image, first face image, and first voice signal, and the first determination module 304 re-executes the steps of determining the quality scores of the first fingerprint image, first face image, and first voice signal respectively.

[0224] In an exemplary embodiment, regarding determining the quality score of a first fingerprint image based on first image parameters of the first fingerprint image, the first determining module 304 is specifically used to determine the quality score of the first fingerprint image based on the first image parameters of the first fingerprint image and the contamination level of the fingerprint sensor used to acquire the first fingerprint image; the quality score of the first fingerprint image is negatively correlated with the contamination level of the fingerprint sensor. Regarding determining the quality score of a first face image based on second image parameters of the first face image, the first determining module 304 is specifically used to determine the quality score of the first face image based on the second image parameters of the first face image and the brightness value of the first face image; the quality score of the first face image is positively correlated with the brightness value of the first face image. Regarding determining the quality score of a first voice signal based on voice parameters of the first voice signal, the first determining module 304 is specifically used to determine the quality score of the first voice signal based on the voice parameters of the first voice signal and the environmental noise index of the first voice signal; the quality score of the first voice signal is negatively correlated with the environmental noise index of the first voice signal.

[0225] In an exemplary embodiment, in determining the release result for the first user based on the first weight, first similarity, second weight, second similarity, third weight, and third similarity, the fourth determining module 310 is specifically used to perform a weighted summation of the first similarity, second similarity, and third similarity based on the first weight, second weight, and third weight to obtain a fused similarity; and determine that the first user is allowed to pass if the fused similarity is greater than or equal to the first similarity threshold.

[0226] In an exemplary embodiment, the fourth determining module 310 is further configured to determine a first similarity threshold based on a preset benchmark value, the abnormality of the first user's passage behavior, and the security level of the monitoring area; the first similarity threshold, with the preset benchmark value remaining unchanged, is positively correlated with the abnormality of the first user's passage behavior and the security level of the monitoring area, respectively; the monitoring area is the area where the first facial image is acquired.

[0227] In one exemplary embodiment, the user release device provided in this application includes:

[0228] The acquisition module 302 is used to acquire the first fingerprint image, the first face image, and the first voice signal of the first user;

[0229] The first determining module 304 is used to determine the quality score of the first fingerprint image based on the clarity of the first fingerprint image, the first fingerprint weight, the ridge integrity, the second fingerprint weight, the effective area ratio, the third fingerprint weight, the contamination level of the fingerprint sensor used to acquire the first fingerprint image, and the fourth fingerprint weight; wherein, the first fingerprint weight corresponds to the clarity of the first fingerprint image, the second fingerprint weight corresponds to the ridge integrity, the third fingerprint weight corresponds to the effective area ratio, the fourth fingerprint weight corresponds to the contamination level of the fingerprint sensor, and the first fingerprint weight is greater than the second fingerprint weight, the second fingerprint weight is greater than the third fingerprint weight, and the third fingerprint weight is greater than the fourth fingerprint weight;

[0230] The first determining module 304 is further configured to determine the quality score of the first face image based on the illumination equalization, first face weight, occlusion rate, second face weight, key point confidence, third face weight, brightness value of the first face image, and fourth face weight of the first face image; wherein, the first face weight corresponds to the illumination equalization of the first face image, the second face weight corresponds to the occlusion rate, the third face weight corresponds to the key point confidence, and the fourth face weight corresponds to the brightness value of the first face image, and the third face weight is greater than the second face weight, the second face weight is greater than the first face weight, and the first face weight is greater than the fourth face weight;

[0231] The first determining module 304 is further configured to determine the quality score of the first speech signal based on the signal-to-noise ratio, the first speech weight, the effective duration of the speech, the second speech weight, the environmental noise index of the first speech signal, and the third speech weight; wherein the first speech weight corresponds to the signal-to-noise ratio of the first speech signal, the second speech weight corresponds to the effective duration of the speech, the third speech weight corresponds to the environmental noise index of the first speech signal, and the first speech weight is greater than the third speech weight, and the third speech weight is greater than the second speech weight.

[0232] The second determining module 306 is used to determine the age range of the first user based on the first facial image;

[0233] The second determining module 306 is further configured to determine the first weight of the first fingerprint image, the second weight of the first face image, and the third weight of the first voice signal based on the quality scores of the first fingerprint image, the first face image, and the first voice signal, respectively.

[0234] The second determining module 306 is further configured to generate a first weight adjustment coefficient, a second weight adjustment coefficient, and a third weight adjustment coefficient based on the age of the first user when the age range of the first user is a child age range or an elderly age range.

[0235] The second determining module 306 is further configured to adjust the first weight of the first fingerprint image based on the first weight adjustment coefficient to obtain the adjusted first weight; adjust the second weight of the first face image based on the second weight adjustment coefficient to obtain the adjusted second weight; and adjust the third weight of the first voice signal based on the third weight adjustment coefficient to obtain the adjusted third weight. Wherein, when the first user's age group is a child, the adjusted first weight is greater than the unadjusted first weight, the adjusted second weight is less than the unadjusted second weight, and the adjusted third weight is less than the unadjusted third weight; when the first user's age group is an elderly person, the adjusted first weight is less than the unadjusted first weight, the adjusted second weight is greater than the unadjusted second weight, and the adjusted third weight is greater than the unadjusted third weight.

[0236] The third determining module 308 is used to determine the first similarity between the first fingerprint image and the second fingerprint image of the second user, the second similarity between the first face image and the second face image of the second user, and the third similarity between the first voice signal and the second voice signal of the second user, wherein the second user is any user stored.

[0237] The fourth determining module 310 is used to perform a weighted summation based on the adjusted first weight, first similarity, adjusted second weight, second similarity, adjusted third weight, and third similarity to obtain a fused similarity; if the fused similarity is greater than or equal to the similarity threshold, it determines to allow access for the first user.

[0238] Each module in the aforementioned user release device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in the processor of a computer device in hardware form or independent of it, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0239] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data related to the user release method. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements a user release method.

[0240] In one exemplary embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, Near Field Communication (NFC), or other technologies. When the computer program is executed by the processor, it implements a user access method. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0241] Those skilled in the art will understand that Figure 5The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0242] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0243] In the above computer device embodiments, any of the computer devices can be a camera.

[0244] In one exemplary embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above-described method embodiments.

[0245] In one exemplary embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0246] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0247] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.

[0248] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0249] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A user release method, characterized in that, include: Acquire the first fingerprint image, the first facial image, and the first voice signal of the first user; The first image parameters of the first fingerprint image include the clarity, ridge integrity, and effective area ratio of the first fingerprint image; the second image parameters of the first face image include the illumination uniformity, occlusion rate, and key point confidence of the first face image; the speech parameters of the first speech signal include the signal-to-noise ratio and effective duration of the speech signal. The quality scores of the first fingerprint image, the first face image, and the first voice signal are determined respectively. Based on the quality scores of the first fingerprint image, the first face image, and the first voice signal, a first weight of the first fingerprint image, a second weight of the first face image, and a third weight of the first voice signal are determined respectively. The first similarity between the first fingerprint image and the second fingerprint image of the second user, the second similarity between the first face image and the second face image of the second user, and the third similarity between the first voice signal and the second voice signal of the second user are determined respectively, wherein the second user is any user stored. Based on the first weight, the first similarity, the second weight, the second similarity, the third weight, and the third similarity, a release result is determined for the first user; The step of determining the quality scores of the first fingerprint image, the first face image, and the first voice signal includes: Based on the first image parameters of the first fingerprint image, the quality score of the first fingerprint image is determined; based on the second image parameters of the first face image, the quality score of the first face image is determined; based on the voice parameters of the first voice signal, the quality score of the first voice signal is determined. Determining the quality score of the first fingerprint image based on the first image parameters of the first fingerprint image includes: Based on the clarity, ridge integrity, and effective area ratio of the first fingerprint image, the quality score of the first fingerprint image is determined. The quality score of the first fingerprint image is positively correlated with the clarity, ridge integrity, and effective area ratio of the first fingerprint image, respectively. Determining the quality score of the first face image based on the second image parameters of the first face image includes: Based on the illumination evenness, occlusion rate, and key point confidence of the first face image, the quality score of the first face image is determined. The quality score of the first face image is positively correlated with the illumination evenness and key point confidence of the first face image, respectively, and negatively correlated with the occlusion rate. Determining the quality score of the first speech signal based on its speech parameters includes: Based on the signal-to-noise ratio and effective duration of the first speech signal, the quality score of the first speech signal is determined. The quality score of the first speech signal is positively correlated with the signal-to-noise ratio and effective duration of the first speech signal, respectively.

2. The method according to claim 1, characterized in that, The acquisition of the first user's first fingerprint image, first facial image, and first voice signal includes: The system acquires the first fingerprint image of the first user through a fingerprint sensor, the first facial image of the first user through a camera, and the first voice signal of the first user through a microphone.

3. The method according to claim 1, characterized in that, The sum of the first weight, the second weight, and the third weight is 1.

4. The method according to claim 1, characterized in that, The method further includes: If the quality score of the first fingerprint image is less than a first score threshold, the quality score of the first face image is less than a second score threshold, or the quality score of the first voice signal is less than a third score threshold, the steps of acquiring the first fingerprint image, the first face image, and the first voice signal of the first user, and determining the quality scores of the first fingerprint image, the first face image, and the first voice signal respectively are executed.

5. The method according to claim 2, characterized in that, Determining the quality score of the first fingerprint image based on the first image parameters of the first fingerprint image includes: Based on the first image parameters of the first fingerprint image and the contamination level of the fingerprint sensor used to acquire the first fingerprint image, the quality score of the first fingerprint image is determined; the quality score of the first fingerprint image is negatively correlated with the contamination level of the fingerprint sensor. The step of determining the quality score of the first face image based on the second image parameters of the first face image includes: Based on the second image parameters of the first face image and the brightness value of the first face image, the quality score of the first face image is determined; there is a positive correlation between the quality score of the first face image and the brightness value of the first face image. Determining the quality score of the first speech signal based on its speech parameters includes: Based on the speech parameters of the first speech signal and the environmental noise index of the first speech signal, the quality score of the first speech signal is determined; the quality score of the first speech signal is negatively correlated with the environmental noise index of the first speech signal.

6. The method according to claim 1, characterized in that, The step of determining the release result for the first user based on the first weight, the first similarity, the second weight, the second similarity, the third weight, and the third similarity includes: Based on the first weight, the second weight, and the third weight, the first similarity, the second similarity, and the third similarity are weighted and summed to obtain the fused similarity. If the fusion similarity is greater than or equal to the first similarity threshold, then it is determined that the first user can be allowed to proceed.

7. The method according to claim 6, characterized in that, The method further includes: The first similarity threshold is determined based on a preset benchmark value, the abnormality of the first user's passage behavior, and the security level of the monitoring area. The first similarity threshold, with the preset benchmark value remaining unchanged, is positively correlated with the abnormality of the first user's passage behavior and the security level of the monitoring area, respectively. The monitoring area is the area where the first facial image is acquired.

8. A user release device, characterized in that, include: The acquisition module is used to acquire the first fingerprint image, the first facial image, and the first voice signal of the first user; The first image parameters of the first fingerprint image include the clarity, ridge integrity, and effective area ratio of the first fingerprint image; the second image parameters of the first face image include the illumination uniformity, occlusion rate, and key point confidence of the first face image; the speech parameters of the first speech signal include the signal-to-noise ratio and effective duration of the speech signal. The first determining module is used to determine the quality scores of the first fingerprint image, the first face image, and the first voice signal, respectively. The second determining module is used to determine a first weight of the first fingerprint image, a second weight of the first face image, and a third weight of the first voice signal based on the quality scores of the first fingerprint image, the first face image, and the first voice signal, respectively. The third determining module is used to determine the first similarity between the first fingerprint image and the second fingerprint image of the second user, the second similarity between the first face image and the second face image of the second user, and the third similarity between the first voice signal and the second voice signal of the second user, wherein the second user is any of the stored users; The fourth determining module is used to determine the release result for the first user based on the first weight, the first similarity, the second weight, the second similarity, the third weight, and the third similarity; Specifically, in determining the quality scores of the first fingerprint image, the first face image, and the first voice signal, the first determining module is used to determine the quality score of the first fingerprint image based on the first image parameters of the first fingerprint image; and to determine the quality score of the first face image based on the second image parameters of the first face image. Based on the speech parameters of the first speech signal, determine the quality score of the first speech signal; Regarding determining the quality score of the first fingerprint image based on the first image parameters of the first fingerprint image, the first determining module is specifically used to determine the quality score of the first fingerprint image based on the clarity, ridge integrity, and effective area ratio of the first fingerprint image. The quality score of the first fingerprint image is positively correlated with the clarity, ridge integrity, and effective area ratio of the first fingerprint image, respectively. Regarding determining the quality score of the first face image based on the second image parameters of the first face image, the first determining module is specifically used to determine the quality score of the first face image based on the illumination evenness, occlusion rate, and keypoint confidence of the first face image. The quality score of the first face image is positively correlated with the illumination evenness and keypoint confidence of the first face image, respectively, and negatively correlated with the occlusion rate. Regarding determining the quality score of the first speech signal based on the speech parameters of the first speech signal, the first determining module is specifically used to determine the quality score of the first speech signal based on the signal-to-noise ratio and effective speech duration of the first speech signal. The quality score of the first speech signal is positively correlated with the signal-to-noise ratio and effective speech duration of the first speech signal, respectively.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Attendance identification method and device

    CN109102582A

  • Payment verification method and device, terminal, server and storage medium

    CN115829575A