Retail cabinet pick-up method and device supporting voice recognition and face recognition

Through the multimodal biometric fusion and dynamic verification strategies of speech recognition and face recognition, the security and environmental adaptability problems of smart retail cabinets are solved, the balance between high security and barrier-free operation is achieved, and the reliability of application scenarios is expanded.

CN120472582APending Publication Date: 2025-08-12SHANGHAI QUZHI NETWORK TECH CO LTD

Patent Information

Application Number
CN202510754540.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-08-12

AI Technical Summary

Technical Problem

The existing smart retail cabinet identity verification technology has the problem of single biometric vulnerability and poor environmental adaptability, and cannot balance high security and barrier-free operations, resulting in bottlenecks in the fields of medical supplies and retail of valuable goods.

Method used

The multimodal biometric fusion and dynamic verification strategy of speech recognition and face recognition is adopted, and the microphone array and binocular camera are activated through infrared sensors, and data is collected in combination with directional beamforming and 3D structured light technology, voiceprint and facial image verification is performed, verification intensity is dynamically adjusted and supplementary verification is performed.

Benefits of technology

It realizes security in high-risk scenarios and convenient operation in ordinary scenarios, taking into account the protection of biometric data privacy and operational robustness in complex environments, and shortens the response time to within 2 seconds, expanding the application reliability of smart retail cabinets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120472582A_ABST
    Figure CN120472582A_ABST
Patent Text Reader

Abstract

The invention discloses a retail cabinet pick-up method and device supporting voice recognition and face recognition. According to the method, an infrared sensor detects that a user enters an interaction area in real time, a voice and face bimodal acquisition module is synchronously activated, a voice instruction is captured by using a directional beam forming technology, and a high-precision face image is obtained in combination with 3D structured light projection. In the verification stage, the voice recognition module extracts voiceprint features and executes living body detection, and the face recognition module achieves anti-fake verification through depth information analysis. Dynamically selecting a dual verification mode or a single-mode verification mode according to a preset security policy, and comprehensively judging a verification result through a weight fusion algorithm; when verification fails, a supplementary verification mechanism is triggered based on failure reasons, and the fault-tolerant capability is improved through cross-modal data compensation and environment self-adaptive adjustment. The problems that a single biological recognition technology is prone to being attacked and poor in environmental adaptability are solved, and balance between safety and operation efficiency is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent retail equipment, and in particular to a retail cabinet pickup method and device supporting voice recognition and face recognition. Background Art

[0002] Current authentication technologies for smart retail lockers primarily focus on single biometric features or traditional digital credentials. In scenarios like express delivery and unmanned vending, pickup codes and QR code scanning remain the mainstream authentication methods, generating a one-time digital credential via text message or mobile app to verify identity. With the increasing adoption of biometric technology, some devices are beginning to incorporate facial recognition modules, using cameras to capture facial features for comparison; other systems are experimenting with voice interaction solutions, enabling contactless operation by analyzing the user's voiceprint or voice commands. These technologies have streamlined the pickup process to a certain extent, and the demand for contactless authentication has driven the implementation of related solutions.

[0003] However, the existing technology system still has insurmountable flaws. Traditional digital credentials rely on the stability of network transmission, which can easily lead to operational interruptions in signal blind spots or when the device is offline, and there are security risks such as credential forwarding and screenshot fraud. Single biometric verification solutions are limited by environmental interference and forgery attacks: pure facial recognition is sensitive to lighting conditions, with recognition rates plummeting in backlight or occlusion, and cannot resist deception using high-resolution photos or dynamic videos; pure voice recognition has an increased error rate in the presence of background noise, and it is difficult to distinguish between real users and recorded playbacks. More critically, existing solutions lack a multimodal collaborative mechanism and are unable to dynamically adjust verification strength based on scenario risks. Mandatory two-factor authentication reduces operational efficiency for people without barriers, while single verification is difficult to meet the needs of high-security scenarios. This technical contradiction has led to application bottlenecks in existing retail counters in areas such as medical supply storage and access and the retail of valuable goods.

[0004] Therefore, how to invent and develop a retail locker pickup method that supports voice recognition and facial recognition, allowing users to pick up items more safely, conveniently and inclusively, while taking into account high security and barrier-free operation, has become an urgent problem that needs to be solved. Summary of the Invention

[0005] To this end, the present invention provides a retail counter pickup method and device that supports both voice and facial recognition. By integrating multimodal biometric features with a dynamic verification strategy, this method effectively addresses the inherent security and environmental adaptability limitations of single biometric technologies. The dual liveness detection mechanism of voice and facial recognition proactively protects against counterfeit attacks. Combined with adaptive verification mode switching based on environmental awareness, this method ensures safety in high-risk scenarios, such as medical supplies, while providing seamless and convenient operation in more common scenarios.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a retail cabinet pickup method supporting voice recognition and face recognition, comprising:

[0007] When the infrared sensor detects that the user enters the set interaction range, it activates the microphone array of the voice recognition module and the binocular camera of the face recognition module, and initializes the multimodal data acquisition protocol;

[0008] Based on the multimodal data acquisition protocol, the voice recognition module captures the user's voice commands through directional beamforming technology; the face recognition module acquires the user's facial image through 3D structured light projection technology;

[0009] Based on the user's voice command, voice recognition verification is performed to obtain a voiceprint matching score and a voice liveness detection result; based on the user's facial image, face recognition verification is performed to obtain a face matching score and a face liveness detection result;

[0010] According to the retail counter verification mode, combined with the voice recognition verification results and the face recognition verification results, the recognition verification conclusion is output; if the recognition verification passes, the counter door is opened and the goods are picked up; if the recognition verification fails, the next step is processed;

[0011] When the identification verification fails, the modal data that has passed the verification is frozen and the reason for the verification failure is recorded; based on the reason for the verification failure, a supplementary verification is performed; if the supplementary verification passes, the cabinet door is opened and the pickup is completed; if the supplementary verification fails, the pickup operation is exited.

[0012] As a preferred solution for a retail locker pickup method that supports voice recognition and face recognition, in the process of obtaining the user's voice instructions and the user's facial image, synchronous collection is performed through timestamps to ensure consistency in the time and space of data collection.

[0013] As a preferred solution for retail locker pickup that supports voice recognition and facial recognition, during the voice recognition verification process:

[0014] Performing noise reduction processing on the user's voice command and extracting Mel-frequency cepstral coefficients as voiceprint features; performing dynamic time warping matching on the voiceprint features and the set encrypted voiceprint template to generate a voiceprint matching score and voice liveness detection results;

[0015] During the face recognition verification process:

[0016] Based on the user's facial image, facial micro-expression changes and pupil reflection features are calculated to obtain the face liveness detection result; after the liveness detection is passed, the sampling occlusion adaptive algorithm is used to extract the features of the area around the eyes, and the face matching score is obtained through cosine similarity calculation.

[0017] As a preferred solution for a retail counter pickup method that supports voice recognition and facial recognition, in the process of outputting the recognition and verification conclusion according to the retail counter verification mode, combining the voice recognition verification results and the facial recognition verification results:

[0018] If the retail counter verification mode is high security mode, the voiceprint matching score must be ≥ 0.9 and the face matching score must be ≥ 0.95, and both voice liveness detection and face liveness detection must be passed;

[0019] If the retail counter verification mode is the convenient mode, the voiceprint matching score is required to be ≥ 0.85 or the face matching score is required to be ≥ 0.9, and the liveness detection of the other modality must be passed.

[0020] As a preferred solution for a retail locker pickup method that supports voice recognition and facial recognition, during the supplementary verification process based on the verification failure reason:

[0021] If the voice recognition verification fails, the microphone array gain is adjusted and the user is prompted to repeat the voice command;

[0022] If facial recognition verification fails, infrared fill light is activated and the user is guided to adjust the facial angle;

[0023] After the supplementary verification is passed, the initial verification and supplementary verification results are integrated and the cabinet opening operation is performed according to the downgraded security policy.

[0024] As a preferred solution for retail locker pickup that supports voice recognition and facial recognition, during the user's voice recognition verification and facial recognition verification process, regardless of whether the verification is passed or not, the collected voiceprint spectrum, facial feature points and verification log will be sharded and encrypted and stored in a local trusted execution environment; after the locker opening operation is completed, the generated biometric intermediate data will be destroyed.

[0025] The present invention also provides a retail cabinet pickup device supporting voice recognition and face recognition, comprising:

[0026] The recognition module activation unit is used to activate the microphone array of the voice recognition module and the binocular camera of the face recognition module when the infrared sensor detects that the user enters the set interaction range, and initialize the multimodal data acquisition protocol;

[0027] A user voice command and facial image acquisition unit is configured to capture user voice commands using directional beamforming technology based on the multimodal data acquisition protocol; and the facial recognition module acquires user facial images using 3D structured light projection technology;

[0028] A voice recognition verification and face recognition verification unit, configured to perform voice recognition verification based on the user's voice command to obtain a voiceprint matching score and a voice liveness detection result; and perform face recognition verification based on the user's facial image to obtain a face matching score and a face liveness detection result;

[0029] The recognition and verification conclusion acquisition and processing unit is used to output the recognition and verification conclusion based on the retail counter verification mode, combined with the voice recognition verification results and the face recognition verification results; if the recognition and verification pass, the cabinet door is opened and the goods are collected; if the recognition and verification fails, the next step of processing is carried out;

[0030] The supplementary verification unit is used to freeze the verified modal data and record the reason for the verification failure when the identification verification fails; perform supplementary verification based on the reason for the verification failure; if the supplementary verification passes, open the cabinet door and complete the pickup; if the supplementary verification fails, exit the pickup operation.

[0031] As a preferred solution for a retail cabinet pickup device that supports voice recognition and face recognition, the user voice command and facial image acquisition unit uses timestamps to synchronize the acquisition during the process of acquiring the user voice command and the user facial image, so that the time and space of data acquisition remain consistent.

[0032] As a preferred solution for a retail cabinet pickup device supporting voice recognition and face recognition, in the voice recognition verification and face recognition verification unit, during the voice recognition verification process:

[0033] Performing noise reduction processing on the user's voice command and extracting Mel-frequency cepstral coefficients as voiceprint features; performing dynamic time warping matching on the voiceprint features and the set encrypted voiceprint template to generate a voiceprint matching score and voice liveness detection results;

[0034] During the face recognition verification process:

[0035] Based on the user's facial image, facial micro-expression changes and pupil reflection features are calculated to obtain the face liveness detection result; after the liveness detection is passed, the sampling occlusion adaptive algorithm is used to extract the features of the area around the eyes, and the face matching score is obtained through cosine similarity calculation.

[0036] As a preferred solution for a retail counter pickup device supporting voice recognition and face recognition, the identification verification conclusion acquisition and processing unit, in the process of outputting the identification verification conclusion according to the retail counter verification mode and combining the voice recognition verification results and the face recognition verification results:

[0037] If the retail counter verification mode is high security mode, the voiceprint matching score must be ≥ 0.9 and the face matching score must be ≥ 0.95, and both voice liveness detection and face liveness detection must be passed;

[0038] If the retail counter verification mode is the convenient mode, the voiceprint matching score is required to be ≥ 0.85 or the face matching score is required to be ≥ 0.9, and the liveness detection of the other modality must be passed.

[0039] As a preferred solution for a retail cabinet pickup device supporting voice recognition and face recognition, the supplementary verification unit, during the supplementary verification process based on the verification failure reason:

[0040] If the voice recognition verification fails, the microphone array gain is adjusted and the user is prompted to repeat the voice command;

[0041] If facial recognition verification fails, infrared fill light is activated and the user is guided to adjust the facial angle;

[0042] After the supplementary verification is passed, the initial verification and supplementary verification results are integrated and the cabinet opening operation is performed according to the downgraded security policy.

[0043] As a preferred solution for a retail locker pickup device that supports voice recognition and facial recognition, during the user's voice recognition verification and facial recognition verification process, regardless of whether the verification is passed or not, the collected voiceprint spectrum, facial feature points and verification log will be sharded and encrypted and stored in a local trusted execution environment; after the locker opening operation is completed, the generated biometric intermediate data will be destroyed.

[0044] The present invention has the following advantages: when the infrared sensor detects that a user enters a set interaction range, the present invention activates the microphone array of the voice recognition module and the binocular camera of the face recognition module, and initializes the multimodal data acquisition protocol; based on the multimodal data acquisition protocol, the voice recognition module captures the user's voice command through directional beamforming technology; the face recognition module obtains the user's facial image through 3D structured light projection technology; based on the user's voice command, voice recognition verification is performed to obtain a voiceprint matching score and a voice liveness detection result; based on the user's facial image, face recognition verification is performed to obtain a face matching score and a face liveness detection result; according to the retail cabinet verification mode, the voice recognition verification result and the face recognition verification result are combined to output the recognition verification conclusion; if the recognition verification passes, the cabinet door is opened and the goods are picked up; if the recognition verification fails, the next step is carried out; when the recognition verification fails, the verified modal data is frozen and the reason for the verification failure is recorded; based on the reason for the verification failure, supplementary verification is performed; if the supplementary verification passes, the cabinet door is opened and the goods are picked up; if the supplementary verification fails, the pickup operation is exited. The present invention effectively solves the inherent defects of single biometric technology in terms of security and environmental adaptability through multimodal biometric fusion and dynamic verification strategies. The dual liveness detection mechanism of voice and face proposed in the present invention can actively defend against counterfeit attacks. Combined with the adaptive verification mode switching based on environmental perception, it ensures the safety of high-risk scenarios such as medical supplies while providing barrier-free and quick operation for ordinary scenarios. The localized encryption calculation and cross-modal fault-tolerant design in the present invention take into account both the privacy protection of biometric data and the robustness of operations in complex environments. The response time is shortened to within 2 seconds, which significantly expands the application reliability of smart retail cabinets in multiple scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments or the description of the prior art. Obviously, the drawings described below are merely exemplary, and those skilled in the art can, without inventive effort, derive other implementation drawings based on the provided drawings.

[0046] The structures, proportions, sizes, etc. illustrated in this specification are intended solely to complement the contents disclosed herein and to facilitate understanding and reading by persons skilled in the art. They are not intended to limit the conditions under which the present invention may be implemented and therefore have no substantive technical significance. Any structural modifications, changes in proportions, or adjustments in sizes, without affecting the efficacy and objectives of the present invention, shall remain within the scope of the technical contents disclosed herein.

[0047] Figure 1This is a flow chart of a method for picking up items from a retail counter that supports voice recognition and face recognition, provided in Example 1 of the present invention;

[0048] Figure 2 This is a schematic diagram of the architecture of a retail cabinet pickup device supporting voice recognition and face recognition provided in Example 2 of the present invention. DETAILED DESCRIPTION

[0049] The following describes the implementation of the present invention using specific embodiments. Those skilled in the art will readily understand the other advantages and benefits of the present invention from the disclosure herein. Obviously, the embodiments described are only a portion of the present invention, not all of it. All other embodiments derived by persons of ordinary skill in the art based on the embodiments of the present invention without inventive effort are intended to fall within the scope of protection of the present invention.

[0050] Example 1

[0051] See also Figure 1 Embodiment 1 of the present invention provides a method for picking up items from a retail counter that supports voice recognition and face recognition, comprising the following steps:

[0052] S1. When the infrared sensor detects that the user enters the set interaction range, it activates the microphone array of the voice recognition module and the binocular camera of the face recognition module, and initializes the multimodal data acquisition protocol;

[0053] S2. Based on the multimodal data acquisition protocol, the voice recognition module captures the user's voice commands through directional beamforming technology; the face recognition module acquires the user's facial image through 3D structured light projection technology;

[0054] S3. Based on the user's voice command, perform voice recognition verification to obtain a voiceprint matching score and a voice liveness detection result; based on the user's facial image, perform face recognition verification to obtain a face matching score and a face liveness detection result;

[0055] S4. Output the recognition verification conclusion based on the retail counter verification mode, combined with the voice recognition verification results and the face recognition verification results; if the recognition verification passes, the counter door is opened and the goods are picked up; if the recognition verification fails, proceed to the next step;

[0056] S5. When the identification verification fails, the modal data that has passed the verification is frozen and the reason for the verification failure is recorded; based on the reason for the verification failure, a supplementary verification is performed; if the supplementary verification passes, the cabinet door is opened and the pickup is completed; if the supplementary verification fails, the pickup operation is exited.

[0057] In this embodiment, in step S1, when the infrared sensor detects that the user enters the set interaction range, the microphone array of the voice recognition module and the binocular camera of the face recognition module are activated, and the multimodal data acquisition protocol is initialized;

[0058] Specifically, when a user approaches a retail counter, the infrared sensor monitors the human body's thermal signature in real time and determines whether the user has entered the effective interaction area based on a preset distance threshold (e.g., 0.5-1.5 meters). At this point, the main control system synchronously activates the high-sensitivity microphone array of the voice recognition module and the binocular camera of the face recognition module, and starts the initialization process of the multimodal data acquisition protocol. This protocol synchronizes the clocks of the acoustic and optical sensors through the hardware driver layer to ensure that the timestamps of voice signal acquisition and image frame capture are aligned, while loading pre-stored noise models and illumination compensation parameters, laying the foundation for subsequent collaborative processing of multimodal data.

[0059] In this embodiment, in step S2, based on the multimodal data acquisition protocol, the voice recognition module captures the user's voice command through directional beamforming technology; the face recognition module acquires the user's facial image through 3D structured light projection technology;

[0060] Specifically, under the control of the data acquisition protocol, the voice recognition module uses directional beamforming technology and a spatial filtering algorithm based on a six-microphone ring array to suppress ambient noise and focus the user's voice beam, capturing voice commands containing voiceprint features (such as "pickup code 2587") in real time. Simultaneously, the facial recognition module activates a 3D structured light projection device, projecting tens of thousands of invisible infrared light spots onto the user's face. The binocular camera collects 3D facial point cloud data with depth information, and fuses it with visible light images to generate anti-glare facial texture information.

[0061] The two sensor systems use a timestamp synchronization mechanism to ensure the consistency of voice commands and facial images in time and space, avoiding multimodal data misalignment caused by acquisition delays.

[0062] In this embodiment, in step S3, based on the user's voice command, voice recognition verification is performed to obtain a voiceprint matching score and a voice liveness detection result; based on the user's facial image, face recognition verification is performed to obtain a face matching score and a face liveness detection result;

[0063] Specifically, the voice verification branch first performs noise reduction on the original voice signal, extracts Mel-frequency cepstral coefficients (MFCC) as voiceprint features, and performs similarity matching with pre-stored voiceprint templates through the dynamic time warping (DTW) algorithm to generate a voiceprint matching score in the range of 0-1.

[0064] At the same time, the voice liveness detection module analyzes the voice signal's biometrics, such as breathing intervals and vocal cord vibration frequency, to determine if the voice is real. The face verification branch uses 3D point cloud data to calculate facial curvature, pupil reflection characteristics of infrared light, and micro-expression changes to perform liveness detection. If liveness verification is passed, an improved ArcFace algorithm is used to extract features of the eye area (such as eyelid contour and iris texture) obscured by the mask, and cosine similarity is used to calculate a face match score with a pre-stored template.

[0065] Among them, the cosine similarity calculation formula is:

[0066]

[0067] Where A and B are the real-time extracted facial feature vector (such as 128-dimensional eye area feature) and the pre-stored facial template feature vector respectively; A i and B i is the value of the i-th dimension in the feature vector; n is the number of dimensions.

[0068] The similarity range is compressed to [-1, 1], and the system uses a linear transformation to map it to the interval [0, 1], where 0 indicates a complete mismatch and 1 indicates a perfect match. When the score exceeds a preset threshold (such as 0.95 in high-security mode), the face verification is considered successful.

[0069] In this embodiment, in step S4, according to the retail counter verification mode, combined with the voice recognition verification results and the face recognition verification results, the recognition verification conclusion is output; if the recognition verification passes, the counter door is opened and the goods are picked up; if the recognition verification fails, the next step is processed;

[0070] Specifically, the system dynamically decides based on the preset verification mode:

[0071] In high-security mode, the voiceprint matching score must be ≥0.9, the face matching score must be ≥0.95, and both liveness detections must pass before the authentication is considered successful and the electronic lock can be driven to open the target cabinet door.

[0072] In convenient mode, if the voiceprint score is ≥0.85 or the face score is ≥0.9, and the other modality does not trigger a liveness alarm (such as detecting a recording or photo attack), a weighted fusion algorithm (such as 0.6×face score + 0.4×voiceprint score) is used for comprehensive judgment. If the verification is successful, the cabinet door opens immediately and a successful pickup prompt sound is played; if the verification fails, the exception handling process is entered, the failure type (such as insufficient voiceprint matching, facial occlusion, etc.) is recorded, and the fault tolerance mechanism is triggered.

[0073] In this embodiment, in step S5, when the identification verification fails, the modal data that has passed the verification is frozen and the reason for the verification failure is recorded; based on the reason for the verification failure, a supplementary verification is performed; if the supplementary verification passes, the cabinet door is opened and the goods are picked up; if the supplementary verification fails, the pickup operation is exited.

[0074] Specifically, when verification fails, the system freezes the verified modal data (e.g., a voiceprint score of 0.88 but face detection fails) and initiates targeted supplementary verification based on the cause of the failure:

[0075] If speech recognition fails due to ambient noise, the microphone array gain is increased and the user is prompted to repeat the command;

[0076] If face recognition fails due to facial occlusion, the infrared fill light is enabled and the user is guided to adjust the angle and re-capture the image.

[0077] During the supplementary verification process, the system combines the initial verification results with supplementary data (such as a voice secondary match score of 0.92) and executes the locker access according to the downgraded security policy (such as a single modality high score). If the supplementary verification still fails, the process is terminated and the user is prompted to contact the administrator.

[0078] In this embodiment, regardless of whether the verification is passed or not, the collected voiceprint spectrum, facial feature points and verification log will be encrypted in pieces and stored in the local trusted execution environment (TEE); after the cabinet opening operation is completed, the temporarily generated intermediate feature vector will be destroyed to ensure that the private data cannot be restored.

[0079] Among them, the encryption process can be encrypted by AES-256.

[0080] In a possible embodiment, an example of picking up items from a community smart drug retail cabinet is provided as follows:

[0081] The hardware configuration of the smart medicine retail cabinet is:

[0082] Main control system: Rockchip RK3588 chip, equipped with dual-core NPU (computing power 6TOPS), running embedded Linux system;

[0083] Voice module: Circular 6-microphone array (XMOS XVF3610 chip), arranged at a 30° elevation angle on the top of the cabinet, supporting 120° directional pickup;

[0084] Face recognition module: Orbbec U3Pro binocular 3D camera with integrated VCSEL infrared structured light projector (resolution 1280×960, frame rate 30fps);

[0085] Secure storage: Local encryption chip uses Infineon OPTIGA TMTPM 2.0, supports SM4 national encryption algorithm.

[0086] Specific implementation process:

[0087] User registration:

[0088] Community patients use the hospital app to input facial data (3D point cloud + visible light texture) and voiceprint samples (reading the dynamic numeric string "2587"). The data is encrypted and stored locally in the TEE. MFCC coefficients are extracted from the voiceprint features, and a 128-dimensional voiceprint vector is generated using the GMM-UBM model. Facial features are extracted using an improved ArcFace algorithm, focusing on extracting 64 feature points around the eyes (such as the distance between the inner canthi and iris texture) when the face is obscured by a mask.

[0089] User pickup process:

[0090] T1: When the user is 1.2 meters away from the cabinet, the infrared sensor (Senba Optoelectronics HX-S5) is triggered and activated, the main control chip loads the multimodal protocol, the microphone and camera are started synchronously, and the noise model is initialized (the noise reduction level is automatically configured based on the current ambient noise of 60dB);

[0091] T2. Data Collection: Voice: Uses beamforming to lock the user's position, suppresses side ambient noise (such as air conditioning), and captures the voice command "pickup code 2587" at a sampling rate of 16kHz. Face: 3D structured light projects 28,000 infrared points, calculates a depth map (accuracy ±1mm), simultaneously collects visible light images, and fuses them to generate anti-backlight facial data.

[0092] T3, Voiceprint Verification: Extract voice MFCC features and compare them with the pre-stored GMM-UBM model, outputting a match of 0.92 (threshold 0.85), while also detecting continuous breathing intervals (liveness detection passed); Face Verification: 3D liveness detection determines non-photo attacks (facial curvature change rate > 0.15), with a periocular feature comparison score of 0.94 (threshold 0.90);

[0093] T4. Because the medicine cabinet is forced to use mode 1 (double verification), the system determines that both the voiceprint (0.92) and the face (0.94) exceed the threshold, and drives the electromagnetic lock to open the corresponding medicine compartment (number A07). The log is recorded:

[0094] "2024-05-2014:30:23 User ID_1032 dual-modal verification successful";

[0095] Fault-tolerant scenarios:

[0096] If the user wears a mask resulting in a face score of 0.88 (below the threshold of 0.90):

[0097] T5. Supplementary verification:

[0098] Freeze the passed voiceprint data (0.92) and record the failure reason as "face occlusion";

[0099] The system prompts "Please adjust the angle to keep both eyes visible to the camera" and activates the infrared fill light to enhance the details around the eyes;

[0100] After collecting facial data for the second time, the score of eye contour features increased to 0.91, triggering the downgrade strategy: combining the high voiceprint score (0.92) and the supplementary face score (0.91), the weighted total score reached 0.914 (formula: 0.6×face score + 0.4×voiceprint score), exceeding the fault tolerance threshold of 0.90, and the cabinet was finally opened.

[0101] In this embodiment, all biometric data are stored in SM4 encrypted fragments. The temporarily generated voiceprint spectrogram (.wav) and facial point cloud data (.ply) are destroyed within 2 seconds after the cabinet is opened, and only the encrypted hash value is retained for auditing.

[0102] To summarize, when the infrared sensor detects that a user enters a set interaction range, the present invention activates the microphone array of the voice recognition module and the binocular camera of the face recognition module, and initializes the multimodal data acquisition protocol; based on the multimodal data acquisition protocol, the voice recognition module captures the user's voice command through directional beamforming technology; the face recognition module obtains the user's facial image through 3D structured light projection technology; based on the user's voice command, voice recognition verification is performed to obtain a voiceprint matching score and a voice liveness detection result; based on the user's facial image, face recognition verification is performed to obtain a face matching score and a face liveness detection result; according to the retail cabinet verification mode, the voice recognition verification result and the face recognition verification result are combined to output the recognition verification conclusion; if the recognition verification passes, the cabinet door is opened and the goods are picked up; if the recognition verification fails, the next step is processed; when the recognition verification fails, the verified modal data is frozen and the reason for the verification failure is recorded; based on the reason for the verification failure, supplementary verification is performed; if the supplementary verification passes, the cabinet door is opened and the goods are picked up; if the supplementary verification fails, the pickup operation is exited. The present invention effectively solves the inherent defects of single biometric technology in terms of security and environmental adaptability through multimodal biometric fusion and dynamic verification strategies. The dual liveness detection mechanism of voice and face proposed in the present invention can actively defend against counterfeit attacks. Combined with the adaptive verification mode switching based on environmental perception, it ensures the safety of high-risk scenarios such as medical supplies while providing barrier-free and quick operation for ordinary scenarios. The localized encryption calculation and cross-modal fault-tolerant design in the present invention take into account both the privacy protection of biometric data and the robustness of operations in complex environments. The response time is shortened to within 2 seconds, which significantly expands the application reliability of smart retail cabinets in multiple scenarios.

[0103] It should be noted that the method of the embodiments of the present disclosure can be performed by a single device, such as a computer or server. The method of the embodiments of the present disclosure can also be applied in a distributed scenario, where multiple devices cooperate to perform the method. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiments of the present disclosure, and the multiple devices will interact with each other to complete the method.

[0104] It should be noted that the above description is limited to some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0105] Example 2

[0106] See also Figure 2 Embodiment 2 of the present invention further provides a retail cabinet pickup device supporting voice recognition and face recognition, comprising:

[0107] The recognition module activation unit 001 is used to activate the microphone array of the voice recognition module and the binocular camera of the face recognition module and initialize the multimodal data acquisition protocol when the infrared sensor detects that the user enters the set interaction range;

[0108] The user voice command and facial image acquisition unit 002 is used to capture the user's voice command through the directional beamforming technology of the voice recognition module based on the multimodal data acquisition protocol; and the face recognition module obtains the user's facial image through the 3D structured light projection technology;

[0109] The voice recognition verification and face recognition verification unit 003 is used to perform voice recognition verification based on the user's voice command to obtain a voiceprint matching score and a voice liveness detection result; and perform face recognition verification based on the user's facial image to obtain a face matching score and a face liveness detection result;

[0110] The recognition verification conclusion acquisition and processing unit 004 is used to output the recognition verification conclusion based on the retail counter verification mode, combined with the voice recognition verification results and the face recognition verification results; if the recognition verification passes, the cabinet door is opened and the goods are collected; if the recognition verification fails, the next step is carried out;

[0111] The supplementary verification unit 005 is used to freeze the verified modal data and record the reason for the verification failure when the identification verification fails; perform supplementary verification based on the reason for the verification failure; if the supplementary verification passes, open the cabinet door and complete the pickup; if the supplementary verification fails, exit the pickup operation.

[0112] In this embodiment, in the user voice command and facial image acquisition unit 002, in the process of acquiring the user voice command and the user facial image, synchronous acquisition is performed through timestamps to keep the time and space of data acquisition consistent.

[0113] In this embodiment, in the voice recognition verification and face recognition verification unit 003, during the voice recognition verification process:

[0114] Performing noise reduction processing on the user's voice command and extracting Mel-frequency cepstral coefficients as voiceprint features; performing dynamic time warping matching on the voiceprint features and the set encrypted voiceprint template to generate a voiceprint matching score and voice liveness detection results;

[0115] During the face recognition verification process:

[0116] Based on the user's facial image, facial micro-expression changes and pupil reflection features are calculated to obtain the face liveness detection result; after the liveness detection is passed, the sampling occlusion adaptive algorithm is used to extract the features of the area around the eyes, and the face matching score is obtained through cosine similarity calculation.

[0117] In this embodiment, in the identification verification conclusion acquisition and processing unit 004, in the process of outputting the identification verification conclusion according to the retail counter verification mode, combining the voice recognition verification result and the face recognition verification result:

[0118] If the retail counter verification mode is high security mode, the voiceprint matching score must be ≥ 0.9 and the face matching score must be ≥ 0.95, and both voice liveness detection and face liveness detection must be passed;

[0119] If the retail counter verification mode is the convenient mode, the voiceprint matching score is required to be ≥ 0.85 or the face matching score is required to be ≥ 0.9, and the liveness detection of the other modality must be passed.

[0120] In this embodiment, in the process of performing supplementary verification based on the verification failure reason, the supplementary verification unit 005:

[0121] If the voice recognition verification fails, the microphone array gain is adjusted and the user is prompted to repeat the voice command;

[0122] If facial recognition verification fails, infrared fill light is activated and the user is guided to adjust the facial angle;

[0123] After the supplementary verification is passed, the initial verification and supplementary verification results are integrated and the cabinet opening operation is performed according to the downgraded security policy.

[0124] In this embodiment, during the process of the user performing voice recognition verification and face recognition verification, regardless of whether the verification is passed or not, the collected voiceprint spectrum, facial feature points and verification log will be fragmented and encrypted and stored in the local trusted execution environment; after the cabinet opening operation is completed, the generated biometric intermediate data will be destroyed.

[0125] It should be noted that the information interaction, execution process, etc. between the modules of the above-mentioned system are based on the same concept as the method embodiment in Example 1 of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown above in the present application, and no further details will be given here.

[0126] Example 3

[0127] Embodiment 3 of the present invention provides a non-transitory computer-readable storage medium, in which a program code for a retail cabinet pickup method that supports voice recognition and face recognition is stored. The program code includes instructions for executing embodiment 1 or any possible implementation thereof, a retail cabinet pickup method that supports voice recognition and face recognition.

[0128] Computer-readable storage media can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state drives (SSDs)).

[0129] Example 4

[0130] Embodiment 4 of the present invention provides an electronic device, including: a memory and a processor;

[0131] The processor and the memory communicate with each other through a bus; the memory stores program instructions that can be executed by the processor, and the processor calls the program instructions to execute a retail cabinet pickup method that supports voice recognition and face recognition in embodiment 1 or any possible implementation thereof.

[0132] Specifically, the processor can be implemented by hardware or by software. When implemented by hardware, the processor can be a logic circuit, an integrated circuit, etc.; when implemented by software, the processor can be a general-purpose processor, which is implemented by reading software code stored in a memory. The memory can be integrated into the processor or located outside the processor and exist independently.

[0133] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present invention is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable systems. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode.

[0134] Obviously, those skilled in the art will appreciate that the various modules or steps of the present invention described above can be implemented using a general-purpose computing system. They can be centralized on a single computing system or distributed across a network of multiple computing systems. Alternatively, they can be implemented using program code executable by a computing system, and thus, they can be stored in a storage system and executed by the computing system. In some cases, the steps shown or described herein can be performed in a different order than that shown, or they can be fabricated into separate integrated circuit modules, or multiple modules or steps can be fabricated into a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.

[0135] Although the present invention has been described in detail above using general descriptions and specific embodiments, it will be apparent to those skilled in the art that modifications and improvements may be made thereto. Therefore, such modifications and improvements, without departing from the spirit of the present invention, are intended to be within the scope of protection claimed herein.

Claims

1. A retail counter pickup method supporting voice recognition and face recognition, characterized in that: include: When the infrared sensor detects that the user enters the set interaction range, it activates the microphone array of the voice recognition module and the binocular camera of the face recognition module, and initializes the multimodal data acquisition protocol; Based on the multimodal data acquisition protocol, the voice recognition module captures the user's voice commands through directional beamforming technology; the face recognition module acquires the user's facial image through 3D structured light projection technology; Based on the user's voice command, voice recognition verification is performed to obtain a voiceprint matching score and a voice liveness detection result; based on the user's facial image, face recognition verification is performed to obtain a face matching score and a face liveness detection result; According to the retail counter verification mode, combined with the voice recognition verification results and the face recognition verification results, the recognition verification conclusion is output; if the recognition verification passes, the counter door is opened and the goods are picked up; if the recognition verification fails, the next step is processed; When the identification verification fails, the modal data that has passed the verification is frozen and the reason for the verification failure is recorded; based on the reason for the verification failure, a supplementary verification is performed; if the supplementary verification passes, the cabinet door is opened and the pickup is completed; if the supplementary verification fails, the pickup operation is exited.

2. A retail counter pickup method supporting voice recognition and face recognition according to claim 1, characterized in that: In the process of acquiring the user voice command and the user facial image, synchronous acquisition is performed through timestamps to keep the time and space of data acquisition consistent.

3. The method for picking up items from a retail counter supporting voice recognition and face recognition according to claim 2, characterized in that: During the voice recognition verification process: Performing noise reduction processing on the user's voice command and extracting Mel-frequency cepstral coefficients as voiceprint features; performing dynamic time warping matching on the voiceprint features and the set encrypted voiceprint template to generate a voiceprint matching score and voice liveness detection results; During the face recognition verification process: Based on the user's facial image, facial micro-expression changes and pupil reflection features are calculated to obtain the face liveness detection result; after the liveness detection is passed, the sampling occlusion adaptive algorithm is used to extract the features of the area around the eyes, and the face matching score is obtained through cosine similarity calculation.

4. The method for picking up items from a retail counter supporting voice recognition and face recognition according to claim 3, wherein: In the process of outputting the recognition and verification conclusion according to the retail counter verification mode, combining the voice recognition and verification results with the face recognition and verification results: If the retail counter verification mode is high security mode, the voiceprint matching score must be ≥ 0.9 and the face matching score must be ≥ 0.95, and both voice liveness detection and face liveness detection must be passed; If the retail counter verification mode is the convenient mode, the voiceprint matching score is required to be ≥ 0.85 or the face matching score is required to be ≥ 0.9, and the liveness detection of the other modality must be passed.

5. The method for picking up items from a retail counter supporting voice recognition and face recognition according to claim 4, characterized in that: During the supplementary verification process based on the reasons for the verification failure: If the voice recognition verification fails, the microphone array gain is adjusted and the user is prompted to repeat the voice command; If facial recognition verification fails, infrared fill light is activated and the user is guided to adjust the facial angle; After the supplementary verification is passed, the initial verification and supplementary verification results are integrated and the cabinet opening operation is performed according to the downgraded security policy.

6. The method for picking up items from a retail counter supporting voice recognition and face recognition according to claim 5, characterized in that: During the process of voice recognition and face recognition verification, regardless of whether the user passes the verification or not, the collected voiceprint spectrum, facial feature points and verification log will be encrypted in pieces and stored in the local trusted execution environment; after the cabinet opening operation is completed, the generated biometric intermediate data will be destroyed.

7. A retail cabinet pickup device that supports voice recognition and face recognition, characterized in that: include: The recognition module activation unit is used to activate the microphone array of the voice recognition module and the binocular camera of the face recognition module when the infrared sensor detects that the user enters the set interaction range, and initialize the multimodal data acquisition protocol; A user voice command and facial image acquisition unit is configured to capture user voice commands using directional beamforming technology based on the multimodal data acquisition protocol; and the facial recognition module acquires user facial images using 3D structured light projection technology; A voice recognition verification and face recognition verification unit, configured to perform voice recognition verification based on the user's voice command to obtain a voiceprint matching score and a voice liveness detection result; and perform face recognition verification based on the user's facial image to obtain a face matching score and a face liveness detection result; The recognition and verification conclusion acquisition and processing unit is used to output the recognition and verification conclusion based on the retail counter verification mode, combined with the voice recognition verification results and the face recognition verification results; if the recognition and verification pass, the cabinet door is opened and the goods are collected; if the recognition and verification fails, the next step of processing is carried out; The supplementary verification unit is used to freeze the verified modal data and record the reason for the verification failure when the identification verification fails; perform supplementary verification based on the reason for the verification failure; if the supplementary verification passes, open the cabinet door and complete the pickup; if the supplementary verification fails, exit the pickup operation.

8. The retail cabinet pickup device supporting voice recognition and face recognition according to claim 7, characterized in that: In the user voice command and facial image acquisition unit, in the process of acquiring the user voice command and the user facial image, synchronous acquisition is performed through timestamps, so that the time and space of data acquisition remain consistent.

9. The retail cabinet pickup device supporting voice recognition and face recognition according to claim 8, characterized in that: In the voice recognition verification and face recognition verification unit, during the voice recognition verification process: Performing noise reduction processing on the user voice command and extracting Mel-frequency cepstral coefficients as voiceprint features; Perform dynamic time-warping matching on the voiceprint feature and the set encrypted voiceprint template to generate a voiceprint matching score and a voice liveness detection result; During the face recognition verification process: Based on the user's facial image, facial micro-expression changes and pupil reflection features are calculated to obtain the face liveness detection result; after the liveness detection is passed, the sampling occlusion adaptive algorithm is used to extract the features of the area around the eyes, and the face matching score is obtained through cosine similarity calculation.

10. The retail cabinet pickup device supporting voice recognition and face recognition according to claim 9, characterized in that: In the identification verification conclusion acquisition and processing unit, in the process of outputting the identification verification conclusion according to the retail counter verification mode, combining the voice recognition verification result and the face recognition verification result: If the retail counter verification mode is high security mode, the voiceprint matching score must be ≥ 0.9 and the face matching score must be ≥ 0.95, and both voice liveness detection and face liveness detection must be passed; If the retail counter verification mode is the convenient mode, the voiceprint matching score is required to be ≥ 0.85 or the face matching score is required to be ≥ 0.9, and the liveness detection of the other modality must be passed.

Citation Information

Patent Citations

  • Voiceprint identification, face identification and synchronous in-vivo detection-based identity authentication method and system

    CN105426723A

  • Interactive authentication pickup system based on face recognition and voice recognition

    CN112466057A

  • Intelligent voice remote controller

    CN119207386A

Cited By

  • Data processing method and device, medium, program product and robot system

    CN120998204A