Face recognition method and device based on mixed modal features and electronic device

By combining multi-camera, full-color, and infrared cameras to achieve hybrid modal feature fusion, the accuracy and success rate issues of facial recognition technology in complex environments were resolved, resulting in highly efficient facial recognition.

CN122116449APending Publication Date: 2026-05-29BEIJING TAI DE HUI ZHI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING TAI DE HUI ZHI TECH CO LTD
Filing Date
2026-04-30
Publication Date
2026-05-29

AI Technical Summary

Technical Problem

Facial recognition technology suffers from low success and accuracy in security and physical access control scenarios due to limitations imposed by shooting angles and complex lighting conditions.

Method used

By employing a combination of multi-camera systems, full-color cameras, and infrared cameras, motion tendency information is generated through scrambling recognition. High-resolution full-color and infrared images are acquired simultaneously, and mixed-modal features are extracted and fused, including color, infrared, and depth facial features, thereby improving recognition accuracy.

Benefits of technology

It improves the success rate and accuracy of facial recognition, reduces computing resources and power consumption, enhances the robustness of image acquisition, overcomes the influence of lighting environment, and does not increase hardware costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122116449A_ABST
    Figure CN122116449A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure disclose a face recognition method and device based on mixed modal features and an electronic device. A specific implementation of the method comprises: performing a moving detection on a first image sequence to generate moving tendency information; in response to the moving tendency information representing that a moving direction of a recognized object is towards an entrance side, synchronously collecting a second image sequence and a third image sequence by using a multi-view camera; performing mixed modal feature extraction on the second image sequence and the third image sequence to obtain a mixed modal feature sequence; performing feature fusion on the mixed modal feature sequence to obtain a fused face feature; and performing face recognition on the fused face feature to obtain a face recognition result. The implementation improves the success rate and accuracy of face recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein relate to the field of computer technology, and more particularly to the field of image processing and recognition, specifically to facial recognition methods, apparatus, and electronic devices based on mixed-modal features. Background Technology

[0002] Facial recognition technology, as an important branch of biometric identification, is widely used, especially in high-security applications such as security and physical access control. However, it is limited by factors such as shooting angle and complex lighting environment, which seriously restricts the success rate and accuracy of facial recognition. Summary of the Invention

[0003] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.

[0004] Some embodiments of this disclosure propose facial recognition methods, apparatuses, and electronic devices based on mixed-modal features to address the technical problems mentioned in the background section above.

[0005] In a first aspect, some embodiments of this disclosure provide a facial recognition method based on hybrid modal features. The method includes: performing intrusion recognition based on a first image sequence to generate movement tendency information, wherein the first image is a low-resolution full-color image captured by a full-color camera included in a multi-view camera system, the multi-view camera being mounted on a one-way gate device and facing the entry side; responding to the movement tendency information indicating that the movement direction of the identified object is towards the entry side, simultaneously acquiring a second image sequence and a third image sequence through the multi-view camera system, wherein the second image is a high-resolution full-color image captured by the full-color camera system, and the third image is an infrared image captured by an infrared camera included in the multi-view camera system; extracting hybrid modal features based on the second image sequence and the third image sequence to obtain a hybrid modal feature sequence, wherein the hybrid modal features consist of color facial features, infrared facial features, and depth-of-field facial features; fusing features in the hybrid modal feature sequence to obtain fused facial features; and performing facial recognition based on the fused facial features to obtain a facial recognition result.

[0006] Secondly, some embodiments of this disclosure provide a facial recognition device based on mixed modal features. The device includes: an intrusion recognition unit configured to perform intrusion recognition based on a first image sequence to generate movement tendency information, wherein the first image is a low-resolution full-color image captured by a full-color camera included in a multi-view camera, the multi-view camera being disposed on a one-way gate device and facing the entry side; and a synchronous acquisition unit configured to, in response to the movement tendency information indicating that the movement direction of the identified object is towards the entry side, synchronously acquire a second image sequence and a third image sequence through the multi-view camera, wherein the second .... The high-resolution full-color image captured by the full-color camera, and the third image is an infrared image captured by the infrared camera included in the multi-view camera; the mixed-modality feature extraction unit is configured to extract mixed-modality features based on the second image sequence and the third image sequence to obtain a mixed-modality feature sequence, wherein the mixed-modality features consist of color facial features, infrared facial features and depth facial features; the feature fusion unit is configured to fuse the mixed-modality feature sequence to obtain fused facial features; the face recognition unit is configured to perform face recognition based on the fused facial features to obtain a face recognition result.

[0007] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.

[0008] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.

[0009] The above embodiments of this disclosure have the following beneficial effects: the facial recognition method based on mixed modal features of some embodiments of this disclosure improves the success rate and accuracy of facial recognition. Specifically, this disclosure first performs intrusion recognition based on a first image sequence to generate movement tendency information, wherein the first image is a low-resolution full-color image captured by a full-color camera, which is mounted on a one-way gate device and faces the entry side. In practice, gate devices, as one of the main physical access control facilities, usually use full-color cameras for image acquisition and facial recognition. When a moving object passes in front of the gate device, it will frequently trigger the camera's image acquisition and facial recognition, thereby causing unnecessary waste of computing resources. Therefore, this disclosure reduces computing resource consumption by combining low-resolution full-color images for movement tendency recognition. Secondly, in response to the movement tendency information indicating that the movement direction of the identified object is towards the entry side, a second image sequence and a third image sequence are simultaneously acquired by the multi-view camera, wherein the second image is a high-resolution full-color image captured by the full-color camera, and the third image is an infrared image captured by an infrared camera, which is included in the multi-view camera. When the identified object moves toward the gate device, the infrared camera and full-color camera are triggered to acquire high-resolution images, thereby reducing device power consumption and computing power consumption. Furthermore, complex lighting environments (such as backlighting, low light, and overexposure) can severely affect image quality; therefore, this disclosure improves the robustness of image acquisition by combining infrared and full-color cameras. Next, mixed-modal feature extraction is performed based on the second and third image sequences to obtain a mixed-modal feature sequence, wherein the mixed-modal features consist of color facial features, infrared facial features, and depth-of-field facial features. By extracting color and infrared facial features, the features are mutually complementary, thereby improving the robustness of facial feature representation. In addition, depth-of-field facial features are typically obtained using dual full-color cameras for depth recovery, but this cannot overcome the influence of the aforementioned lighting environment. Therefore, this disclosure achieves depth recovery by combining infrared and full-color images, thereby achieving effective depth recovery without increasing hardware costs. Following this, feature fusion is performed on the aforementioned mixed-modal feature sequence to obtain fused facial features. Finally, facial recognition is performed based on the fused facial features to obtain the facial recognition result. In practice, due to the influence of environmental factors and the movement state of the object being recognized, the accuracy of facial recognition based on a single image is difficult to guarantee. Therefore, this disclosure improves the expressive power of features by combining color facial features, infrared facial features, mental facial features, and temporal features through feature fusion, thereby improving the success rate and accuracy of facial recognition. Attached Figure Description

[0010] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.

[0011] Figure 1 This is a flowchart of some embodiments of the face recognition method based on mixed modal features according to the present disclosure;

[0012] Figure 2 This is a schematic diagram of an intrusion recognition scenario;

[0013] Figure 3 This is a schematic diagram illustrating the process of generating facial features.

[0014] Figure 4 This is another schematic diagram illustrating the generation process that integrates facial features;

[0015] Figure 5 This is a schematic diagram of the structure of some embodiments of a face recognition device based on mixed modal features according to the present disclosure;

[0016] Figure 6 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation

[0017] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.

[0018] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.

[0019] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.

[0020] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0021] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.

[0022] Before performing any of the operations involving the images (such as the first image, the second image, and the third image) disclosed herein, the relevant organizations or individuals shall fulfill their obligations, including conducting information security impact assessments, informing the information subjects, and obtaining prior authorization and consent from the information subjects.

[0023] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.

[0024] refer to Figure 1 The diagram illustrates a flow 100 of some embodiments of a face recognition method based on mixed-modal features according to the present disclosure. This face recognition method based on mixed-modal features includes the following steps: Step 101: Perform scrambling identification based on the first image sequence to generate movement tendency information.

[0025] In some embodiments, the implementer of the face recognition method based on mixed modal features (e.g., a computing device) can perform scrambling identification based on a first image sequence to generate movement tendency information.

[0026] The first image is a low-resolution full-color image captured by a full-color camera included in the multi-view camera system. For example, the image resolution of the first image may be 320 × 240. The frame rate of the first image sequence may be 15 FPS. The multi-view camera is installed on a one-way turnstile and faces the entrance side. The one-way turnstile is a turnstile used to control one-way passage. The multi-view camera is a binocular camera that includes a full-color camera and an infrared camera. Motion tendency information characterizes the movement tendency of the identified object. Intrusion recognition refers to the recognition of a moving object that has entered a fixed recognition area and is moving towards the one-way turnstile. For example, the moving object may be a pedestrian.

[0027] In practice, a lightweight object detection model (such as the Tiny-YOLO model) can be used to perform continuous object detection on the first image in the first image sequence, and the SORT (Simple Online and Realtime Tracking) algorithm can be used for object tracking to determine the movement tendency of the identified object and obtain movement tendency information. In particular, since only the identified object needs to be identified, object detection can be performed using only key parts such as the head or shoulders. Furthermore, since the first image is a low-resolution image and only the movement tendency of the identified object needs to be captured, the SORT algorithm only requires the use of Kalman filtering and Hungarian matching algorithms for object tracking, eliminating the need for forced matching of appearance features. This reduces unnecessary data processing and improves the speed of ramming detection.

[0028] As an example, see Figure 2 The diagram shown illustrates an intrusion recognition scenario, in which... Figure 2 Two one-way turnstiles are shown: one-way turnstile A1 and one-way turnstile A2. One-way turnstile A1 is equipped with a multi-view camera B1. One-way turnstile A2 is equipped with a multi-view camera B2. Multi-view camera B1 faces the entry side of one-way turnstile A1. Multi-view camera B2 faces the entry side of one-way turnstile A2. First, taking one-way turnstile A1 as an example, low-resolution full-color images are continuously acquired using the full-color camera included in multi-view camera B1 to obtain a first image sequence. An object C1 is identified using a disturbance recognition method, and movement tendency information representing the object C1 moving towards the entry side of one-way turnstile A1 is generated. Next, taking one-way turnstile A2 as an example, low-resolution full-color images are continuously acquired using the full-color camera included in multi-view camera B2 to obtain a first image sequence. The system identifies the object C2 by means of intrusion recognition and generates movement tendency information indicating that the object C2 does not move toward the entry side corresponding to the one-way gate device A1.

[0029] It should be noted that the aforementioned computing device can be either hardware or software. When the computing device is hardware, it can be implemented as a distributed cluster consisting of multiple servers or terminal devices, or as a single server or a single terminal device. When the computing device is software, it can be installed within the hardware devices listed above. It can be implemented as, for example, multiple software programs or software modules used to provide distributed services, or as a single software program or software module. No specific limitations are made here. In particular, the aforementioned execution entity can be an edge computing device embedded within a one-way gate device.

[0030] In some optional implementations of certain embodiments, the execution entity performs scrambling identification based on the first image sequence to generate movement tendency information, including: Step S1: Extract frames from the first image sequence to obtain the target image sequence.

[0031] The target image is the first image obtained by frame extraction.

[0032] In practice, the frame extraction frequency of image framing is lower than the acquisition frequency of the first image sequence. Because object movement is continuous, image framing can further compress the number of images processed, improving the speed of disturbance recognition. For example, the frame extraction frequency can be 10 FPS.

[0033] Step S2: For each target image in the above target image sequence, perform the following first processing step: Step S21: Extract object blur features from the above target image to obtain object blur features.

[0034] Among them, the object fuzzy feature represents the image semantic features corresponding to the target image.

[0035] In practice, since only invasive behavior detection is required, it is not necessary to extract detailed object features (such as facial texture) while ensuring the effectiveness of invasive behavior detection. Therefore, when extracting fuzzy object features, this disclosure adopts a lightweight feature extraction model. Specifically, the feature extraction model consists of three parallel convolutional units and a feature stacking layer. The first convolutional unit consists of three sequential convolutional layers, each with a kernel size of 1×1, a stride of 1, and padding of 0. The second convolutional unit consists of three sequential convolutional layers, each with a kernel size of 3×3, a stride of 1, and padding of 1. The third convolutional unit consists of three sequential convolutional layers, each with a kernel size of 5×5, a stride of 1, and padding of 2. This approach utilizes three convolutional units to achieve lightweight feature extraction across different receptive fields. Simultaneously, by adjusting the stride and padding, the consistent output size of the three convolutional units is ensured. Assuming the target image size is H×W, the output size of the first convolutional unit can be H×W×C1, where C1 is the number of output channels. The output size of the second convolutional unit can be H×W×C2, where C2 is the number of output channels. The output size of the third convolutional unit can be H×W×C3, where C3 is the number of output channels. Therefore, the feature stacking layer can stack the outputs of the three convolutional units along the channel dimension to obtain the object blurring feature. The feature dimension of the object blurring feature can be H×W×(α×C1+β×C2+γ×C3), where α+β+γ=1, α is the stacking weight of the first convolutional unit's output, β is the stacking weight of the second convolutional unit's output, and γ is the stacking weight of the third convolutional unit's output. For example, α=0.3, β=0.5, γ=0.2.

[0036] Step S22: Determine the region of interest of the object based on the above fuzzy features of the object.

[0037] The region of interest is the image region that is suspected of containing the object to be identified.

[0038] In practice, the Region Proposal Network (RPN) can be used to locate targets based on fuzzy object features, thereby obtaining the region of interest.

[0039] Step S23: Perform one-dimensional feature mapping on the above-mentioned fuzzy features of the object to obtain the fuzzy feature matching vector of the object.

[0040] The object blur feature matching vector is a feature vector used to match and associate multiple target images in a target image sequence that contain the same identified object. Specifically, the vector dimension of the object blur feature matching vector can be 1×512.

[0041] In practice, a one-dimensional feature mapping of the fuzzy features of an object can be performed using three fully connected layers to obtain the object's fuzzy feature matching vector.

[0042] Step S3: Based on the object fuzzy feature matching vector corresponding to the target image, perform image association on the target images in the above target image sequence to obtain the associated image sequence.

[0043] Among them, the associated image sequence consists of multiple target images that correspond to the same identified object.

[0044] In practice, since each target image has been processed in the first processing step, when the target object contains the identified object, there will be a corresponding object fuzzy feature matching vector. Therefore, by calculating the vector similarity (e.g., cosine similarity) between the object fuzzy feature matching vectors corresponding to each two target images, it can be determined whether the identified objects contained in the two target images are the same, thereby realizing image association and obtaining the associated image sequence.

[0045] Step S4: Generate a spatiotemporal search region sequence based on the region of interest of the objects corresponding to the associated images in the above associated image sequence.

[0046] In this context, the spatiotemporal search regions in the aforementioned spatiotemporal search region sequence correspond one-to-one with the first images in the aforementioned first image sequence. The spatiotemporal search region is the region of interest within each first image that is suspected of containing the object to be identified.

[0047] In practice, since the associated image sequence consists of multiple target images corresponding to the same identified object, each associated image corresponds to a region of interest (ROI). The target image sequence is obtained by extracting frames from the first image sequence, and the movement of the identified object is continuous. Therefore, by using the two ROIs corresponding to two adjacent target images, sequential ROI interpolation can be performed on multiple first images between adjacent target images to obtain the spatiotemporal search region corresponding to each first image.

[0048] Step S5: For each first image in the first image sequence, local features are extracted from the first image within the spatiotemporal search region corresponding to the first image to obtain local image features.

[0049] Specifically, the lightweight feature extraction model in step S21 can be used to extract local features of the first image within the spatiotemporal search area corresponding to the first image, thereby further compressing the recognition area and reducing feature processing of irrelevant image areas.

[0050] Step S6: Perform optical flow tracing based on the obtained local image feature sequence to generate object motion features and object movement trajectory features.

[0051] Among them, object motion features can characterize the motion state of the identified object. Object movement trajectory features can characterize the movement trajectory of the identified object.

[0052] In practice, since local image features can characterize the image position of the object being identified in the first image, the movement trajectory of the object can be extracted using optical flow tracing, thus obtaining the object's movement trajectory features. Based on these features, the direction, speed, and acceleration of the object's movement trajectory between every two first images are determined, resulting in multiple directions, speeds, and accelerations, which serve as the object's motion features. This allows for the characterization of the object's movement behavior from both the trajectory and motion state dimensions.

[0053] Step S7: Generate movement tendency information based on the above object motion characteristics and the above object movement trajectory characteristics.

[0054] In practice, firstly, the object's motion features and the aforementioned object movement trajectory features can be concatenated (Concat concatenation) to construct a concatenated vector. Next, using this concatenated vector as input, a multi-classifier outputs movement tendency information. Specifically, the multi-classifier can be a binary classifier, and the output can be a classification result indicating whether the identified object's movement direction is towards the gate or not. For example, a classification result indicating that the identified object's movement direction is towards the gate can be represented by "1," and a classification result indicating that the identified object's movement direction is not towards the gate can be represented by "0."

[0055] Step 102: In response to the movement tendency information indicating that the movement direction of the identified object is toward the gate side, a second image sequence and a third image sequence are simultaneously acquired by a multi-view camera.

[0056] In some embodiments, the aforementioned execution entity may, in response to the movement tendency information indicating that the movement direction of the identified object is toward the gate side, simultaneously acquire a second image sequence and a third image sequence through a multi-view camera.

[0057] The second image is a high-resolution full-color image captured by the aforementioned full-color camera. The third image is an infrared image captured by the infrared camera included in the multi-view camera system. Both the infrared camera and the full-color camera included in the multi-view camera system have been pre-calibrated. The second and third images have the same image size.

[0058] In practice, since facial feature recognition is required subsequently, high-resolution images are necessary to ensure sufficient and accurate extraction of facial features. Therefore, this disclosure utilizes a multi-camera system, including an infrared camera and a full-color camera, to simultaneously acquire high-resolution full-color and infrared images, thereby obtaining a second image sequence and a third image sequence. In particular, the synchronous acquisition method ensures sequence alignment between the second and third image sequences, facilitating subsequent feature fusion. Specifically, the synchronous image acquisition of the multi-camera system, including the infrared camera and the full-color camera, can be controlled via an FSYNC (Frame Synchronization) signal.

[0059] Step 103: Extract mixed modal features based on the second image sequence and the third image sequence to obtain a mixed modal feature sequence.

[0060] In some embodiments, the aforementioned execution entity may perform mixed-modal feature extraction based on the second image sequence and the third image sequence to obtain a mixed-modal feature sequence.

[0061] The mixed-modal features consist of color facial features, infrared facial features, and depth-of-field facial features. Color facial features represent facial features extracted from the second image. Infrared facial features represent facial features extracted from the third image. Depth-of-field facial features represent depth-of-field features extracted from both the second and third images.

[0062] In practice, firstly, for color facial features and infrared facial features, a dual-path convolutional neural network can be used to extract facial features in parallel from the second image in the second image sequence and the third image in the third image sequence, obtaining the mixed modality features including color facial features and infrared facial features. The dual-path convolutional neural network consists of two parallel convolutional neural networks. The two parallel convolutional neural networks share parameters. Each of the two parallel convolutional neural networks is connected to a regressor for locating the facial region. By extracting the local features corresponding to the facial region, color facial features and infrared facial features are obtained. Secondly, since both the infrared camera and the full-color camera have been pre-calibrated, meaning their intrinsic parameters (focal length, principal point position, and distortion coefficients) and extrinsic parameters (rotation matrix and translation vector) are known, and since the infrared camera and the full-color camera are included within a multi-camera system, meaning their relative positions are fixed, and the sequences of the second and third image sequences are aligned. Therefore, by extracting feature points and matching features, the matching points of the second and third images that have a corresponding relationship in the facial region can be determined, and the pixel disparity between pairs of pixels can be calculated to recover the distance from the pixel to the camera, thereby obtaining the depth facial features.

[0063] In some optional implementations of certain embodiments, the execution entity performs mixed-modal feature extraction based on the second image sequence and the third image sequence to obtain a mixed-modal feature sequence, including: Step S1: For each second image in the second image sequence and the corresponding third image in the third image sequence, perform the following second processing step: Step S11: Perform distortion correction on the second image and the third image respectively to obtain the corrected second image and the corrected third image.

[0064] The second image after correction is the second image after distortion correction. The third image after correction is the third image after distortion correction.

[0065] In practice, since both infrared cameras and full-color cameras are pre-calibrated, distortion correction can be performed on the second and third images respectively based on the calibration parameters to obtain the corrected second and third images.

[0066] Step S12: Perform multi-scale image feature extraction on the corrected second image and the corrected third image respectively to obtain the first multi-scale image feature and the second multi-scale image feature.

[0067] The first multi-scale image feature is the multi-scale image feature representation corresponding to the corrected second image. The second multi-scale image feature is the multi-scale image feature representation corresponding to the corrected third image.

[0068] In practice, a dual-path feature pyramid network can be used to extract multi-scale image features from the corrected second and third images, yielding first and second multi-scale image features. The dual-path feature pyramid network comprises a first feature pyramid network and a second feature pyramid network. The first and second feature pyramid networks have identical structures and share parameters between corresponding layers. The first feature pyramid network is used for multi-scale image feature extraction from the corrected second image. The second feature pyramid network is used for multi-scale image feature extraction from the corrected third image. Both the first and second feature pyramid networks employ a "many-to-many" structure. Both the first and second feature pyramid networks include 5 downsampling layers. Multi-scale image feature extraction can be achieved through the feature pyramid network. Furthermore, by sharing parameters, not only can the training cost of the model be reduced, but the receptive field of view under full-color and infrared images can also be shared.

[0069] Step S13: Perform facial region search based on the first multi-scale image features to obtain the facial region of interest.

[0070] Among them, the facial region of interest represents the image region that is suspected to contain the face of the object to be identified.

[0071] In practice, the dual-path feature pyramid network is connected to the RPN network. The RPN network is used to search for facial regions of interest in the first multi-scale image features.

[0072] Step S14: Extract the local features of the first multi-scale image features and the second multi-scale image features in the facial region of interest, respectively, to obtain the first local features and the second local features, which are the color facial features and infrared facial features included in the mixed modality features.

[0073] The first local feature is a local feature located within the region of interest on the face, which is part of the first multi-scale image features. The second local feature is a local feature located within the region of interest on the face, which is part of the second multi-scale image features.

[0074] In practice, since the second and third images have the same image size, and the first and second multi-scale image features are located in the same feature space, local feature extraction can be performed on the first and second multi-scale image features by combining the facial region of interest with feature projection, resulting in first and second local features. This method retains only facial-related local features, thereby reducing the amount of subsequent feature processing.

[0075] Step S15: Generate a set of colored facial feature points based on the first local features described above.

[0076] Among them, colored facial feature points are facial feature points under the view of a colored image.

[0077] In practice, a binary classifier can be used to determine whether each feature point in the first local feature is a facial feature point. If the classification result is a facial feature point, the feature point is used as a colored facial feature point.

[0078] Step S16: Generate an infrared facial feature point set based on the second local features described above.

[0079] Among them, infrared facial feature points are facial feature points under the infrared image view.

[0080] In practice, a binary classifier can be used to determine whether each feature point in the second local feature is a facial feature point. If the classification result is a facial feature point, the feature point is used as an infrared facial feature point.

[0081] Step S17: Perform facial feature point matching on the above-mentioned set of colored facial feature points and the above-mentioned set of infrared facial feature points to obtain a set of facial feature point pairs.

[0082] Among them, there are two pairs of color facial feature points and infrared facial feature points that did not match successfully.

[0083] In practice, facial feature point matching can be performed on the above-mentioned set of colored facial feature points and the above-mentioned set of infrared facial feature points by similarity calculation to obtain a set of facial feature point pairs.

[0084] Step S18: Based on the facial region of interest, perform image segmentation on the second image and the third image to obtain a first facial region image and a second facial region image.

[0085] The first facial region image is a local image within the region of interest of the face in the second image. The second facial region image is a local image within the region of interest of the face in the third image.

[0086] Step S19: Based on the above set of facial feature points, perform image registration on the first facial region image and the second facial region image to obtain the registered region.

[0087] The registration region is the overlapping area of ​​the first facial region image and the second facial region image, with facial feature point pairs as alignment points.

[0088] In practice, since infrared images and color images represent features in different dimensions, image alignment and calibration are performed again by combining facial feature points to ensure the effectiveness of subsequent depth extraction.

[0089] Step S20: Generate a disparity feature map based on the above-mentioned registration region, the above-mentioned first facial region image, and the above-mentioned second facial region image.

[0090] The disparity feature map represents the pixel difference between corresponding pixels in the registration area of ​​the first facial region image and the second facial region image.

[0091] In practice, the pixel difference between pixels at the same position in the registration area of ​​the first facial region image and the aforementioned second facial region image can be calculated to obtain multiple pixel differences, which can be used as a disparity feature map.

[0092] Step S21: Perform depth transformation based on the above disparity feature map to obtain the depth facial features included in the mixed modality features.

[0093] Wherein, depth of field = (focal length × baseline distance) / parallax value.

[0094] In practice, since both the infrared and full-color cameras are pre-calibrated, their focal lengths and other camera parameters are known. Furthermore, their relative positions are fixed, meaning the optical center distance (baseline distance) between them is known. Therefore, multiple depth values ​​corresponding to the registration area can be calculated using the aforementioned formula, serving as depth-of-field facial features. Compared to the conventional method based on two full-color images, combining infrared and color images effectively overcomes the impact of complex lighting conditions on facial recognition. Moreover, this depth extraction method using both infrared and full-color images eliminates the need for an additional full-color camera, thus reducing hardware costs.

[0095] Step 104: Perform feature fusion on the mixed modality feature sequence to obtain fused facial features.

[0096] In some embodiments, the aforementioned execution entity may perform feature fusion on the mixed modality feature sequence to obtain fused facial features.

[0097] Among them, the fused facial features can be facial features obtained by fusing color facial features, infrared facial features, depth facial features, and temporal mixed modal feature sequences.

[0098] As an example, see Figure 3 The diagram illustrates the generation process of fused facial features. The mixed modal feature sequence 301 can include K mixed modal features. Taking the first mixed modal feature in the mixed modal feature sequence 301 as an example, its included color facial features, infrared facial features, and depth facial features are independently encoded in parallel by three encoders (encoder T1, encoder T2, and encoder T3) to expand the color facial features, infrared facial features, and depth facial features included in the mixed modal feature to the same feature dimension. Specifically, encoder T1 independently encodes the color facial features, obtaining the first encoding result 302. Encoder T2 independently encodes the infrared facial features, obtaining the second encoding result 303. Encoder T3 independently encodes the depth facial features, obtaining the third encoding result 304. Encoders T1, T2, and T3 are all Transformer encoder heads. Next, since the first encoding result 302, the second encoding result 303, and the third encoding result 304 have the same feature dimensions, they can be superimposed using feature superposition to obtain superimposed features 305. By processing each mixed modality feature in the mixed modality feature sequence 301, K superimposed features 305 can be obtained, and these K superimposed features 305 are superimposed again to obtain the fused facial features.

[0099] In some optional implementations of certain embodiments, the execution entity performs feature fusion on the aforementioned mixed-modal feature sequence to obtain fused facial features, including: Step S1: For each hybrid modality feature in the above hybrid modality feature sequence, perform intra-frame feature fusion on the color facial features, infrared facial features and depth facial features included in the above hybrid modality features to obtain the first fused facial feature.

[0100] The first fused facial feature is the fused feature obtained by fusing the color facial features, infrared facial features and depth facial features included in the mixed modality features within the frame.

[0101] In practice, the color facial features, infrared facial features, and depth facial features included in the hybrid modality features can be sampled to the same feature dimension and then superimposed as the first fused facial feature.

[0102] Step S2: Perform inter-frame feature fusion on the color facial features included in the above mixed modal feature sequence to obtain the second fused facial features.

[0103] The second fused facial feature is the fused feature obtained by fusing various colored facial features.

[0104] In practice, a recurrent neural network model with a "many to one" architecture can be used to perform inter-frame feature fusion on the color facial features included in the above mixed modality feature sequence to obtain the second fused facial features.

[0105] Step S3: Perform inter-frame feature fusion on the infrared facial features included in the above mixed modal feature sequence to obtain the third fused facial feature.

[0106] The third fused facial feature is the fused feature obtained by fusing various infrared facial features.

[0107] In practice, a recurrent neural network model with a "many to one" architecture can be used to perform inter-frame feature fusion on the infrared facial features included in the above mixed modal feature sequence to obtain the third fused facial feature.

[0108] Step S4: Perform inter-frame feature fusion on the depth-of-field facial features included in the above mixed modal feature sequence to obtain the fourth fused facial feature.

[0109] Among them, the fourth fused facial feature is the fused feature obtained by fusing facial features of each depth.

[0110] In practice, a recurrent neural network model with a "many to one" architecture can be used to perform inter-frame feature fusion on the infrared facial features included in the above mixed modality feature sequence to obtain the fourth fused facial feature.

[0111] Step S5: Dynamically fuse the obtained first fused facial feature set, the second fused facial feature, the third fused facial feature, and the fourth fused facial feature to obtain fused facial features.

[0112] Among them, the first fused facial feature set, the second fused facial feature set, the third fused facial feature set and the fourth fused facial feature set are all set with corresponding dynamic fusion weights, and the four dynamic fusion weights are updated through self-learning.

[0113] As an example, see Figure 4The diagram illustrates another process for generating fused facial features. The mixed-modality feature sequence 301 can include K mixed-modality features, corresponding to K color facial features 401, K infrared facial features 402, and K depth-sensing facial features 403. First, taking the first mixed-modality feature in the mixed-modality feature sequence 301 as an example, its included color facial features, infrared facial features, and depth-sensing facial features are sampled to the same feature dimension and then superimposed as the first fused facial feature. This is repeated K times within a frame to fuse the K mixed-modality features, resulting in K first fused facial features (first fused facial feature set 404). Second, the K color facial features 401 undergo inter-frame feature fusion through a recurrent neural network model R1 to obtain the second fused facial feature 405. Next, the K infrared facial features 402 undergo inter-frame feature fusion through a recurrent neural network model R2 to obtain the third fused facial feature 406. Furthermore, the K depth-of-field facial features 403 are fused across frames using a recurrent neural network model R3 to obtain a fourth fused facial feature 407. Finally, the first fused facial feature set 404, the second fused facial feature 405, the third fused facial feature 406, and the fourth fused facial feature 407 are dynamically fused using a weight matrix containing 1×4 dynamic fusion weights to obtain a fused facial feature 408. Recurrent neural network models R1, R2, and R3 all employ the Tiny RNN model.

[0114] Step 105: Perform facial recognition based on the fused facial features to obtain the facial recognition result.

[0115] In some embodiments, the aforementioned executing entity can perform facial recognition based on fused facial features to obtain facial recognition results.

[0116] The facial recognition result represents the recognition outcome for the identified object. Specifically, the facial recognition result may include a facial matching identifier and a result timestamp. The facial matching identifier indicates whether a pre-recorded facial feature matches the aforementioned fused facial features. For example, the facial matching identifier can be represented by "0" and "1". When the facial matching identifier is "0", it indicates that no pre-recorded facial feature matches the aforementioned fused facial features. When the facial matching identifier is "1", it indicates that a pre-recorded facial feature matches the aforementioned fused facial features.

[0117] In practice, the aforementioned execution entity can perform a database search based on the fused facial features and a pre-built facial feature database. When a facial feature matching the fused facial features exists, the output represents a facial recognition result indicating the existence of a pre-recorded facial feature matching the fused facial features. When no facial feature matching the fused facial features exists, the output represents a facial recognition result indicating the absence of a pre-recorded facial feature matching the fused facial features.

[0118] In some optional implementations of certain embodiments, the execution entity performs facial recognition based on the fused facial features to obtain a facial recognition result, including: Step S1: Perform one-dimensional feature mapping on the above fused facial features to obtain the facial feature retrieval vector.

[0119] The facial feature retrieval vector has a vector dimension of 1×512.

[0120] In practice, the fused facial features can be mapped in one dimension using multiple serially connected fully connected layers to obtain a facial feature retrieval vector. Feature mapping enables feature compression, facilitating subsequent vector similarity calculations.

[0121] Step S2: Perform asymmetric vector encryption on the above facial feature retrieval vector to obtain the encrypted facial feature retrieval vector.

[0122] Among them, the encrypted facial feature retrieval vector is the facial feature retrieval vector after data encryption.

[0123] In practice, facial feature retrieval vectors can be encrypted using a pre-allocated asymmetric key to obtain encrypted facial feature retrieval vectors.

[0124] Step S3: Send the encrypted facial feature retrieval vector to the cloud server through the communication tunnel.

[0125] The aforementioned cloud server stores a facial feature vector database, which is used for vector comparison with the encrypted facial feature retrieval vector. Specifically, the facial feature vector database stores the facial feature vectors of registered users. The communication tunnel is a dedicated communication channel between the one-way gate device and the cloud server.

[0126] Step S4: In response to receiving the search feedback information sent by the cloud server, perform anti-counterfeiting verification on the search feedback information.

[0127] The retrieval feedback information consists of vector comparison results sent by a cloud server, targeting the encrypted facial feature retrieval vectors. Specifically, the retrieval feedback information includes: the retrieval feedback result, user access permissions, and server verification certificate. The retrieval feedback result represents the vector comparison result. User access permissions represent the access rights of the identified object. For example, user access permissions may include: access level, permitted locations, and access validity period.

[0128] In practice, the server verification certificate can be obtained by parsing the above retrieval feedback information, and then the server verification certificate can be verified for anti-counterfeiting through a CA (Certificate Authority) center.

[0129] Step S5: In response to passing the anti-counterfeiting verification, generate facial recognition results based on the above retrieval feedback information.

[0130] In practice, the aforementioned executing entity can parse the retrieval feedback information to obtain the retrieval feedback results, which can then be used as the facial recognition results.

[0131] In some optional implementations of some embodiments, the above method further includes: Step S1: In response to the above facial recognition result indicating recognition failure, control the speaker included in the one-way gate device to play an identity abnormality prompt message.

[0132] Among them, the identity anomaly alert message is used to indicate an abnormality in the facial recognition of the identified object. For example, the identity anomaly alert message could be "Identity recognition failed".

[0133] Step S2: In response to the above facial recognition result indicating successful recognition, determine the object's permission information based on the above retrieval feedback information.

[0134] In practice, the aforementioned executing entity can parse the retrieval feedback information to obtain user access permissions, which can then be used as object permission information.

[0135] Step S3: Perform permission matching based on the above object permission information and the access permissions corresponding to the above one-way gate device.

[0136] The one-way turnstile is equipped with a minimum level of access permission.

[0137] In practice, permission matching can be performed based on the permission level, passable location, and pass validity period included in the object's permission information, and the pass permissions corresponding to the aforementioned one-way gate device, in order to determine whether the identified object is used for the pass permissions of the one-way gate device.

[0138] Step S4: In response to the authorization matching, control the above-mentioned one-way gate device to open the barrier gate; Step S5: In response to the failure to match permissions, control the above speaker to play abnormal permission prompt information.

[0139] The permission exception message is used to indicate that the access permission of the identified object is abnormal. For example, the permission exception message could be "Access permissions do not match, please confirm permissions".

[0140] In some optional implementations of some embodiments, the above method further includes: Step S1: Generate real-time log records based on the facial recognition results, object permission information, and gate control information.

[0141] The real-time log records are used to store facial recognition results, the aforementioned object permission information, and the barrier gate control information. The barrier gate control information represents the control operations of the one-way gate device.

[0142] In practice, the facial recognition results, the aforementioned object permission information, and the gate control information can be filled into the log recording template through field filling to obtain real-time log records.

[0143] Step S2: Encrypt the above real-time log records to obtain encrypted log records.

[0144] In practice, real-time log records can be encrypted using a pre-created log encryption key to obtain encrypted log records. The log encryption key can be an asymmetric key that is updated periodically.

[0145] Step S3: Store the encrypted log records into the local log table to obtain the updated local log table.

[0146] The local log table is a log table stored locally on the one-way gate device.

[0147] In practice, encrypted log records can be written into a local log table to obtain an updated local log table.

[0148] Step S4: Generate gate traffic based on the updated local log table.

[0149] Among them, the gate flow rate represents the number of people passing through a one-way gate device per unit time.

[0150] In practice, the number of log records in the updated local log table within a unit of time can be counted as the gate traffic.

[0151] Step S5: Synchronize the above gate traffic flow to the cloud-based pedestrian flow map.

[0152] Among them, the cloud-based pedestrian flow map is a flow map set up in the cloud to visualize changes in pedestrian flow.

[0153] In practice, the traffic flow from the turnstiles can be synchronized to the corresponding location of the one-way turnstile in the cloud-based pedestrian flow map to achieve map updates of the cloud-based pedestrian flow map.

[0154] The various embodiments of this disclosure have the following beneficial effects: the facial recognition method based on mixed modal features in some embodiments of this disclosure improves the success rate and accuracy of facial recognition. In practice, this disclosure first performs intrusion recognition based on a first image sequence to generate movement tendency information. The first image is a low-resolution full-color image captured by a full-color camera, which is mounted on a one-way gate device and faces the entry side. In practice, gate devices, as one of the main physical access control facilities, typically use full-color cameras for image acquisition and facial recognition. When a moving object passes in front of the gate device, it frequently triggers image acquisition and facial recognition by the camera, resulting in unnecessary waste of computing resources. Therefore, this disclosure reduces computing resource consumption by combining low-resolution full-color images for movement tendency recognition. Secondly, in response to the aforementioned movement tendency information indicating that the identified object's movement direction is towards the gate, a second image sequence and a third image sequence are simultaneously acquired by the aforementioned multi-view camera. The second image is a high-resolution full-color image acquired by the aforementioned full-color camera, and the third image is an infrared image acquired by the infrared camera included in the multi-view camera. When the identified object moves towards the gate device, the infrared camera and the full-color camera are triggered to acquire high-resolution images, thereby reducing device power consumption and computing power consumption. Furthermore, complex lighting environments (e.g., backlight, low light, overexposure, etc.) can severely affect image quality; therefore, this disclosure improves the robustness of image acquisition by combining the infrared camera and the full-color camera. Next, mixed-modal feature extraction is performed based on the aforementioned second and third image sequences to obtain a mixed-modal feature sequence, wherein the mixed-modal features consist of color facial features, infrared facial features, and depth-of-field facial features. By extracting color facial features and infrared facial features, mutual complementarity between features is achieved, thereby improving the robustness of facial feature representation. Furthermore, depth-of-field facial features are typically obtained using dual full-color cameras for depth recovery, but this method cannot overcome the influence of the aforementioned lighting environment. Therefore, this disclosure achieves depth recovery by combining infrared and full-color images, thus achieving effective depth recovery without increasing hardware costs. Next, feature fusion is performed on the aforementioned mixed-modality feature sequence to obtain fused facial features. Finally, facial recognition is performed based on the fused facial features to obtain the facial recognition result. In practice, the accuracy of facial recognition based on a single image is difficult to guarantee due to environmental factors and the movement state of the identified object. Therefore, this disclosure improves the expressive power of features by combining color facial features, infrared facial features, mental facial features, and temporal features through feature fusion, thereby increasing the success rate and accuracy of facial recognition.

[0155] Further reference Figure 5As an implementation of the methods shown in the above figures, this disclosure provides some embodiments of a face recognition device based on mixed modal features. These device embodiments are similar to... Figure 1 Corresponding to the method embodiments shown, this facial recognition device based on hybrid modal features can be specifically applied to various electronic devices.

[0156] like Figure 5 As shown, a face recognition device 500 based on hybrid modal features in some embodiments includes: an intrusion recognition unit 501, a synchronous acquisition unit 502, a hybrid modal feature extraction unit 503, a feature fusion unit 504, and a face recognition unit 505. The intrusion recognition unit 501 is configured to perform intrusion recognition based on a first image sequence to generate movement tendency information. The first image is a low-resolution full-color image acquired by a full-color camera included in a multi-view camera system, which is mounted on a one-way gate device and faces the entry side. The synchronous acquisition unit 502 is configured to, in response to the movement tendency information indicating that the movement direction of the identified object is towards the entry side, synchronously acquire a second image using the multi-view camera. The system comprises an image sequence and a third image sequence, wherein the second image is a high-resolution full-color image captured by the aforementioned full-color camera, and the third image is an infrared image captured by the infrared camera included in the aforementioned multi-view camera; a mixed-modality feature extraction unit 503 is configured to extract mixed-modality features based on the aforementioned second image sequence and the aforementioned third image sequence to obtain a mixed-modality feature sequence, wherein the mixed-modality features consist of color facial features, infrared facial features, and depth-of-field facial features; a feature fusion unit 504 is configured to fuse the aforementioned mixed-modality feature sequence to obtain fused facial features; and a facial recognition unit 505 is configured to perform facial recognition based on the aforementioned fused facial features to obtain a facial recognition result.

[0157] It is understandable that the units described in the face recognition device 500 based on mixed modal features are similar to those in the reference device. Figure 1 The steps in the described method correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method are also applicable to the face recognition device 500 based on mixed-modal features and the units contained therein, and will not be repeated here.

[0158] The following is for reference. Figure 6 It shows a schematic diagram of the structure of an electronic device (e.g., a computing device) 600 suitable for implementing some embodiments of the present disclosure. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.

[0159] like Figure 6As shown, the electronic device 600 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory 602 or a program loaded from a storage device 608 into a random access memory 603. The random access memory 603 also stores various programs and data required for the operation of the electronic device 600. The processing unit 601, the read-only memory 602, and the random access memory 603 are interconnected via a bus 604. An input / output interface 605 is also connected to the bus 604.

[0160] Typically, the following devices can be connected to the input / output interface 605: input devices 606 including, for example, a touchscreen, touchpad, keyboard, mouse, camera, microphone, accelerometer, gyroscope, etc.; output devices 607 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; storage devices 608 including, for example, magnetic tape, hard disk, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 6 Each box shown can represent a device or multiple devices as needed.

[0161] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a read-only memory 602. When the computer program is executed by the processing device 601, it performs the functions defined above in the methods of some embodiments of this disclosure.

[0162] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0163] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.

[0164] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device. The aforementioned computer-readable medium carries one or more programs. When the electronic device executes the aforementioned one or more programs, the electronic device causes the following actions: It performs scrambling identification based on a first image sequence to generate movement tendency information, wherein the first image is a low-resolution full-color image captured by a full-color camera, which is mounted on a one-way gate device and faces the entry side; in response to the movement tendency information indicating that the movement direction of the identified object is towards the entry side, it simultaneously acquires a second image sequence and a third image sequence through the aforementioned multi-view camera, wherein the second image is a high-resolution full-color image captured by the aforementioned full-color camera, and the third image is an infrared image captured by an infrared camera, which is included in the aforementioned multi-view camera; it performs mixed-modal feature extraction based on the aforementioned second image sequence and the aforementioned third image sequence to obtain a mixed-modal feature sequence, wherein the mixed-modal features consist of color facial features, infrared facial features, and depth-of-field facial features; it performs feature fusion on the aforementioned mixed-modal feature sequence to obtain fused facial features; and it performs facial recognition based on the aforementioned fused facial features to obtain a facial recognition result.

[0165] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0166] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0167] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.

[0168] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described inventive concept. For example, technical solutions formed by substituting the above-described features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.

Claims

1. A face recognition method based on hybrid modal features, characterized in that, include: Intrusion recognition is performed based on the first image sequence to generate movement tendency information, wherein the first image is a low-resolution full-color image captured by a full-color camera, which is part of a multi-view camera and is mounted on a one-way gate device and faces the entrance side. In response to the movement tendency information indicating that the movement direction of the identified object is toward the gate side, a second image sequence and a third image sequence are simultaneously acquired by the multi-view camera, wherein the second image is a high-resolution full-color image acquired by the full-color camera, and the third image is an infrared image acquired by the infrared camera included in the multi-view camera; Mixed modal features are extracted based on the second image sequence and the third image sequence to obtain a mixed modal feature sequence, wherein the mixed modal features consist of color facial features, infrared facial features and depth-of-field facial features; The hybrid modality feature sequence is fused to obtain fused facial features; Facial recognition is performed based on the fused facial features to obtain facial recognition results.

2. The face recognition method based on mixed modal features according to claim 1, characterized in that, The step of performing facial recognition based on the fused facial features to obtain facial recognition results includes: A one-dimensional feature mapping is performed on the fused facial features to obtain a facial feature retrieval vector, wherein the vector dimension of the facial feature retrieval vector is 1×512. The facial feature retrieval vector is subjected to asymmetric vector encryption to obtain the encrypted facial feature retrieval vector. The encrypted facial feature retrieval vector is sent to a cloud server through a communication tunnel. The cloud server stores a facial feature vector database, which is used to perform vector comparison with the encrypted facial feature retrieval vector. In response to receiving the search feedback information sent by the cloud server, the search feedback information is subjected to anti-counterfeiting verification; In response to passing the anti-counterfeiting verification, a facial recognition result is generated based on the retrieval feedback information.

3. The face recognition method based on mixed-modal features according to claim 2, characterized in that, The method further includes: In response to the facial recognition result indicating recognition failure, the speaker included in the one-way gate device is controlled to play an identity abnormality prompt message; In response to the facial recognition result indicating successful recognition, the object permission information is determined based on the retrieval feedback information; Permission matching is performed based on the object permission information and the access permission corresponding to the one-way gate device; In response to permission matching, the one-way gate device is controlled to open the barrier gate; In response to a failure to match permissions, the speaker is controlled to play an error message indicating an abnormal permission.

4. The face recognition method based on mixed modal features according to claim 3, characterized in that, The method further includes: Based on the facial recognition results, the object permission information, and the gate control information, a real-time log record is generated. The real-time log records are encrypted to obtain encrypted log records; The encrypted log records are stored in a local log table to obtain an updated local log table; Based on the updated local log table, generate the gate traffic; The gate traffic flow is synchronized to the cloud-based pedestrian flow map.

5. The face recognition method based on mixed modal features according to claim 4, characterized in that, The step of performing scrambling detection based on the first image sequence to generate movement tendency information includes: The first image sequence is subjected to frame extraction to obtain the target image sequence; For each target image in the target image sequence, perform the following first processing step: The target image is subjected to object blur feature extraction to obtain object blur features; Based on the fuzzy features of the object, determine the region of interest of the object; Perform one-dimensional feature mapping on the fuzzy features of the object to obtain the fuzzy feature matching vector of the object; Based on the object fuzzy feature matching vector corresponding to the target image, image association is performed on the target images in the target image sequence to obtain an associated image sequence; Based on the region of interest corresponding to the associated image in the associated image sequence, a spatiotemporal search region sequence is generated, wherein the spatiotemporal search regions in the spatiotemporal search region sequence correspond one-to-one with the first image in the first image sequence; For each first image in the first image sequence, local features are extracted from the first image within the spatiotemporal search region corresponding to the first image to obtain local image features; Optical flow tracing is performed based on the obtained local image feature sequence to generate object motion features and object movement trajectory features; Based on the object's motion characteristics and the object's movement trajectory characteristics, movement tendency information is generated.

6. The face recognition method based on mixed modal features according to claim 5, characterized in that, The step of extracting mixed modal features based on the second image sequence and the third image sequence to obtain a mixed modal feature sequence includes: For each second image in the second image sequence and the corresponding third image in the third image sequence, the following second processing step is performed: Distortion correction is performed on the second image and the third image respectively to obtain the corrected second image and the corrected third image; Multi-scale image feature extraction is performed on the corrected second image and the corrected third image respectively to obtain the first multi-scale image feature and the second multi-scale image feature; Based on the features of the first multi-scale image, a facial region search is performed to obtain the facial region of interest. Local features of the first multi-scale image features and the second multi-scale image features are extracted from the region of interest on the face to obtain the first local features and the second local features, which are included as the color facial features and infrared facial features in the mixed modality features. Based on the first local feature, a set of colored facial feature points is generated; Based on the second local feature, generate a set of infrared facial feature points; Facial feature point matching is performed on the set of colored facial feature points and the set of infrared facial feature points to obtain a set of facial feature point pairs. Based on the facial region of interest, the second image and the third image are segmented to obtain a first facial region image and a second facial region image. Based on the set of facial feature point pairs, image registration is performed on the first facial region image and the second facial region image to obtain the registered region; A disparity feature map is generated based on the registration region, the first facial region image, and the second facial region image. Depth conversion is performed based on the disparity feature map to obtain the depth facial features included in the mixed modality features.

7. The face recognition method based on mixed modal features according to claim 6, characterized in that, The feature fusion of the mixed modality feature sequence to obtain fused facial features includes: For each hybrid modality feature in the hybrid modality feature sequence, intra-frame feature fusion is performed on the color facial features, infrared facial features, and depth facial features included in the hybrid modality feature to obtain a first fused facial feature; Inter-frame feature fusion is performed on the color facial features included in the mixed modality feature sequence to obtain a second fused facial feature; Inter-frame feature fusion is performed on the infrared facial features included in the mixed modality feature sequence to obtain a third fused facial feature; Inter-frame feature fusion is performed on the depth-of-field facial features included in the hybrid modal feature sequence to obtain a fourth fused facial feature; The first fused facial feature set, the second fused facial feature set, the third fused facial feature set, and the fourth fused facial feature set are dynamically fused to obtain fused facial features.

8. A facial recognition device based on hybrid modal features, characterized in that, include: The intrusion detection unit is configured to perform intrusion detection based on a first image sequence to generate movement tendency information, wherein the first image is a low-resolution full-color image captured by a full-color camera, which is part of a multi-view camera and is mounted on a one-way gate device and faces the entrance side. The synchronous acquisition unit is configured to, in response to the movement tendency information indicating that the movement direction of the identified object is toward the gate side, synchronously acquire a second image sequence and a third image sequence through the multi-view camera, wherein the second image is a high-resolution full-color image acquired by the full-color camera, and the third image is an infrared image acquired by the infrared camera included in the multi-view camera; The mixed modality feature extraction unit is configured to extract mixed modality features based on the second image sequence and the third image sequence to obtain a mixed modality feature sequence, wherein the mixed modality features consist of color facial features, infrared facial features and depth-of-field facial features; A feature fusion unit is configured to perform feature fusion on the mixed modality feature sequence to obtain fused facial features; A facial recognition unit is configured to perform facial recognition based on the fused facial features to obtain a facial recognition result.

9. An electronic device, characterized in that, include: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 7.

10. A computer-readable medium, characterized in that, It stores a computer program thereon, wherein the computer program, when executed by a processor, implements the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Control method and system of channel gate

    CN113393603A

  • Multi-mode face recognition method and device and intelligent door lock

    CN117197851A

  • Artificial intelligence visual detection method based on feature face method

    CN119007261A

  • Face recognition method, device, system and electronic instrument, as well as readable storage medium

    JP2024005748A