Face recognition method and device based on biological characteristics and time-space fusion

Through the CNN-Transformer hybrid architecture face recognition model integrates PPG signals, face images and network environment information, it solves the accuracy and robustness of face recognition in complex environments, and achieves higher identity authentication accuracy and security.

CN120375439AActive Publication Date: 2025-07-25BEIJING GZT NETWORK TECH +1

Patent Information

Application Number
CN202510412186.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-25
Estimated Expiration
2045-04-02

AI Technical Summary

Technical Problem

The existing facial recognition technology lacks accuracy and robustness in complex environments, mainly due to its dependence on single-modal image information and unfusion dynamic multimodal biometrics and space-time context information.

Method used

The face recognition model with a hybrid architecture of CNN-Transformer is adopted to integrate PPG signals, face image information and access network environment information, and obtain global dependencies through the multi-head self-attention mechanism and perform feature fusion to generate recognition results.

Benefits of technology

It improves the accuracy and security of identity authentication, effectively prevents the forgery and tampering of static biometric data, and enhances the robustness of face recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375439A_ABST
    Figure CN120375439A_ABST
Patent Text Reader

Abstract

The invention discloses a face recognition method and device based on biological characteristics and space-time fusion. The face recognition method based on biological characteristics and space-time fusion comprises the following steps: acquiring a PPG signal and face image information through a camera device; obtaining access network environment information; obtaining a trained face recognition model; and inputting the PPG signal, the face image information and the access network environment information into the face recognition model to obtain a recognition result. According to the method, multiple biological characteristics (such as face recognition and PPG signals) are dynamically fused, so that the accuracy and safety of identity authentication are improved. By dynamically monitoring and analyzing changes of different biological characteristics, counterfeiting and tampering of static biological characteristic data can be effectively prevented.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and particularly to a face recognition method based on biometrics and spatio-temporal fusion, and a face recognition device based on biometrics and spatio-temporal fusion. Background Art

[0002] In the prior art, for face recognition, it mainly relies on image processing and machine learning algorithms. The technical solutions usually include the following steps:

[0003] Face image acquisition: Acquire face images through devices such as cameras.

[0004] Preprocessing: Perform preprocessing operations such as grayscale conversion, denoising, and normalization on the acquired face images to improve the accuracy of subsequent processing.

[0005] Feature extraction: Use algorithms to extract information that can represent individual features from the preprocessed face images, such as facial contours, textures, key points, etc.

[0006] Feature matching and recognition: Compare the extracted face features with the feature templates in the database, calculate the similarity score, and make a decision based on the score to determine the identity.

[0007] The disadvantages of the prior art are as follows:

[0008] In complex environments (such as light changes, occlusion, angle changes, etc.), due to relying only on single-modal image information, the accuracy of face recognition will be greatly affected.

[0009] Failure to fuse dynamic multi-modal biometrics (such as facial expressions, head postures, etc.) and spatio-temporal context information (such as changes between consecutive frames, changes in the position of the face in the image, etc.) results in insufficient robustness of the recognition system when facing complex scenarios. Summary of the Invention

[0010] The purpose of the present invention is to provide a face recognition method based on biometrics and spatio-temporal fusion to at least solve one of the above technical problems.

[0011] In one aspect of the present invention, there is provided a face recognition method based on biometrics and spatio-temporal fusion, and the face recognition method based on biometrics and spatio-temporal fusion includes:

[0012] Obtain PPG signals and face image information through a camera device;

[0013] Obtain access network environment information;

[0014] Obtain a trained face recognition model;

[0015] Input the PPG signal, face image information, and access network environment information into the face recognition model to obtain a recognition result.

[0016] Optionally, the face recognition model is a CNN-Transformer hybrid architecture face recognition model.

[0017] Optionally, the CNN-Transformer hybrid architecture face recognition model includes:

[0018] An input layer for obtaining the PPG signal, face image information, and access network environment information;

[0019] A CNN feature extraction module for extracting face image features of the face image information;

[0020] An LSTM sequence processing module for obtaining PPG features based on the PPG signal and network environment information features based on the access network environment information;

[0021] A Transformer global dependency modeling module for receiving the face image features, PPG features, and network environment information features, and generating global features after obtaining global dependency relationships through the multi-head self-attention mechanism;

[0022] A feature fusion module for fusing the face image features, PPG features, network environment information features, and global features to obtain fused features;

[0023] A classification module for obtaining the fused features and obtaining a recognition result based on the fused features.

[0024] Optionally, the face recognition method based on biometric and spatio-temporal fusion further includes:

[0025] Training the CNN-Transformer hybrid architecture face recognition model.

[0026] Optionally, the training of the CNN-Transformer hybrid architecture face recognition model includes:

[0027] Obtaining training data;

[0028] Obtaining the CNN-Transformer hybrid architecture face recognition model;

[0029] Training the CNN-Transformer hybrid architecture face recognition model with the training data.

[0030] Optionally, the CNN-Transformer hybrid architecture face recognition model adopts the following loss function:

[0031] Wherein,

[0032] L represents the loss function, x represents the feature vector output by the model, usually a high-dimensional vector containing the features extracted from the input data, y represents the true label, C represents the number of categories, and y i represents the i-th element in the true label vector y. If the sample belongs to the i-th category, then y i = 1, otherwise y i = 0, W i represents the weight vector of category i, represents the dot product of the weight vector W i and the feature vector x, used to measure the similarity between them, ||W i || represents the L2 norm of the weight vector W i The L2 norm of ||x|| represents the L2 norm of the feature vector x, ∈ represents a positive number less than 10, used to prevent the denominator from being zero and ensure the stability of the calculation, α represents a positive coefficient, used to control the influence of the adaptive weight term on the overall loss function, and (1 - y i ) is a constraint condition to ensure that only the categories with non-true labels are considered. If the sample belongs to the i-th category, then (1 - y i ) = 0, otherwise (1 - y i ) = 1, represents the square of the L2 norm of the weight vector W (here W is the set of all category weight vectors W i ), that is, the sum of the squares of all weight elements, and β represents a positive coefficient.

[0033] This application also provides a face recognition device based on biometric and spatio-temporal fusion. The face recognition device based on biometric and spatio-temporal fusion includes:

[0034] An information acquisition module, which is used to acquire the face image information, PPG signal, and access network environment information transmitted by the camera device;

[0035] A face recognition model acquisition module, which is used to acquire the trained face recognition model;

[0036] A recognition result acquisition module, which is used to input the PPG signal, face image information, and access network environment information into the face recognition model to obtain the recognition result.

[0037] Beneficial effects

[0038] The face recognition method based on biometrics and spatio-temporal fusion of the present application has the following advantages:

[0039] Dynamic multimodal biometric recognition: The present invention proposes to dynamically fuse multiple biometric features (such as face recognition, PPG signals, etc.) to improve the accuracy and security of identity authentication. By dynamically monitoring and analyzing the changes of different biometric features, it is possible to effectively prevent the forgery and tampering of static biometric data.

[0040] Spatio-temporal context fusion: Introduce spatio-temporal context information to conduct a deeper understanding and analysis of face images. By modeling the time series and spatial relationships of face images, it is possible to effectively capture the dynamic changes and context information in face images and improve the robustness of face recognition. Description of the drawings

[0041] Figure 1 is a schematic flowchart of a face recognition method based on biometrics and spatio-temporal fusion according to an embodiment of the present application.

[0042] Figure 2 is a schematic diagram of an electronic device for implementing the face recognition method based on biometrics and spatio-temporal fusion according to an embodiment of the present application. Detailed implementation manners

[0043] To make the purpose, technical solutions, and advantages of the implementation of the present application clearer, the technical solutions in the embodiments of the present application will be described in more detail below with reference to the drawings in the embodiments of the present application. In the drawings, the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions from beginning to end. The described embodiments are some but not all of the embodiments of the present application. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application. The embodiments of the present application will be described in detail below with reference to the drawings.

[0044] Such as Figure 1 The face recognition method based on biometrics and spatio-temporal fusion shown includes:

[0045] Step 1: Obtain PPG signals and face image information through a camera device;

[0046] Step 2: Obtain access network environment information;

[0047] Step 3: Obtain a trained face recognition model;

[0048] Step 4: Input the PPG signal, face image information, and access network environment information into the face recognition model to obtain a recognition result.

[0049] The face recognition method based on biometric and spatio-temporal fusion of the present application has the following advantages:

[0050] Dynamic multi-modal biometric recognition: The present invention proposes to dynamically fuse multiple biometric features (such as face recognition, PPG signal, etc.) to improve the accuracy and security of identity authentication. By dynamically monitoring and analyzing the changes in different biometric features, it is possible to effectively prevent the forgery and tampering of static biometric data.

[0051] Spatio-temporal context fusion: Introduce spatio-temporal context information to conduct a deeper understanding and analysis of face images. By modeling the time series and spatial relationships of face images, it is possible to effectively capture the dynamic changes and context information in face images and improve the robustness of face recognition.

[0052] In this embodiment, the PPG signal can be obtained by the following method:

[0053] The basic principle of PPG signal acquisition is to use an imaging device to obtain a video of the surface skin area of the human body. The change in blood volume will cause a change in the brightness of the surface skin area, and this change in brightness will cause a change in the grayscale value of each frame of the video. By analyzing these changes in grayscale values, the PPG signal can be obtained.

[0054] In this embodiment, the access network environment information may include information such as time stamps, geographical locations, device fingerprints, and network environments.

[0055] In this embodiment, the face recognition model is a CNN-Transformer hybrid architecture face recognition model.

[0056] In this embodiment, the CNN-Transformer hybrid architecture face recognition model includes:

[0057] An input layer for obtaining the PPG signal, face image information, and access network environment information;

[0058] A CNN feature extraction module for extracting face image features of the face image information;

[0059] An LSTM sequence processing module for obtaining PPG features according to the PPG signal and obtaining network environment information features according to the access network environment information;

[0060] Transformer global dependency modeling module, which is used to receive the face image features, PPG features, and network environment information features, and generate global features after obtaining global dependency relationships through the multi-head self-attention mechanism;

[0061] Feature fusion module, which is used to fuse the face image features, PPG features, network environment information features, and global features to obtain fused features;

[0062] Classification module, which is used to obtain the fused features and obtain recognition results based on the fused features.

[0063] In this embodiment, the face recognition method based on biometric and spatio-temporal fusion further includes:

[0064] Training the CNN-Transformer hybrid architecture face recognition model.

[0065] In this embodiment, the training of the CNN-Transformer hybrid architecture face recognition model includes:

[0066] Obtaining training data;

[0067] Obtaining the CNN-Transformer hybrid architecture face recognition model;

[0068] Training the CNN-Transformer hybrid architecture face recognition model with the training data.

[0069] In this embodiment, the CNN-Transformer hybrid architecture face recognition model adopts the following loss function:

[0070] Among them,

[0071] L represents the loss function, x represents the feature vector output by the model, usually a high-dimensional vector containing features extracted from the input data, y represents the true label, C represents the number of classes, y i represents the i-th element in the true label vector y. If the sample belongs to the i-th class, then y i = 1, otherwise y i = 0, W i represents the weight vector of class i, represents the dot product of the weight vector W i and the feature vector x, which is used to measure the similarity between them, ||W i || represents the weight vector W iThe L2 norm of, ||x|| represents the L2 norm of the feature vector x, ∈ represents a positive number less than 10, which is used to prevent the denominator from being zero and ensure the stability of the calculation, α represents a positive coefficient, which is used to control the influence of the adaptive weight term on the overall loss function, (1 - y i ) is a constraint condition to ensure that only the categories with non-real labels are considered. If the sample belongs to the i-th category, then (1 - y i ) = 0, otherwise (1 - y i ) = 1. represents the square of the L2 norm of the weight vector W (here W is the set of all category weight vectors W i ), that is, the sum of the squares of all weight elements. β represents a positive coefficient, and its value can be set according to needs.

[0072] By using the loss function of the present application, the following advantages are obtained:

[0073] 1. By converting the cross-entropy loss function into an angle-based form, the angle information in the face recognition task is better utilized, which helps to improve the recognition accuracy of the model.

[0074] 2. The adaptive weight mechanism enables the model to dynamically adjust its contribution to the update of model parameters according to the difficulty of samples. This helps the model to learn difficult samples more efficiently, thereby improving the overall performance.

[0075] 3. The addition of the L2 regularization term helps to prevent the model from overfitting and improve the generalization ability of the model.

[0076] The present application also provides a face recognition device based on the fusion of dynamic multi-modal biometrics and spatio-temporal context. The face recognition device based on the fusion of dynamic multi-modal biometrics and spatio-temporal context includes an information acquisition module, a face recognition model acquisition module, and a recognition result acquisition module, where

[0077] The information acquisition module is used to acquire the face image information, PPG signal, and access network environment information transmitted by the camera device;

[0078] The face recognition model acquisition module is used to acquire the trained face recognition model;

[0079] The recognition result acquisition module is used to input the PPG signal, face image information, and access network environment information into the face recognition model to obtain the recognition result

[0080] It should be noted that the foregoing explanations of the method embodiments also apply to the device of this embodiment, and will not be repeated here.

[0081] The present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and capable of running on the processor. When the processor executes the computer program, the face recognition method based on biometrics and spatio-temporal fusion as described above is implemented.

[0082] The present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the face recognition method based on biometrics and spatio-temporal fusion as described above can be implemented.

[0083] Figure 2 It is an exemplary structural diagram of an electronic device capable of implementing the face recognition method based on biometrics and spatio-temporal fusion provided by an embodiment of the present application.

[0084] As Figure 2 shown, the electronic device includes an input device 501, an input interface 502, a central processing unit 503, a memory 504, an output interface 505, and an output device 506. Among them, the input interface 502, the central processing unit 503, the memory 504, and the output interface 505 are connected to each other through a bus 507. The input device 501 and the output device 506 are respectively connected to the bus 507 through the input interface 502 and the output interface 505, and then connected to other components of the electronic device. Specifically, the input device 501 receives input information from the outside and transmits the input information to the central processing unit 503 through the input interface 502; the central processing unit 503 processes the input information based on computer-executable instructions stored in the memory 504 to generate output information, temporarily or permanently stores the output information in the memory 504, and then transmits the output information to the output device 506 through the output interface 505; the output device 506 outputs the output information to the outside of the electronic device for the user to use.

[0085] That is to say, Figure 2 the electronic device shown can also be implemented as including: a memory storing computer-executable instructions; and one or more processors, and when the one or more processors execute the computer-executable instructions, the face recognition method based on biometrics and spatio-temporal fusion described in conjunction with Figure 1 can be implemented.

[0086] In one embodiment, Figure 2 the electronic device shown can be implemented as including: a memory 504 configured to store executable program code; one or more processors configured to run the executable program code stored in the memory 504 to execute the face recognition method based on biometrics and spatio-temporal fusion in the above embodiment.

[0087] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0088] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0089] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0090] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, which includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the figures. For example, two consecutive blocks marked may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or overall flowcharts, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0091] In this embodiment, the so-called processor may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor, or the processor may also be any conventional processor, etc.

[0092] The memory can be used to store computer programs and / or modules. The processor realizes various functions of the device / terminal device by running or executing the computer programs and / or modules stored in the memory, and by calling the data stored in the memory. The memory may mainly include a program storage area and a data storage area. Among them, the program storage area may store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created according to the use of the mobile phone (such as audio data, phone book, etc.). In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0093] In this embodiment, if the modules / units integrated in the device / terminal device are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. Although this application is disclosed above in preferred embodiments, it is not actually used to limit this application. Any person skilled in the art can make possible changes and modifications without departing from the spirit and scope of this application. Therefore, the protection scope of this application should be subject to the scope defined by the claims of this application.

[0094] Those skilled in the art should understand that the embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, this application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0095] In addition, obviously, the term "including" does not exclude other units or steps. The multiple units, modules, or devices stated in the apparatus claims can also be implemented by one unit or a general apparatus through software or hardware.

[0096] Although the present invention has been described in detail above with general descriptions and specific embodiments, based on the present invention, some modifications or improvements can be made, which are obvious to those skilled in the art. Therefore, these modifications or improvements made without departing from the spirit of the present invention all fall within the scope of protection required by the present invention.

Claims

1. A face recognition method based on biometric and spatio-temporal fusion, characterized in that, The face recognition method based on biometric and spatio-temporal fusion includes: Obtaining PPG signals and face image information through a camera device; Obtaining access network environment information; Obtaining a trained face recognition model; Inputting the PPG signals, face image information, and access network environment information into the face recognition model to obtain a recognition result.

2. The face recognition method based on biometrics and spatio-temporal fusion according to claim 1, wherein The face recognition model is a CNN-Transformer hybrid architecture face recognition model.

3. The face recognition method based on biometrics and spatio-temporal fusion according to claim 2, wherein The CNN-Transformer hybrid architecture face recognition model includes: An input layer for obtaining the PPG signals, face image information, and access network environment information; A CNN feature extraction module for extracting face image features of the face image information; An LSTM sequence processing module for obtaining PPG features based on the PPG signals and network environment information features based on the access network environment information; A Transformer global dependency modeling module for receiving the face image features, PPG features, and network environment information features, and generating global features after obtaining global dependency relationships through a multi-head self-attention mechanism; A feature fusion module for fusing the face image features, PPG features, network environment information features, and global features to obtain fused features; A classification module for obtaining the fused features and obtaining a recognition result based on the fused features.

4. The face recognition method based on biometrics and spatio-temporal fusion according to claim 3, wherein The face recognition method based on biometric and spatio-temporal fusion further includes: Training the CNN-Transformer hybrid architecture face recognition model.

5. The face recognition method based on biometrics and spatio-temporal fusion according to claim 4, wherein, The training of the CNN-Transformer hybrid architecture face recognition model includes: Obtaining training data; Obtaining the CNN-Transformer hybrid architecture face recognition model; Training the CNN-Transformer hybrid architecture face recognition model with the training data.

6. The face recognition method based on biometrics and spatio-temporal fusion according to claim 5, wherein The CNN-Transformer hybrid architecture face recognition model uses the following loss function: Among them, L represents the loss function, x represents the feature vector output by the model, usually a high-dimensional vector containing features extracted from the input data, y represents the true label, C represents the number of categories, and y i : represents the i-th element in the true label vector y. If the sample belongs to the i-th category, then y i = 1, otherwise y i = 0, W i represents the weight vector of category i, represents the weight vector W i and the dot product of the feature vector x, used to measure the similarity between them, ||W i || represents the L2 norm of the weight vector W i of, ||x|| represents the L2 norm of the feature vector x, ∈ represents a positive number less than 10, used to prevent the denominator from being zero and ensure the stability of the calculation, α represents a positive coefficient, used to control the influence of the adaptive weight term on the overall loss function, (1 - y i ) is a constraint condition to ensure that only the categories with non-true labels are considered; if the sample belongs to the i-th category, then (1 - y i ) = 0, otherwise (1 - y i ) = 1, represents the square of the L2 norm of the weight vector W (here W is the set of all category weight vectors W i )), that is, the sum of the squares of all weight elements, and β represents a positive coefficient.

7. A face recognition device based on biometric and spatio-temporal fusion, characterized in that, The face recognition device based on biometric and spatio-temporal fusion includes: An information acquisition module for acquiring face image information, PPG signals, and access network environment information transmitted by a camera device; A face recognition model acquisition module for acquiring a trained face recognition model; A recognition result acquisition module for inputting the PPG signals, face image information, and access network environment information into the face recognition model to obtain a recognition result.

Citation Information

Patent Citations

  • Video-based identity authentication method and video equipment

    CN111325118A

  • Model training method and device and identity verification method and device

    CN115497146A

  • Photoelectric pulse signal enhanced face recognition method

    CN117034238A

  • Face recognition method and intelligent computer all-in-one machine applying face recognition method

    CN119559685A

  • Method and system for automatic biometric authentication based on facial spatio-temporal features

    KR1020170045093A

Cited By

  • Multi-modal identity authentication method and system based on PPG characteristic spectrum construction

    CN121350560A