A multi-scale feature aggregation method and system for liveness detection
By converting RGB images into HSV images and performing multi-scale feature aggregation, the computational complexity problem caused by the increase in parameters in existing liveness detection algorithms is solved, and efficient and highly generalized liveness detection is achieved.
Patent Information
- Application Number
- CN202211153579.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-21
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2042-09-21
AI Technical Summary
When existing liveness detection algorithms use RGB and HSV image features, the increase in network parameters leads to increased computational complexity, reduced running speed, and insufficient feature aggregation, which affects detection efficiency and generalization.
A multi-scale feature aggregation method is used to convert RGB images into HSV images and fuse them into RGB-HSV images. After extracting features through the backbone network, the feature depth expansion module and the multi-feature extraction module are used to obtain more contextual information. Training is performed under the constraints of the cross-entropy loss function to reduce the network computational complexity and expand the receptive field.
While maintaining the same running speed, the detection performance and efficiency are improved, more contextual information is obtained through multi-scale feature aggregation, and lightweight and highly generalized liveness detection is achieved.
Smart Images

Figure CN115546907B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image detection, and in particular to a method and system for multi-scale feature aggregation liveness detection. Background Art
[0002] Liveness detection is a method used in identity verification scenarios to determine a subject's true physiological characteristics. In facial recognition applications, liveness detection can verify that the user is truly alive by using facial landmark location and tracking technologies, including blinking, opening the mouth, shaking the head, and nodding. It effectively protects against common attacks such as photos, face swaps, masks, occlusions, and screen capture, helping users identify fraudulent activity and protecting their interests. With the development of deepfakes and facial recognition countermeasures, liveness detection technology has become a research hotspot.
[0003] Facial recognition systems primarily consist of four components: face detection, keypoint detection, liveness detection, and face matching. With the advancement of deepfake technology and facial countermeasures, liveness detection technology faces unprecedented challenges. Designing lightweight and highly generalizable algorithms is a current research hotspot. Existing liveness detection algorithms primarily utilize massive amounts of data and the powerful feature extraction capabilities of convolutional neural networks to train a binary classifier. While existing algorithms have proven to improve the robustness of liveness detection by using RGB and HSV spatial representations of images, they often utilize a two-way network for feature extraction, which introduces additional parameters and slows down performance.
[0004] Existing liveness detection algorithms have demonstrated that utilizing both RGB and HSV image representations improves accuracy and generalization. These algorithms primarily utilize a two-branch network to extract features from RGB and HSV images, respectively. These features are then aggregated at the end of the convolutional neural network before being fed into a classifier for binary classification. However, this doubles the number of network parameters required for both branches, and simply aggregating features at the end fails to optimize network training. This increased number of network parameters increases computational complexity and slows network performance. Summary of the Invention
[0005] In order to solve the above technical problems existing in the prior art, the present invention proposes a multi-scale feature aggregation liveness detection method and system to solve the above technical problems.
[0006] According to one aspect of the present invention, a multi-scale feature aggregation method for liveness detection is proposed, comprising:
[0007] S1: Convert the RGB image into an HSV image through image transformation, fuse the RGB image and HSV image into an RGB-HSV image and send it to the backbone network;
[0008] S2: The features extracted by the backbone network are fed into the feature depth expansion module, and the output is fed into the multi-feature extraction module to obtain more context information;
[0009] S3: The final output is passed through the pooling layer and the classification layer, and trained under the constraints of the cross entropy loss function.
[0010] In some specific embodiments, the backbone network may use ResNet18 or MobileNetV2.
[0011] In some specific embodiments, the number of channels of the RGB-HSV image is 6, and the number of input channels of the first convolutional layer of the backbone network is 3.
[0012] In some specific embodiments, the feature depth expansion module includes a 3*3 convolution, a BN layer and a ReLU activation function.
[0013] In some specific embodiments, the multi-feature extraction module includes two branches, the upper branch includes a simulated 3*3 convolution composed of 3*1 and 1*3 convolutions, and the lower branch includes a 3*3 dilated convolution with a dilation rate of 3. The upper and lower branches of the multi-feature extraction module are fused and then passed through a BN layer and ReLU to prevent overfitting.
[0014] In some specific embodiments, there are seven multi-feature extraction modules, and the output of the feature depth expansion module is connected to the output of the fourth and seventh multi-feature extraction modules respectively.
[0015] According to a second aspect of the present invention, a computer-readable storage medium is provided, on which one or more computer programs are stored. When the one or more computer programs are executed by a computer processor, any one of the above methods is implemented.
[0016] According to a third aspect of the present invention, a multi-scale feature aggregation liveness detection system is proposed, the system comprising:
[0017] An image fusion unit is configured to convert the RGB image into an HSV image through image transformation, fuse the RGB image and the HSV image into an RGB-HSV image, and send the image to the backbone network;
[0018] A feature depth expansion unit is configured to feed the features extracted by the backbone network into the feature depth expansion module, and feed the output into the multi-feature extraction module to obtain more context information;
[0019] Training unit: The configuration is used to pass the final output through the pooling layer and the classification layer, and train under the constraints of the cross entropy loss function.
[0020] In some specific embodiments, the backbone network can be ResNet18 or MobileNetV2, the number of channels of the RGB-HSV image is 6, and the number of input channels of the first convolutional layer of the backbone network is 3.
[0021] In some specific embodiments, the feature depth expansion module includes a 3*3 convolution, a BN layer and a ReLU activation function.
[0022] In some specific embodiments, the multi-feature extraction module includes two branches, the upper branch includes a simulated 3*3 convolution composed of 3*1 and 1*3 convolutions, and the lower branch includes a 3*3 dilated convolution with a dilation rate of 3. The upper and lower branches of the multi-feature extraction module are fused and then passed through a BN layer and ReLU to prevent overfitting.
[0023] In some specific embodiments, there are seven multi-feature extraction modules, and the output of the feature depth expansion module is connected to the output of the fourth and seventh multi-feature extraction modules respectively.
[0024] The present invention proposes a multi-scale feature aggregation liveness detection method and system, which has the characteristics of small parameter number, large receptive field of image, and multi-scale feature aggregation. It uses void convolution to expand the receptive field of the network and obtain more contextual information. It uses RGB images and HSV images as 6-channel images as input. The running speed is consistent with the running speed of using only RGB images, achieving a good balance between performance and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The accompanying drawings are included to provide a further understanding of the embodiments and are incorporated into and constitute a part of this specification. The accompanying drawings illustrate the embodiments and, together with the description, serve to explain the principles of the invention. Other embodiments and many of the intended advantages of the embodiments will be readily apparent as they become better understood by reference to the following detailed description. Other features, objects, and advantages of the present application will become more apparent by reading the detailed description of the non-limiting embodiments made with reference to the following drawings:
[0026] Figure 1 This is a flowchart of a method for liveness detection using multi-scale feature aggregation according to an embodiment of the present application;
[0027] Figure 2 This is an algorithm framework diagram of a multi-scale feature aggregation liveness detection method according to a specific embodiment of the present application;
[0028] Figure 3 This is a framework diagram of a feature depth expansion module of a specific embodiment of the present application;
[0029] Figure 4This is a framework diagram of a multi-feature extraction module of a specific embodiment of the present application;
[0030] Figure 5 This is a framework diagram of a multi-scale feature aggregation liveness detection system according to an embodiment of the present application;
[0031] Figure 6 It is a structural diagram of a computer system suitable for implementing the electronic device of the embodiment of the present application. DETAILED DESCRIPTION
[0032] The present application will be further described in detail below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the relevant invention and are not intended to limit the invention. It should also be noted that, for ease of description, only portions relevant to the relevant invention are shown in the accompanying drawings.
[0033] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0034] According to a multi-scale feature aggregation method for liveness detection according to an embodiment of the present application, Figure 1 FIG. 1 shows a flow chart of a method for detecting a living body by multi-scale feature aggregation according to an embodiment of the present application. Figure 1 As shown, the method includes:
[0035] S101: Convert the RGB image into an HSV image through image transformation, fuse the RGB image and the HSV image into an RGB-HSV image and send it to the backbone network.
[0036] In a specific embodiment, the backbone network can be ResNet18 or MobileNetV2. The number of channels of the RGB-HSV image is 6, and the number of input channels of the first convolutional layer of the backbone network is 3.
[0037] S102: The features extracted by the backbone network are fed into the feature depth expansion module, and the output is fed into the multi-feature extraction module to obtain more contextual information.
[0038] In a specific embodiment, the feature depth expansion module includes a 3*3 convolution, a batch normalization layer, and a ReLU activation function. The multi-feature extraction module includes two branches: the upper branch includes a simulated 3*3 convolution composed of 3*1 and 1*3 convolutions, and the lower branch includes a 3*3 dilated convolution with a dilation rate of 3. The upper and lower branches of the multi-feature extraction module are fused and then passed through a batch normalization layer and ReLU to prevent overfitting. There are seven multi-feature extraction modules, and the outputs of the feature depth expansion module are connected to the outputs of the fourth and seventh multi-feature extraction modules, respectively.
[0039] S103: The final output is passed through the pooling layer and the classification layer, and trained under the constraints of the cross entropy loss function.
[0040] Figure 2 FIG. 4 shows an algorithm framework diagram of a living body detection method based on multi-scale feature aggregation according to a specific embodiment of the present invention, as shown in FIG. Figure 2 As shown, the following steps are included:
[0041] Step S1: Convert the RGB image to an HSV image through image transformation, fuse it into an RGB-HSV image with 6 channels, and send it to the backbone network (ResNet18 or MobileNetV2 can be used). Since the previous backbone network receives a 3-channel image, it is necessary to change the input channel number of the first convolutional layer of the backbone network to 3.
[0042] Step S2: After the partial convolution block extraction of the backbone network, the number of channels of the feature map is too shallow, and a small number of feature maps are not suitable for subsequent high-level feature extraction. Therefore, the features output by the backbone network in step S1 are sent to the feature depth expansion module. The framework structure of the feature depth expansion module is as follows: Figure 3 The framework diagram of the feature depth expansion module according to a specific embodiment of the present application is shown in FIG. Figure 3 As shown in the figure, the feature depth expansion module mainly undergoes a 3*3 convolution, a BN layer, and a ReLU activation function to make the back propagation of the network smoother.
[0043] Step S3: Send the output of step S2 to the multi-feature extraction module, Figure 4 FIG shows a framework diagram of a multi-feature extraction module according to a specific embodiment of the present application, as shown in FIG. Figure 4 As shown in the figure, the multi-feature extraction module is divided into two branches. The upper branch mainly consists of 3*1 and 1*3 convolutions to simulate the operation of 3*3 convolution, while reducing the number of parameters by 1 / 3.
[0044] Step S4: The lower branch of the multi-feature extraction module is composed of a 3x3 dilated convolution with a dilation ratio of 3 to expand the network's receptive field and obtain more contextual feature information. The upper and lower branches of the multi-feature extraction module are fused and then passed through a batch normalization layer and ReLU to prevent overfitting.
[0045] Step S5: The multi-feature extraction module is used 7 times, and the idea of residual network is introduced. The output of step S2 is connected to the output of the fourth and seventh multi-feature extraction modules respectively, so that the network can obtain more context information.
[0046] Step S6: Finally, the output features are passed through the pooling layer and the classification layer, and trained under the constraints of the cross entropy loss function.
[0047] This paper proposes a liveness detection method based on multi-scale feature aggregation, using a six-channel image synthesized from RGB and HSV as input. This method reduces network computational complexity, increases the network's nonlinear expression capability, and broadens the network's receptive field. It captures and aggregates features at different scales, thus obtaining more contextual information. The running speed is comparable to that of a method using only RGB images, achieving a good balance between performance and efficiency.
[0048] Continue to refer Figure 5 , Figure 5 The framework diagram of a liveness detection system with multi-scale feature aggregation according to an embodiment of the present invention is shown. The system specifically includes an image fusion unit 501, a feature depth expansion unit 502, and a training unit 503. The image fusion unit 501 is configured to convert an RGB image into an HSV image through image transformation, fuse the RGB image and the HSV image into an RGB-HSV image, and feed the image into the backbone network; the feature depth expansion unit 502 is configured to feed the features extracted by the backbone network into the feature depth expansion module, and feed the output into the multi-feature extraction module to obtain more context information; the training unit 503 is configured to pass the final output through a pooling layer and a classification layer, and train under the constraints of a cross-entropy loss function.
[0049] Reference below Figure 6 , which shows a structural diagram of a computer system suitable for implementing an electronic device of an embodiment of the present application. Figure 6 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0050] like Figure 6 As shown, the computer system includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a storage unit 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the system 600 are also stored in the RAM 603. The CPU 601, ROM 602, and RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0051] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, a mouse, and the like; an output section 607 including a liquid crystal display (LCD) and speakers; a storage section 608 including a hard disk; and a communication section 609 including a network interface card such as a LAN card or a modem. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. Removable media 611, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 610 as needed, so that computer programs read therefrom can be installed into the storage section 608 as needed.
[0052] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 609, and / or installed from the removable medium 611. When the computer program is executed by the central processing unit (CPU) 601, the above-mentioned functions defined in the method of the present application are executed. It should be noted that the computer-readable storage medium of the present application can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including, but not limited to, an electromagnetic signal, an optical signal, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable storage medium other than a computer-readable storage medium that can transmit, propagate, or transfer a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0053] Computer program code for performing the operations of the present application can be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In cases involving a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0054] The flow charts and block diagrams in the accompanying drawings illustrate the possible architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the module, program segment or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented by a dedicated hardware-based system that performs the specified function or operation, or can be implemented by a combination of dedicated hardware and computer instructions.
[0055] The modules described in the embodiments of the present application may be implemented by software or hardware.
[0056] As another aspect, the present application also provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment; or it may exist independently without being assembled into the electronic device. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device: converts the RGB image into an HSV image through image transformation, fuses the RGB image and the HSV image into an RGB-HSV image and feeds it into a backbone network; feeds the features extracted by the backbone network into a feature depth expansion module, and feeds the output into a multi-feature extraction module to obtain more contextual information; passes the final output through a pooling layer and a classification layer, and trains it under the constraints of a cross-entropy loss function.
[0057] The above description is merely a preferred embodiment of the present application and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in this application is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also encompasses other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned inventive concept. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this application.
Claims
1. A multi-scale feature aggregation method for liveness detection, characterized in that: include: S1: Convert the RGB image into an HSV image through image transformation, fuse the RGB image and the HSV image into an RGB-HSV image and send it to the backbone network; S2: The features extracted by the backbone network are fed into the feature depth expansion module, and the output is fed into the multi-feature extraction module to obtain more context information; S3: The final output is passed through the pooling layer and the classification layer, and trained under the constraints of the cross entropy loss function; The feature depth expansion module includes a 3*3 convolution, a BN layer and a ReLU activation function. The multi-feature extraction module includes two branches. The upper branch includes a simulated 3*3 convolution composed of 3*1 and 1*3 convolutions, and the lower branch includes a 3*3 void convolution with a void rate of 3. The upper and lower branches of the multi-feature extraction module are fused and then passed through the BN layer and ReLU to prevent overfitting. There are 7 multi-feature extraction modules. The output of the feature depth expansion module is connected to the output of the fourth and seventh multi-feature extraction modules respectively.
2. The method for liveness detection based on multi-scale feature aggregation according to claim 1, characterized in that: The backbone network can use ResNet18 or MobileNetV2.
3. The method for liveness detection based on multi-scale feature aggregation according to claim 1, characterized in that: The number of channels of the RGB-HSV image is 6, and the number of input channels of the first convolutional layer of the backbone network is 3.
4. A computer-readable storage medium having one or more computer programs stored thereon, characterized in that: When the one or more computer programs are executed by a computer processor, the method according to any one of claims 1 to 3 is implemented.
5. A multi-scale feature aggregation liveness detection system, characterized in that: The system comprises: An image fusion unit, configured to convert an RGB image into an HSV image through image transformation, fuse the RGB image and the HSV image into an RGB-HSV image, and send the image into a backbone network; A feature depth expansion unit is configured to feed the features extracted by the backbone network into a feature depth expansion module, and feed the output into a multi-feature extraction module to obtain more context information; Training unit: configured to pass the final output through the pooling layer and classification layer, and train under the constraints of the cross entropy loss function; The feature depth expansion module includes a 3*3 convolution, a BN layer and a ReLU activation function. The multi-feature extraction module includes two branches. The upper branch includes a simulated 3*3 convolution composed of 3*1 and 1*3 convolutions, and the lower branch includes a 3*3 void convolution with a void rate of 3. The upper and lower branches of the multi-feature extraction module are fused and then passed through the BN layer and ReLU to prevent overfitting. There are 7 multi-feature extraction modules. The output of the feature depth expansion module is connected to the output of the fourth and seventh multi-feature extraction modules respectively.
6. The multi-scale feature aggregation living body detection system according to claim 5, characterized in that: The backbone network can be ResNet18 or MobileNetV2, the number of channels of the RGB-HSV image is 6, and the number of input channels of the first convolutional layer of the backbone network is 3.
Citation Information
Patent Citations
Multi-feature and multi-model living body face recognition method
CN111160216A
Saliency target detection algorithm for aggregating dense and attention multi-scale features
CN114299305A