A hyperspectral-based live face detection system and method thereof
By simplifying the optical path structure and selecting appropriate wavelengths, and combining rainbow surface coding and single-class support vector machines, the problem of high computational complexity in hyperspectral imaging systems for live face detection is solved, achieving low-cost and efficient live face detection, and improving detection accuracy and user experience.
Patent Information
- Application Number
- CN202510065521.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2045-01-16
AI Technical Summary
Existing hyperspectral imaging systems suffer from high computational complexity and poor real-time performance in live face detection, and existing detection methods are not user-friendly. Improvements are needed to enhance the ease of use and accuracy of detection.
Design a hyperspectral-based live face detection system. Employ an objective lens, diffraction grating, focusing lens, mechanical slit, and sensor. Spectral modulation is achieved through rainbow surface coding, and feature discrimination is performed using a single-class support vector machine. The optical path structure is simplified, and a wavelength of 500nm±10nm is selected for detection. The cheek area is selected for detection, and classification is performed using gray-level difference-ratio features.
It achieves low-cost and efficient live face detection, reduces user occlusion and intrusion, improves the practicality and user experience of detection, reduces the impact of ambient light, and improves the accuracy and comfort of detection.
Smart Images

Figure CN119992622B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision technology, and in particular to a hyperspectral-based live human face detection system and method. Background Technology
[0002] In today's big data era, the diverse and redundant forms of identity verification are often unbearable for users. Facial recognition technology, however, targets the human face—a biometric feature that is difficult to forget and change—and offers advantages such as ease of operation, intuitive results, non-invasiveness, and contactlessness. Currently, in the field of liveness detection for human faces, spectral domain imaging forgery detection technology has made some progress.
[0003] With the continuous development of physiology and optics, researchers have discovered that hemoglobin in the dermis of human skin has the ability to absorb incident light, and its reflection characteristics exhibit unique properties that distinguish it from other objects as the wavelength changes. It stably shows a "W"-shaped trend at the 550nm wavelength, providing a theoretical basis for spectral imaging liveness detection. Based on this, Kim et al., through analysis of a large amount of facial spectral information, proposed using spectral data at 850nm and 685nm as feature bands to distinguish between real and fake faces and different ethnicities. First, they acquired image information of the forehead region in the selected bands. Then, they selected a fixed-size area and calculated its grayscale mean. The calculated results formed a two-dimensional feature vector for liveness detection. This method sacrifices system flexibility to obtain high-quality facial skin information. Although it achieved good experimental results, the mandatory exposure of the forehead region and the strictly fixed measurement distance led to a very unfriendly user experience. As an improvement, Zhang Zhiwei et al. ("Face liveness detection by learning multispectral reflectance distributions." 2011 IEEE International Conference on Automatic Face & Gesture Recognition (FG). IEEE, 2011.) set the active light source illuminating the face to 850nm and 1450nm LEDs, and used photosensitive devices to acquire the light reflected from the image. Finally, they used the values acquired by the photoelectric sensor as feature vectors to determine the face using the LDA method. Although this scheme improves human-computer interaction, the close acquisition magnification distance of 20-34cm inevitably leads to user aversion, reducing the naturalness and non-invasiveness of the acquisition.
[0004] To achieve live face detection using spectral information, hyperspectral information needs to be acquired. Among the existing solutions, SD-CASSI (Single Disperser-Coded Aperture Spectral Imaging) and DD-CASSI (Double Disperser-Coded Aperture Spectral Imaging) are relatively mature. However, these systems have large amounts of spectral data, high computational complexity, long computation time, and lack real-time performance, making them unsuitable for live face detection applications. Summary of the Invention
[0005] To address the various shortcomings in current practical applications of live face detection, this invention provides a hyperspectral-based live face detection system and method to further improve the simplicity and accuracy of spectral detection devices.
[0006] The technical solution adopted in this invention is as follows:
[0007] A hyperspectral-based live face detection system includes an objective lens, a diffraction grating, a focusing lens, a mechanical slit, a sensor, and a data processing device. The objective lens focuses the light reflected from the face onto the surface of the diffraction grating, which disperses the light into different wavelengths. The focusing lens converges the dispersed light onto the sensor. The mechanical slit, located between the focusing lens and the sensor, is used to filter the wavelengths. The sensor transmits the received information to the data processing device, which performs feature discrimination on the input information to obtain the detection result.
[0008] Furthermore, the mechanical slit is located on the R-plane, i.e. the rainbow plane. At the rainbow plane, scattered light rays with specific wavelengths emitted from all points in the scene will reach a specific area on the R-plane, thereby converting the wavelength into a spatial position range, which is then spectrally filtered and modulated by the mechanical slit.
[0009] Furthermore, the data processing device performs feature discrimination on the input information using a single-class support vector machine. The single-class support vector machine maps high-density regions in a single class of sample data to a high-dimensional space using a specific geometric shape. By finding the geometric shape that can surround most samples and has the smallest volume, the optimal support domain is obtained.
[0010] This invention also provides a focusing method for a hyperspectral-based live face detection system, the method comprising the following steps:
[0011] S1. First, focus the objective lens onto the sensor to determine the imaging distance.
[0012] S2, remove the sensor and place the diffraction grating at the sensor focus position determined in step S1. Then adjust the sensor angle and the distance from the diffraction grating to determine the dispersion direction of the first-order dispersion surface of the scene.
[0013] S3. Based on the dispersion direction determined in step S2, adjust the mechanical slit gate position to select the band, place a focusing lens behind the slit and adjust the position of the sensor.
[0014] S4, the sensor receives imaging information and transmits it to the data processing device, which then uses a feature recognition algorithm to determine the focusing result.
[0015] This invention also provides a detection method for a hyperspectral live face detection system. The method specifically includes: first, the detection system performs band filtering on the scene; after the sensor collects information, it extracts feature points from the data; the face region is selected as the feature extraction region; then, the gray-level difference-ratio feature is calculated; and the single-class support vector machine is used for classification; finally, the result is output.
[0016] The technical solution of the present invention has the following advantages over the prior art:
[0017] (1) The CASSI system was simplified according to the actual application scenario, and an efficient spectral image acquisition optical path was designed along the design concept of agile spectral imaging. The live face detection system of the present invention has low cost, simple optical path and strong practicality.
[0018] (2) Spectral images with a wavelength of 500nm±10nm were selected for detection, and their performance in identifying real and fake faces was superior to that of the bands near 550nm and 600nm.
[0019] (3) This invention selects the cheek area, which is less obscured in the case of facial recognition. Compared with the existing technology that forces detection in the forehead area, removing the obscuration of the cheek area is more convenient and easier, and less invasive. In addition, the skin on the cheek is thicker, which can more clearly reflect the spectral reflectance characteristics of the skin.
[0020] (4) In the prior art, the face recognition detection and measurement distance is strictly fixed and the acquisition distance is only 20-34cm. The method of the present invention proposes a specific area grayscale difference-ratio as face feature for classifier training based on the changes in distance and light in real scene. The obtained classifier is used to classify real and fake faces. The improved scheme can not only reduce the pollution of ambient light to CCD, but also ensure that the distance between the subject and the system is not excessively restricted, reducing the user's deliberate interaction and improving the user's comfort. Attached Figure Description
[0021] Figure 1This is a schematic diagram of the optical path principle of the system of the present invention;
[0022] Figure 2 This is a diagram showing the actual optical path setup of the system of the present invention;
[0023] Figure 3 The images are schematic diagrams of each optical path component (to demonstrate a more specific and obvious optical path spectral selection capability, all images were captured by a PointGrey Grasshopper3 GS3-U3-51S5C-C industrial color camera equipped with a Sony IMX250 sensor).
[0024] Figure 4 This is a schematic diagram of the detection structure of the system of the present invention;
[0025] Figure 5 The reflectance spectra of real human faces, human body models, and masks;
[0026] Figure 6 A schematic diagram of facial feature extraction locations. Detailed Implementation
[0027] The invention will be further described below with reference to the accompanying drawings.
[0028] The technical concept of this invention, a hyperspectral-based live face detection system, is as follows: By introducing a diffraction grating and a series of lenses into the optical path, a plane is generated in which light rays from all points in a scene of a specific wavelength intersect at a certain point. Based on the one-to-one correspondence between the wavelengths and spatial positions of the light rays on this plane, a mechanical slit is introduced to selectively block different wavelengths. Then, the light rays are refocused onto the camera's sensor through the lenses, achieving modulation of the spectrum of all points in the image through the slit to generate the final image.
[0029] The design principle of the system of this invention is as follows: Figure 1 As shown. Lens L1 precisely focuses the scene point X onto the diffraction grating plane P. Each image point X on the grating plane P... P Each ray corresponds to a cone of incident light with an angle θ. The diffraction grating disperses each ray into its constituent wavelengths. For each ray in the cone of incident light from a scene point, the grating creates a cone of outgoing light, each with a different wavelength and a dispersion angle α (determined by the type of diffraction grating). For ease of understanding, Figure 1 In the above example, θ is magnified; in the real system, θ << α. Under the focusing effect of lens L2, the sensor plane S and the diffraction grating plane P are conjugate, so the scene point is imaged onto the X-axis of the sensor plane. S Place.
[0030] In this imaging system, Figure 1 The R-plane (Rainbow plane) is a rainbow plane. Dispersed light rays with specific wavelengths emitted from all points in the scene will reach specific regions on the R-plane, thus converting the wavelength into a spatial location range. By placing slits in regions corresponding to specific wavelengths on this R-plane, bands can be filtered from the entire image spectrum to achieve hyperspectral modulation, i.e., high spatial resolution spectral imaging is achieved using rainbow plane coding.
[0031] pass Figure 1 Analysis of the case with finite apertures yields the following results:
[0032]
[0033] Where R θ Let θ be the width of the same wavelength in the rainbow plane on R, a1 be the aperture diameter of lens L1, f2 be the focal length of L2, and d and p be the distances from L1 to plane P and from plane P to L2, respectively. Clearly, the smaller R is, the less overlap there is between regions of different wavelengths in the R plane, resulting in more precise band selection and more accurate spectral modulation. Furthermore, considering that the type of diffraction grating P fixes the grating constant α, the system design requires θ1 << d. A small-aperture objective lens L1 can achieve this requirement.
[0034] The optical path structure of the face detection system built based on the above-mentioned finite aperture condition is as follows: Figure 2 As shown, in practical applications, this system uses an industrial monochrome camera CCD equipped with a sensor. To expand the field of view, the light from the scene is focused onto the grating surface by the industrial lens L1. In this embodiment, L1 is a 25mm industrial high-definition fixed-focus lens. The transmission diffraction grating is model THORLABS GT50-03, with 300 grooves / mm, a groove angle of 17.5°, and a size of 50mm x 50mm. The industrial lens L2 is a megapixel fixed-focus lens. Considering the ease of adjustment and manufacturing difficulty in actual operation, this embodiment uses a THORLABS VA100C / M30 mechanical slit to screen different wavelength spectra. A GCM-TD50MX rack and pinion translation stage is used to move the slit, and in conjunction with the adjustment of the slit size, spectral selection is achieved.
[0035] The focusing process during the setup of the optical path in this embodiment is as follows:
[0036] S1. First, focus the lens L1 onto the industrial camera's CCD to determine the distance from which the objective lens forms the image.
[0037] S2, remove the CCD and place the diffraction grating at the CCD focusing position determined in step S1. At this time, the light path is deflected due to the dispersion effect of the grating, so the subsequent light path needs to be deflected by a certain angle accordingly. Using an industrial camera equipped with a lens, the dispersion direction of the first-order dispersion plane of the scene is determined by adjusting the angle and the distance from the grating. Figure 3 As shown, this process should aim to make the dispersive bands as clear and non-overlapping as possible.
[0038] S3. Based on the dispersion direction determined in step S2, band selection is performed by fixing the slit using formulas 1 and 2. Lens L2 is placed behind the slit and the CCD position is adjusted. The dispersed image is converged by the eyepiece and projected onto the target surface of the grayscale camera's sensor. The image is received using SpinView professional software, and the focus result is determined through a feature recognition algorithm.
[0039] The structure of the system using this embodiment during detection is as follows: Figure 4 As shown, the face 101 to be detected is placed in front of the detection system. The fixed-focus lens 102 is used to focus the light reflected from the face onto the surface of the grating 103. The grating 103 disperses the collected light into different wavelengths. The pixel fixed-focus lens 104 collects the dispersed image onto the sensor target surface of the CCD camera 106. A mechanical slit 105 is placed between the pixel fixed-focus lens 104 and the CCD camera 106 to filter the wavelengths. The CCD camera 106 acts as a sensor to receive optical path information and transmits the information to the computer 107. The computer 107 performs data analysis on the collected spectral information and finally draws a recognition conclusion. The optical platform 108 is used to fix the above optical devices.
[0040] The overall system's live face detection process is as follows: First, the scene is directly filtered by the hardware optical path. After the CCD collects information, the open-source face_recogition library is used to extract feature points. The face region is selected as the feature extraction region. Then, the gray-level difference-ratio feature is calculated and classified by ONE-CLASS-SVM (single-class support vector machine). Finally, the results are output.
[0041] Figure 5The reflectance spectra of a real face, a PVC mannequin, and two silicone masks in the 380nm to 1100nm wavelength range were obtained using an Avantes AvaSpec-128 ultrafast fiber optic spectrometer. It can be more clearly observed that in the 470nm to 550nm and 700nm to 900nm wavelength ranges, the reflectance of both the silicone masks and the PVC mannequin differs consistently from that of human skin. Furthermore, although the skin reflectance is not high in the 470nm to 550nm range, the differences between the various forged materials and human skin are more significant. Therefore, this embodiment selects the 470nm to 550nm wavelength range for live face detection. However, in actual operation, due to limitations in the precision of optical components, the selection of the spectrum cannot be ideal. Therefore, a wavelength around 500nm is selected within the 470nm to 550nm range.
[0042] Considering that the spectral reflectance curves of various facial materials do not intersect in the selected wavelength band of 490nm-550nm, meaning that for the same light source, the concrete representation of the light reflected by different facial materials in the spectral dimension—the grayscale—will have a unique and definite magnitude relationship. Therefore, taking the average grayscale value of the cheek region in the 490nm-510nm wavelength band of the facial image as the original feature not only simplifies the experimental operation and reduces limitations and requirements, but also has good practicality.
[0043] In practical applications, the detection system of this embodiment is susceptible to contamination from the ambient light spectrum. Furthermore, different acquisition distances and exposure times can lead to varying CCD responses for the same reflectivity. Assuming the ambient light is a globally consistent light source during system acquisition, meaning it does not change with spatial location, the CCD response at face point a is I. a (λ i Then, the formula can be written as follows:
[0044]
[0045] Where λ i Represents the minimum wavelength λ imin up to the maximum wavelength λ imax Let λ be the wavelength of band i, H() be a step function, and c(λ), r(λ), and p(λ) be the image sensor response, the reflectivity of the object surface, and the global light source intensity, respectively, for wavelength λ. min and λ max These represent the minimum and maximum values of the wavelength λ range controlled by the mechanical slit. Considering the differences in reflectivity at different locations on the face due to variations in the proportion of skin and subcutaneous tissue, locations b and c are further selected so that the reflectivity at locations a, b, and c are all different. The following calculation is then performed to obtain the grayscale difference-ratio feature:
[0046]
[0047] Clearly, in the grayscale difference ratio formula, by adjusting I(λ) at different positions... i The subtraction and division method solves the problem of the influence of CCD response differences and ambient light differences on grayscale features, ensuring that the environment of the subject and the hard requirements such as acquisition distance and time are not excessively restricted, thus improving the comfort of use.
[0048] For the classifier, this embodiment chooses a single-class support vector machine (SVM). This is because in live face detection tasks, any non-face can be a negative sample, resulting in a significant imbalance between the positive and negative sample sizes. This causes the SVM's classification hyperplane to deviate, greatly reducing classification accuracy. Therefore, a single-class SVM is chosen because it maps high-density regions in a single class of sample data to a high-dimensional space using a specific geometry. By finding the geometry that encloses most samples and has the smallest volume, the optimal support region is obtained. For a given test sample, the optimal classification function is calculated to determine whether it belongs to that class. This approach achieves good classification accuracy even with limited samples and noise interference, making it suitable for tasks driven by live face detection.
[0049] In specific testing, such as Figure 6 As shown, it is necessary to reasonably select regions with different spectral reflectance characteristics for different objects. This embodiment considers that the skin thickness on the cheeks and bridge of the nose is inconsistent, resulting in different reflectance. Furthermore, the skin thickness and reflectance also differ between the areas near and far from the tip of the nose. Since the uniform material of a human face photograph and a silicone face does not cause significant changes in reflectance, four regions each on the left and right cheeks, and two regions from the center of the eyebrows to the tip of the nose, totaling 10 regions, are selected to form the feature extraction region. The average gray values of the feature extraction region on the left cheek are I1, I2, I3, and I4, and similarly, the average gray values of the feature extraction region on the right cheek are I5, I6, I7, and I8. The average gray values of the feature extraction region from the center of the eyebrows to the tip of the nose (i.e., the areas near and far from the tip of the nose) are I9, I1, I2, I3, and I4, respectively. 10 Further analysis revealed that I1 to I8 all belong to the cheek area, and their reflectivity can be considered consistent. I9 and I... 10 These are two regions with reflectivity inconsistent with the above, and their reflectivity is also inconsistent. Therefore, we traverse I1 to I8 and I9 and I... 10 Combining these elements to generate features:
[0050]
[0051] Finally, the input to the classifier is processed by normalization.
[0052] In setting the parameters of the classifier, experiments were conducted to compare and analyze the parameters of the currently popular REF kernel function, Linear kernel function, and Poly kernel function. The results showed that the Poly kernel function has better performance; its real face recognition rate decays more slowly, and the recognition rate of fake faces using both materials reached above 0.65 when the training error was 0.2, with a uniform recognition rate distribution. Therefore, the Poly kernel function was chosen for parameter setting.
[0053] In this embodiment, using the Poly kernel function and setting the training error to 0.27, the system achieves accuracy of 0.75 for recognizing real faces, 0.71 for recognizing face photographs, and 0.80 for recognizing silicone masks. Furthermore, testing shows that this system can acquire 18 frames per second using an industrial monochrome camera, achieving high-resolution scene sampling at 2016×2016 image resolution. Therefore, this system meets the requirements for live face detection. In summary, this invention demonstrates strong practicality in live face detection and possesses unique advantages compared to other current live face detection systems.
[0054] Based on the inconvenience of existing spectral imaging systems and considering that face liveness detection only requires examining feature values of a few bands and does not require reconstruction of the entire hyperspectral spectrum, this invention designs and builds a hyperspectral face detection system based on rainbow surface coding and support vector machines. This system skips the spectral reconstruction step and directly uses optical equipment to filter scene bands. Features are directly extracted from the grayscale images sampled by the two-dimensional image sensor to determine whether a face is real or fake, achieving a WYSIWYG (what you see is what you get) result.
Claims
1. A hyperspectral-based live human face detection system, comprising an objective lens, a diffraction grating, a focusing lens, a mechanical slit, a sensor, and a data processing device, characterized in that, The objective lens focuses the light reflected from the face onto the surface of the diffraction grating, which disperses the light into different wavelengths. The focusing lens converges the dispersed light onto the sensor. The mechanical slit is located between the focusing lens and the sensor to filter the wavelengths. The sensor transmits the received information to a data processing device, which performs feature discrimination on the input information and obtains the detection result. First, the detection system performs band filtering on the scene. After the sensor collects information, it extracts feature points from the data, selects the face region as the feature extraction region, calculates the gray-level difference-ratio feature, and then classifies it using a single-class support vector machine. Finally, the result is output. The formula for calculating the grayscale difference-ratio feature is as follows: Where, λ i Represents the minimum wavelength λ imin up to the maximum wavelength λ imax The wavelength of band i; λ min and λ max The minimum and maximum values of the wavelength λ range controlled by the mechanical slit; I a (λ i ), I b (λ i ) and I c (λ i ) represents the sensor response at different feature extraction regions a, b, and c of the face; H() is the step function; r a (λ), r b (λ) and r c (λ) is the surface reflectance at different feature extraction regions a, b, and c of the face, and these three reflectances are different from each other.
2. The hyperspectral-based live face detection system according to claim 1, characterized in that, The mechanical slit is located on the R-plane, i.e. the rainbow plane. At the rainbow plane, scattered light rays with specific wavelengths emitted from all points in the scene will reach a specific area on the R-plane, thereby converting the wavelength into a spatial position range, which is then spectrally filtered and modulated by the mechanical slit.
3. The hyperspectral-based live face detection system according to claim 1, characterized in that, The data processing device performs feature discrimination on the input information using a single-class support vector machine. The single-class support vector machine maps high-density regions in a single class of sample data to a high-dimensional space using a specific geometric shape. By finding the geometric shape that can surround most samples and has the smallest volume, the optimal support domain is obtained.
4. The hyperspectral-based live face detection system according to claim 1, characterized in that, The selected wavelength range is 490-510nm.
5. The hyperspectral-based live face detection system according to claim 1, characterized in that, The feature extraction region is specifically selected from the cheek area of the face.
6. The hyperspectral-based live face detection system according to claim 1, characterized in that, The feature extraction regions specifically selected were four regions on each of the left and right cheeks and two regions from the center of the eyebrows to the tip of the nose.
7. The hyperspectral-based live face detection system according to claim 1, characterized in that, Where c(λ) and p(λ) are the sensor response and global light source intensity for wavelength λ, respectively.
8. The focusing method of a hyperspectral-based live face detection system as described in claim 1, characterized in that, The method includes the following steps: S1. First, focus the objective lens onto the sensor to determine the imaging distance. S2, remove the sensor and place the diffraction grating at the sensor focus position determined in step S1. Then adjust the sensor angle and the distance from the diffraction grating to determine the dispersion direction of the first-order dispersion surface of the scene. S3. Based on the dispersion direction determined in step S2, adjust the mechanical slit gate position to select the band, place a focusing lens behind the slit and adjust the position of the sensor. S4, the sensor receives imaging information and transmits it to the data processing device, which then uses a feature recognition algorithm to determine the focusing result.
Citation Information
Patent Citations
Flow type imaging system based on spectral labeling method and optical frequency sweeping method
CN112858191A
Space-time hybrid modulation spectrum polarization imaging device and method based on grating-crystal
CN118464191A