Dynamic facial highlight material extraction algorithm based on time-varying light field of lightweight device
By combining lightweight devices and algorithms, utilizing the facial skin reflection model parameters, and designing the light source and camera position, we solved the problems of high hardware cost and large differences in lighting patterns in extracting highlight materials in dynamic scenes, and achieved efficient and low-cost extraction of dynamic facial highlight materials.
Patent Information
- Application Number
- CN202310212243.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-01
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2043-03-01
AI Technical Summary
Existing technologies make it difficult to effectively use lightweight devices to extract the highlight texture of the face in dynamic scenes, especially when the lighting patterns vary greatly. This leads to high hardware costs, difficulty in construction and maintenance, and is not suitable for the temporal correlation of dynamic scenes.
A time-varying light field based on a lightweight device is used, combined with the Torrance-Sparrow model and Rusinkiewicz Half-vector parameterization. Through the design of a camera array and a small number of light sources, the highlight material of the face is extracted. The parameters of the facial skin reflection model are used to design the light source and camera position to achieve the extraction of the highlight component.
It achieves efficient extraction of facial highlight materials in dynamic scenes, reduces hardware cost and complexity, adapts to the timing correlation of dynamic scenes, and only requires a small number of light source modes and lighting modes to provide uniform lighting effects.
Smart Images

Figure CN116246006B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of dynamic PBR material reconstruction of human faces, and in particular to an algorithm for extracting dynamic human face highlight materials based on a time-varying light field using a lightweight device. Background Art
[0002] In fields such as augmented reality (AR), virtual reality (VR), the metaverse, and virtual photography, there is a significant demand for 3D reconstruction of the human body, with the human face being a particular focus. Currently, a common method involves using multi-view cameras and a camera array to capture 3D facial information. In addition to 3D information, facial skin appearance is also captured. Based on its reflective properties, facial skin appearance is divided into diffuse and specular. Diffuse reflectance represents skin tone and is a low-frequency signal, while specularity reflects details down to the pore level and is a high-frequency signal. Specularity is a key element in achieving highly realistic facial features. To capture only diffuse reflectance data, the capture device typically uses a uniform, brightly lit environment. To capture high-frequency specular data, the capture device uses a time-varying light source to capture the face. A representative example of this type of device is the Light Stage series, featuring spherical gradient polarized light. Light Stage requires a spherical capture device with hundreds of distributed light sources to capture the effects of four different lighting modes on the face. Photometric stereo technology is then used to extract the physical reflective texture of the face. However, when extended to the dynamic field, the Light Stage device is not conducive to the temporal correlation of dynamic scenes due to the large differences between different lighting modes.
[0003] Therefore, technical personnel in this field are committed to developing a dynamic facial highlight material extraction algorithm based on a time-varying light field of a lightweight device, with the following requirements: a. The lighting modes should be as few as possible and illuminate the face as evenly as possible to adapt to dynamic scenes; b. The algorithm and device should be highly coupled, fully considering the characteristics of the facial skin reflection model itself; c. The device should be designed to be lightweight to reduce hardware costs, construction and maintenance difficulties. Summary of the Invention
[0004] To overcome the above problems, the present invention provides a dynamic facial highlight material extraction algorithm based on the time-varying light field of a lightweight device. In conjunction with the time-varying light field of the lightweight device, on the basis of fully analyzing the parameters of the facial skin reflection model, the algorithm uses as few light source modes as possible and as homogeneous ambient light as possible to extract the highlight component of the dynamic skin material without affecting the temporal correlation.
[0005] A first aspect of the present invention provides a dynamic facial highlight material extraction algorithm based on a time-varying light field using a lightweight device, comprising the following steps:
[0006] (1) According to the open-source Merl / ETH Skin database of facial skin reflectance parameters, facial highlights are expressed using the Torrance-Sparrow model, with a roughness parameter with a mean of 0.3031 and a standard deviation of 0.0891. This parameter has very little variation among individuals and can be considered a constant. Therefore, the changes in the highlight part are fully reflected in the variable bumpnormal.
[0007] (2) The Torrance-Sparrow skin model is converted into a Rusinkiewicz Half-vector parameterized expression, where the Half-vector is the angle between the camera's line of sight and the directional light. When the Half-vector is determined and used as the Z-axis, the reflectance of the face decreases with increasing altitude and is independent of the azimuth.
[0008] (3) After converting the parameterized expression, the component intensities of the highlights under three different half-vectors are obtained. By querying the fixed half-vector and the roughness and the calculated reflection model parameters, the angle between the normal and the half-vector is obtained; the normal is solved by at least three angles with different half-vectors;
[0009] (4) After obtaining the normal, the Torrance-sparrow model is used to calculate the reflection intensity of the highlight.
[0010] Furthermore, in step 1, based on the open source facial skin reflectance parameter statistical database Merl / ETH Skin, facial highlights are expressed using the Torrance-Sparrow model as follows:
[0011]
[0012] Where f is the reflection coefficient, ρ s is the weight of the highlight component, G is the geometric attenuation factor, D is the microsurface distribution factor, F r is the Fresnel factor; G, D, F r The calculation formula is as follows:
[0013]
[0014]
[0015]
[0016] Among them, m is the roughness, δ is the angle between the surface normal N and the half-vector H, and R0 is the light intensity reflectivity of the face highlight material at vertical incidence.
[0017] Furthermore, in step 2, the Torrance-Sparrow reflection model is converted according to the parameterization method of Rusinkiewicz Half-vector; for the constructed device, the directional light direction ω i and the camera sight normal ω o is known, θ d is H and ω i The angle between d is known;
[0018] θ h is the angle between the normal and H, Yes i The projection angle on the plane perpendicular to h; θ h and The value ranges of these two parameters are as follows:
[0019]
[0020] Furthermore, the θ h and The value ranges of these two parameters are not completely traversable, and combinations with angles greater than 90° with the normal N need to be removed.
[0021] Furthermore, in step 3, by pre-fixing the roughness parameter and θ d For each highlight measurement, we can find its corresponding Rusinkiewicz Half-vector parameterized Torrance-Sparrow reflectance model, and accordingly find its elevation angle with H. We find three elevation angles relative to different H and use the following equations to calculate the normal parameters of the face:
[0022]
[0023] Among them, each row of L is the direction of the normalized directional light, n is the normal to be determined, and s is the inner product of the normal obtained by the above table lookup and the half-vector.
[0024] Furthermore, in step 4, the known normal is substituted into the Torrance-Sparrow reflection model to fit the high light intensity parameter ρ s .
[0025] A second aspect of the present invention provides a face acquisition device based on a dynamic face highlight material extraction algorithm of a time-varying light field of a lightweight device, comprising a camera array, a side camera, a front camera, a softbox, a side directional light source, a front directional light source, and a light source;
[0026] The camera array includes multiple rows and columns of cameras arranged in an array to record different angles of the face. The camera array is arranged in a semicircular shape when viewed from above and is one meter away from the face. The camera array covers the front part of the face from the left ear to the right ear. The camera array is divided into four layers in the vertical direction, covering the range from the neck to the chin to the forehead of the face.
[0027] The soft box is arranged around the rear side of the camera array, and multiple light sources are arranged in the soft box, and the multiple light sources form uniform ambient lighting; the camera array and the directional light source are placed in front of the soft box, one meter away from the face of the person;
[0028] The side directional light sources are provided with three on each side of the face, and the three side directional light sources are spaced apart in the vertical direction to cover the height of the side face; the front directional light sources are provided with two in front of the face, and the two front directional light sources are spaced apart in the vertical direction to cover the height of the front face;
[0029] The side cameras are arranged in three on each side of the face, and the three side cameras are arranged at intervals in the vertical direction; wherein the two cameras located at the top and bottom of the three side cameras are at the same position as the directional light sources located at the top and bottom of the three side directional light sources; the side camera located in the middle of the three side cameras is arranged opposite to the directional light source located in the middle of the three side directional light sources, and all the side cameras and the side directional light sources form a diamond shape;
[0030] There are three frontal cameras set in front of the face. The three frontal cameras are set at intervals in the vertical and horizontal directions, and are 1m away from the captured face. From the front view, the three frontal cameras form a triangular distribution, so that the captured image is sufficient to cover the front area of the face.
[0031] Furthermore, the three side directional light sources are arranged at intervals of 20 to 30 cm in the vertical direction. Among the three side directional light sources, there is no horizontal distance between the side directional light source located at the top and the side directional light source located at the bottom, and the side directional light source located in the middle is located at the edge of the camera array and has a horizontal distance of 15 to 20 cm from the side directional light source located at the top and the side directional light source located at the bottom.
[0032] Furthermore, two front-directional light sources are pointed at the collected face, with a relative height of no less than 20 cm in the vertical direction; the placement of the front-directional light sources avoids direct exposure to the human eye area to reduce discomfort during the collection process.
[0033] The beneficial effects of the present invention are: the number of cameras and light sources required is less than that of other solutions, and it is a lightweight solution; the designed lighting mode can achieve uniform lighting in time sequence without affecting it, and only affects the highlight components representing details and high-frequency signals, achieving an effect that previously required large and expensive light stage devices; it fully utilizes the reflection model of the human face skin material, and the extracted normal has a physical basis; no more than three lighting modes can be used to extract highlight parameters, with a small time interval, and is suitable for dynamic facial scenes. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 Schematic diagram of the placement of the side camera and side directional light source in the present invention.
[0035] Figure 2 Schematic diagram of the placement of the front camera and the front directional light source in the present invention.
[0036] Figure 3 is θ d =50° as an example, the Torrance-Sparrow visualization of the Rusinkiewicz Half-vector parameterization.
[0037] Figure 4 Flowchart of the algorithm of the present invention.
[0038] Figure 5 This is the third embodiment of the light field array of the present invention. The three groups of light sources formed by the three blocks are independent from each other and are turned on and off in sequence within the groups.
[0039] Figure 6 The invention discloses a flow chart for obtaining an input image sequence of a face acquisition device. DETAILED DESCRIPTION
[0040] The following will clearly and completely describe the technical solution of the present invention in conjunction with the accompanying drawings. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of them. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0041] In the description of the present invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer" and the like, indicating orientations or positional relationships, are based on the orientations or positional relationships shown in the accompanying drawings and are intended solely to facilitate and simplify the description of the present invention. They are not intended to indicate or imply that the devices or components referred to must have, be constructed, or operate in a specific orientation, and are therefore not to be construed as limitations on the present invention. Furthermore, the terms "first," "second," and "third" are used for descriptive purposes only and are not to be construed as indicating or implying relative importance.
[0042] In the description of the present invention, it should be noted that, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention in specific contexts.
[0043] Example 1
[0044] Refer to the attached Figure 4 The dynamic face highlight material extraction algorithm based on the time-varying light field of a lightweight device includes the following steps:
[0045] (1) Based on the open-source Merl / ETH Skin database of facial skin reflectance parameters, facial highlights are expressed using the Torrance-Sparrow model as follows:
[0046]
[0047] Where f is the reflection coefficient, ρ s is the weight of the highlight component, G is the geometric attenuation factor, D is the microsurface distribution factor, F r is the Fresnel factor; G, D, F r The calculation formula is as follows:
[0048]
[0049]
[0050]
[0051] Among them, m is the roughness, δ is the angle between the surface normal N and the half-vector H, and R0 is the light intensity reflectivity of the face highlight material at vertical incidence.
[0052] (2) According to the parameterization of Rusinkiewicz Half-vector, the above Torrance-Sparrow reflection model is transformed. For the built device, the directional light direction ω i and the camera sight normal ω o is known, so θ d (H and ω i The angle between the two is known, so the remaining parameters become θ h (the angle between the normal and H), and ω i The projection angle on the plane perpendicular to h The value ranges of these two parameters are as follows:
[0053]
[0054] The above range cannot be completely traversed, and it is also necessary to remove some combinations with an angle of more than 90° with the normal N. Figure 3 The black area with countless values of the bow below.
[0055] With θ d =50° as an example, the Rusinkiewicz Half-vector parameterized Torrance-Sparrow is visualized as Figure 3 , the vertical direction represents θ h , the horizontal direction is The arc below indicates that the value range exceeds the visible range and the reflection parameters are not calculated. According to the isotropic property, the reflection coefficient intensity is only related to θ h related to Not relevant.
[0056] (3) According to the statistical parameter of the Merl / ETH Skin database, 0.3031, the standard deviation is 0.0891, that is, the difference between different people is very small and can be regarded as a constant. Therefore, the change of the highlight part can be fully reflected in the normal (bump normal). By pre-fixing the roughness parameter and θ d For each highlight measurement, we can find its corresponding value in the corresponding Rusinkiewicz Half-vector parameterized Torrance-Sparrow reflectance model, and accordingly find its altitude angle with H. As long as we can find three altitude angles relative to different H, we can solve the following equations to calculate the normal parameters of the face:
[0057]
[0058] Among them, each row of L is the direction of the normalized directional light, n is the normal to be determined, and s is the inner product of the normal obtained by the above table lookup and the half-vector.
[0059] (4) Substitute the known normal into the Torrance-Sparrow reflection model and fit the high light intensity parameter ρ s .
[0060] A dynamic facial texture extraction algorithm based on the time-varying light field of a lightweight device. The algorithm is coupled with the hardware, inputs a multi-view and multi-light highlight image sequence from a face acquisition device, and outputs the face's normal (Bump Normal) layer and highlight intensity.
[0061] Example 2
[0062] Refer to the attached Figure 1-2 , a face acquisition device that implements a dynamic face highlight material extraction algorithm based on a time-varying light field of a lightweight device. According to the algorithm principle, the face acquisition device needs to: a. Try to ensure that three cameras are used to capture highlights at the same time, and the halfvectors of the cameras and light sources need not be collinear; b. The directional light distribution needs to cover the entire face. Because the face has certain fluctuations, more than one light source may be required in a local area, but the number of light sources should be kept within three to adapt to dynamic scenes; c. To create relatively uniform ambient lighting, a uniform ambient lighting with a lower brightness than the directional light is required. The specific implementation of the device is as follows:
[0063] (1) The acquisition device is distributed in a semicircular shape when viewed from above, about one meter away from the face; the camera covers the front part within the range of the left and right ears; the camera array is divided into four layers in the vertical direction, covering the range from the neck, chin to the forehead of the face.
[0064] (2) The device is surrounded by a softbox, with some light sources placed behind the softbox to create a uniform ambient light. The camera and directional light source are placed in front of the softbox, about one meter away from the face. Three directional light sources are placed on each side of the face, covering the height of the side face. The three light sources are spaced about 30 cm apart in the vertical direction, with no horizontal distance between the upper and lower light sources. The middle light source is located at the edge of the device, with a horizontal distance of about 20 cm between the upper and lower light sources. The front part has two directional lights, one above and one below, covering the height of the front face, with a height difference of about 80 cm.
[0065] (3) Three side cameras are arranged on each side, of which the upper and lower side cameras are at the same position as the upper and lower side direction light sources, and the middle side camera is opposite to the middle side direction light source. All side cameras and side direction light sources form a diamond shape; three frontal cameras are placed on the front, one of which is located 1m in front of the face and about 20cm high; the remaining two frontal cameras are located on the left and right sides, with a horizontal distance of about 50cm. The rest are at the same height as the nose of the face, and two are about 20cm lower than the nose.
[0066] Example 3
[0067] Refer to the attached Figure 5 , as in the face acquisition device described in the second embodiment, the input image sequence acquired is obtained as follows:
[0068] Step 1: Turn on and off the upper, middle and lower light sources on the two sides in sequence, and the camera shoots synchronously;
[0069] Step 2: Turn on and off the upper and lower light sources in sequence, and the camera takes pictures simultaneously;
[0070] Step 3: Adjust the ambient light intensity from strong to weak, and the camera will shoot simultaneously;
[0071] The above steps are independent of each other and are not performed in any particular order.
[0072] The multi-view and multi-illumination image sequences input to the algorithm of the present invention are collected from a lightweight face acquisition device. The key time-varying light field design of the device is designed based on a thorough analysis of the face physics-based reflection (PBR) model and is closely integrated with the algorithm. The device only requires a uniform ambient light and several directional lights to achieve dynamic face PBR material extraction. Compared with other face material extraction directions, the present invention can achieve the following: a. The number of time-varying illumination modes is small, and the overall uniform illumination is conducive to time-series correlation; the local directionality is conducive to the extraction of the highlight component of the PBR material; c. The algorithm fully utilizes the statistical parameters of the face skin reflection model as a priori conditions, thereby designing the relative position and illumination mode of the light source and the camera, and efficiently restoring the PBR highlight parameters required for the face; c. The light field required by the device is simple to build and has low cost.
[0073] The contents described in the embodiments of this specification are merely an enumeration of the implementation forms of the inventive concept. The scope of protection of the present invention should not be regarded as limited to the specific forms described in the embodiments. The scope of protection of the present invention also extends to equivalent technical means that can be conceived by those skilled in the art based on the inventive concept.
Claims
1. A dynamic facial highlight material extraction algorithm based on a time-varying light field of a lightweight device, characterized by: The following steps are involved: (1) According to the open-source Merl / ETH Skin database of facial skin reflectance parameters, facial highlights are expressed using the Torrance-Sparrow model, with a roughness parameter mean of 0.3031 and a standard deviation of 0.0891. This parameter has very little variation between individuals and can be considered a constant. Therefore, the changes in the highlight part are fully reflected in the variable "bump normal"; (2) The Torrance-Sparrow skin model is converted into a Rusinkiewicz Half-vector parameterized expression, where the Half-vector is the angle between the camera's line of sight and the directional light. When the Half-vector is determined and used as the Z-axis, the reflectance of the face decreases with increasing altitude and is independent of the azimuth. (3) After converting the parameterized expression, the component intensities of the highlights under three different half-vectors are obtained. By querying the fixed half-vector and the roughness and the calculated reflection model parameters, the angle between the normal and the half-vector is obtained; the normal is solved by at least three angles with different half-vectors; (4) After obtaining the normal, the Torrance-sparrow model is used to calculate the reflection intensity of the highlight.
2. The dynamic facial highlight texture extraction algorithm based on a time-varying light field using a lightweight device as claimed in claim 1, characterized in that: In step 1, based on the open-source facial skin reflectance parameter statistical database Merl / ETH Skin, facial highlights are expressed using the Torrance-Sparrow model as follows: Where f is the reflection coefficient, ρ s is the weight of the highlight component, G is the geometric attenuation factor, D is the microsurface distribution factor, F r is the Fresnel factor; G, D, F r The calculation formula is as follows: Among them, m is the roughness, δ is the angle between the surface normal N and the half-vector H, and R0 is the light intensity reflectivity of the face highlight material at vertical incidence.
3. The dynamic facial highlight texture extraction algorithm based on a time-varying light field using a lightweight device as claimed in claim 1, characterized in that: In step 2, the Torrance-Sparrow reflection model is converted according to the parameterization method of Rusinkiewicz Half-vector; for the built device, the directional light direction ω i and the camera sight normal ω o is known, θ d is H and ω i The angle between d is known; θ h is the angle between the normal and H, Yes i The projection angle on the plane perpendicular to h; θ h and The value ranges of these two parameters are as follows:
4. The dynamic facial highlight texture extraction algorithm based on a time-varying light field using a lightweight device as claimed in claim 3, characterized in that: The θ h and The value ranges of these two parameters are not completely traversable, and combinations with angles greater than 90° with the normal N need to be removed.
5. The dynamic facial highlight texture extraction algorithm based on a time-varying light field using a lightweight device as claimed in claim 1, characterized in that: In step 3, by pre-fixing the roughness parameter and θ d For each highlight measurement, we can find its corresponding Rusinkiewicz Half-vector parameterized Torrance-Sparrow reflectance model, and accordingly find its elevation angle with H. We find three elevation angles relative to different H and use the following equations to calculate the normal parameters of the face: Among them, each row of L is the direction of the normalized directional light, n is the normal to be determined, and s is the inner product of the normal obtained by the above table lookup and the half-vector.
6. The dynamic facial highlight texture extraction algorithm based on a time-varying light field using a lightweight device as claimed in claim 1, characterized in that: In step 4, the known normal is substituted into the Torrance-Sparrow reflection model to fit the high light intensity parameter ρ s .
7. A face acquisition device implementing the dynamic face highlight texture extraction algorithm based on a time-varying light field of a lightweight device according to any one of claims 1 to 6, characterized in that: Includes camera array, side camera, front camera, softbox, side directional light source, front directional light source and light source; The camera array includes multiple rows and columns of cameras arranged in an array to record different angles of the face. The camera array is arranged in a semicircular shape when viewed from above and is one meter away from the face. The camera array covers the front part of the face from the left ear to the right ear. The camera array is divided into four layers in the vertical direction, covering the range from the neck to the chin to the forehead of the face. The soft box is arranged around the rear side of the camera array, and multiple light sources are arranged in the soft box, and the multiple light sources form uniform ambient lighting; the camera array and the directional light source are placed in front of the soft box, one meter away from the face of the person; The side directional light sources are provided with three on each side of the face, and the three side directional light sources are spaced apart in the vertical direction to cover the height of the side face; the front directional light sources are provided with two in front of the face, and the two front directional light sources are spaced apart in the vertical direction to cover the height of the front face; The side cameras are arranged in three on each side of the face, and the three side cameras are arranged at intervals in the vertical direction; wherein the two cameras located at the top and bottom of the three side cameras are at the same position as the directional light sources located at the top and bottom of the three side directional light sources; the side camera located in the middle of the three side cameras is arranged opposite to the directional light source located in the middle of the three side directional light sources, and all the side cameras and the side directional light sources form a diamond shape; There are three frontal cameras set in front of the face. The three frontal cameras are set at intervals in the vertical and horizontal directions, and the distance from the captured face is 1m. From the front view, the three frontal cameras form a triangular distribution, so that the captured image is sufficient to cover the front area of the face.
8. The face acquisition device according to claim 7, wherein: The three side directional light sources are arranged at intervals of 20 to 30 cm in the vertical direction. Among the three side directional light sources, there is no horizontal distance between the side directional light source located at the top and the side directional light source located at the bottom. The side directional light source located in the middle is located at the edge of the camera array and has a horizontal distance of 15 to 20 cm from the side directional light source located at the top and the side directional light source located at the bottom.
9. The face acquisition device according to claim 7, wherein: Two front-facing directional light sources are pointed at the collected face, with a relative height of no less than 20 cm in the vertical direction; the front-facing directional light sources are placed to avoid direct exposure to the human eye area to reduce discomfort during the collection process.
Citation Information
Patent Citations
Face skin material calculation method, device and equipment based on face three-dimensional model
CN112149578A
Human face liveness detection method based on polarization imaging
WO2021217764A1