Lens fitting head modeling method, system, device and medium based on multi-view images
Patent Information
- Application Number
- CN202610901433.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-09-11
AI Technical Summary
但在实际部署中仍存在明显局限:渲染算力负载较高,平台兼容性较差,且生成的模型本质为离散点云而非连续实体网格表面,无法支持眼镜试戴所需的物理碰撞检测,导致虚拟试戴的适配性与交互稳定性难以满足实际应用对碰撞精度与可靠性的要求
[0015]This invention relates to a method, system, electronic device, and computer-readable storage medium for head modeling based on multi-view image sequence analysis of head views and corresponding depth maps; extraction of 2D facial key points, combined with depth maps and camera intrinsic parameters for back-projection to obtain 3D key points; multi-stage optimization of a preset head model based on 3D key points and intrinsic/extrinsic parameters to obtain a 3D head model; generation of standard texture maps using texture space, optimized camera parameters, and various views; identification of the scalp region, registration of the hairstyle point cloud, and construction of an integrated hairstyle model with the head model and texture map. Therefore, this invention proposes a multi-view image-based method for facial modeling in eyeglass fitting, which, through the construction of a multi-view analysis and step-by-step optimization architecture for the preset head model, enables high-fidelity head geometry and texture reconstruction using a regular camera without relying on dedicated scanning equipment. This solves the problems of inaccurate dimensions in monocular reconstruction, gaps in multi-view texture seams, and poor hairstyle fit, improving the geometric realism and fitting accuracy of virtual try-on in intelligent eyeglass fitting scenarios.
Smart Images

Figure CN122737367A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of optometry technology, and in particular to a method, system, device and medium for lens modeling based on multi-view images. Background Technology
[0002] Traditional 3D head modeling solutions often employ dedicated scanning equipment or monocular RGB reconstruction methods. Dedicated scanning equipment relies on structured light or depth cameras, which can acquire high-precision depth information, but the hardware costs are high, and the reconstructed model is a static mesh, making it impossible to adjust facial pose and expression. While monocular video stream or single RGB image reconstruction has a low acquisition threshold, video methods are computationally intensive and time-consuming, making it difficult to meet the needs of real-time try-on. Single image methods have low accuracy in facial structure dimensions, and neither supports editable facial texture maps and hair geometry, limiting overall modeling accuracy and interactive flexibility.
[0003] In the current field of smart eyeglass fitting, to achieve high-fidelity head modeling, some solutions have shifted to simulated face reconstruction based on 3D Gaussian splashing, utilizing discrete point clouds to achieve photorealistic textures and lighting effects. However, significant limitations remain in practical deployments: high rendering computational load, poor platform compatibility, and the generated model is essentially a discrete point cloud rather than a continuous solid mesh surface, making it unable to support the physical collision detection required for eyeglass try-on. This results in the virtual try-on's adaptability and interaction stability failing to meet the collision accuracy and reliability requirements of real-world applications. Summary of the Invention
[0004] This invention provides a method, system, device, and medium for lens fitting facial modeling based on multi-view images, with the main purpose of improving the fitting accuracy of virtual try-on in intelligent lens fitting scenarios.
[0005] To achieve the above objectives, the present invention provides a lens facial modeling method based on multi-view images, comprising: The head is captured by a camera with multiple viewpoint image sequences. The multiple viewpoint image sequences are parsed to obtain multiple head views. Multiple original depth maps of the multiple head views are obtained, wherein each head view corresponds to one original depth map. Two-dimensional facial key points are obtained from multiple head views, as well as the camera intrinsics of the camera. Based on multiple original depth maps and the camera intrinsics, the two-dimensional facial key points are back-projected onto a preset three-dimensional camera coordinate system to obtain three-dimensional facial key points. Based on the three-dimensional facial key points, the camera intrinsic parameters, and the preset camera extrinsic parameters, a multi-stage distribution optimization is performed on the preset head model to obtain a 3D head model. Obtain the predefined texture space of the preset head model and the optimization parameters of the camera, and generate a standard head texture map based on the texture space, multiple head views, the 3D head model and the optimization parameters; Identify the scalp region of the 3D head model, register the preset hairstyle point cloud based on the scalp region to obtain a registered hairstyle model, and construct an integrated head hairstyle model based on the registered hairstyle model, the 3D head model and the standard head texture map.
[0006] Optionally, parsing the multi-view image sequence to obtain multiple head views includes: Camera pose estimation is performed based on the multi-view image sequence to obtain the initial pose data of the camera; The head deflection angle is calculated based on the initial posture data, and the multi-view image sequence is filtered based on the head deflection angle and a preset viewpoint classification threshold to obtain image frames that meet the viewpoint requirements. The image frames that meet the viewing angle requirements are calibrated and categorized to obtain multiple head views.
[0007] Optionally, the step of performing multi-stage distribution optimization on the preset head model based on the three-dimensional facial key points, the camera intrinsic parameters, and preset camera extrinsic parameters to obtain a 3D head model includes: Obtain an initial geometric function, construct an ear-specific loss term based on preset ear key points, and construct a combined geometric function based on the ear-specific loss term and the initial geometric function; The ear-specific loss term in the combined geometric function is adjusted to obtain an optimized geometric function; Based on the optimized geometric function, the 3D facial key points, the camera intrinsic parameters, and the preset camera extrinsic parameters, the preset head model is subjected to multi-stage distribution optimization to obtain a 3D head model.
[0008] Optionally, adjusting the ear-specific loss term in the combined geometric function to obtain an optimized geometric function includes: Obtain the current image frame in the multi-view image sequence, and determine the spatial location of the key points of the ear; Calculate the head deflection angle corresponding to the current image frame to obtain the current image viewpoint, and calculate the visibility of the ear region based on the current image viewpoint and the spatial position; Based on the current image viewpoint and the visibility of the ear region, calculate the adaptive weight coefficients corresponding to the ear-specific loss term; Based on the adaptive weighting coefficients, the constraint strength of the ear-specific loss term in the combined geometric function is dynamically adjusted to obtain the optimized geometric function.
[0009] Optionally, generating a standard head texture map based on the texture space, multiple head views, the 3D head model, and the optimization parameters includes: Obtain the three-dimensional mesh points of the 3D head model, and the pixel values of each pixel point in the multiple head views; Based on the optimization parameters, the mapping relationship between grid points and pixels is found. Based on the mapping relationship, the pixel values of each pixel point are mapped to the texture space to obtain the initial texture map. The initial texture map is fused from multiple perspectives to obtain a standard head texture map.
[0010] Optionally, the step of performing multi-view fusion on the initial texture map to obtain a standard head texture map includes: Obtain the normal vector of each facet in the 3D head model, and calculate the angle value corresponding to each facet based on the facet normal vector and the gaze vector of the optimization parameters; The texture fusion weight of each facet is calculated based on the included angle value corresponding to each facet, and the texture fusion weight of each facet is mapped to the texture space to generate a texture weight map with multiple perspectives. An intermediate texture map is generated based on the texture weight map from the multiple perspectives and the initial texture map; Obtain the occlusion area of the intermediate texture map, perform texture synthesis and hole filling on the occlusion area to obtain the standard head texture map.
[0011] Optionally, the step of registering the preset hairstyle point cloud based on the scalp region to obtain a registered hairstyle model includes: Obtain the scalp region point cloud of the scalp area, and calculate the root mean square radius ratio of the preset hairstyle point cloud and the scalp region point cloud; Based on the root mean square radius ratio, the preset hairstyle point cloud is initially scaled to obtain an initial scaled hairstyle point cloud; A rigid transformation is performed on the initial scaled hairstyle point cloud to obtain the registered hairstyle model.
[0012] To address the aforementioned problems, the present invention also provides a lens appearance modeling system based on multi-view images, the system comprising: The image processing module is used to acquire a multi-view image sequence of the head based on the camera, parse the multi-view image sequence to obtain multiple head views, and generate multiple original depth maps of the multiple head views, with one head view corresponding to one original depth map. Two-dimensional facial key points are obtained from multiple head views, as well as the camera intrinsics of the camera. Based on multiple original depth maps and the camera intrinsics, the two-dimensional facial key points are back-projected onto a preset three-dimensional camera coordinate system to obtain three-dimensional facial key points. The head model calculation module is used to perform multi-stage distribution optimization on the preset head model based on the three-dimensional facial key points, the camera intrinsic parameters and the preset camera extrinsic parameters to obtain a 3D head model. The texture synthesis module is used to obtain the predefined texture space of the preset head model and the optimization parameters of the camera, and generate a standard head texture map based on the texture space, multiple head views, the 3D head model and the optimization parameters; The model unification module is used to identify the scalp region of the 3D head model, register the preset hairstyle point cloud based on the scalp region to obtain the registered hairstyle model, and construct an integrated head hairstyle model based on the registered hairstyle model, the 3D head model and the standard head texture map.
[0013] To address the above problems, the present invention also provides an electronic device, the electronic device comprising: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, which enables the at least one processor to perform the lens part modeling method based on multi-view images as described above.
[0014] To address the aforementioned problems, the present invention also provides a computer-readable storage medium, including a data storage area and a program storage area, wherein the data storage area stores created data and the program storage area stores a computer program; wherein, when the computer program is executed by a processor, it implements the lens facial modeling method based on multi-view images as described above.
[0015] This invention relates to a method, system, electronic device, and computer-readable storage medium for head modeling based on multi-view image sequence analysis of head views and corresponding depth maps; extraction of 2D facial key points, combined with depth maps and camera intrinsic parameters for back-projection to obtain 3D key points; multi-stage optimization of a preset head model based on 3D key points and intrinsic / extrinsic parameters to obtain a 3D head model; generation of standard texture maps using texture space, optimized camera parameters, and various views; identification of the scalp region, registration of the hairstyle point cloud, and construction of an integrated hairstyle model with the head model and texture map. Therefore, this invention proposes a multi-view image-based method for facial modeling in eyeglass fitting, which, through the construction of a multi-view analysis and step-by-step optimization architecture for the preset head model, enables high-fidelity head geometry and texture reconstruction using a regular camera without relying on dedicated scanning equipment. This solves the problems of inaccurate dimensions in monocular reconstruction, gaps in multi-view texture seams, and poor hairstyle fit, improving the geometric realism and fitting accuracy of virtual try-on in intelligent eyeglass fitting scenarios. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating a lens feature modeling method based on multi-view images, provided in an embodiment of the present invention. Figure 2 A schematic diagram of a lens part modeling system based on multi-view images provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the internal structure of an electronic device that implements a lens part modeling method based on multi-view images, according to an embodiment of the present invention.
[0017] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0018] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.
[0019] This application provides a method for lens appearance modeling based on multi-view images. The executing entity of this method includes, but is not limited to, at least one electronic device configured to execute the method provided in this application, such as a server or a terminal. The server can be a standalone server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. In other words, the method for lens appearance modeling based on multi-view images can be executed by software or hardware installed on a remote device or server-side device; the software can be a blockchain platform. The server includes, but is not limited to, a single server, a server cluster, a cloud server, or a cloud server cluster.
[0020] Reference Figure 1 The diagram shown is a flowchart illustrating a lens feature modeling method based on multi-view images according to an embodiment of the present invention. In this embodiment, the lens feature modeling method based on multi-view images includes the following steps S1-S5: S1. Acquire a multi-view image sequence of the head based on the camera, parse the multi-view image sequence to obtain multiple head views, and obtain multiple original depth maps of the multiple head views, wherein one head view corresponds to one original depth map.
[0021] Understandably, the system automatically parses multiple head views (e.g., front view, left view, right view, etc.) from a multi-view image sequence and generates a corresponding original depth map for each view. This provides complementary facial geometry information from multiple angles for subsequent modeling, avoiding the tedious manual image selection while ensuring the integrity of the head contour.
[0022] The camera used is an RGB-D camera, and the multi-view image sequence refers to a continuous set of images captured around the head by an RGB-D camera (such as a mobile phone depth camera or a dedicated depth camera). Each frame contains both color (RGB) and depth information. This sequence covers multiple angles from the front to the side, providing a data foundation for subsequent selection of the best view and reconstruction of the complete head geometry.
[0023] Among them, multiple head views refer to a set of image frames selected from a multi-view image sequence that have different viewing directions (such as front view, left view, right view, etc.). Each view corresponds to a specific camera pose (pitch, yaw, roll angle) to provide complementary visual information for subsequent 3D reconstruction.
[0024] Among them, multiple raw depth maps refer to the unprocessed depth data extracted from the depth channels corresponding to multiple selected head views from different perspectives (such as front view, left view, right view, etc.). Each depth map records the distance information from the camera to various points in the scene. These raw depth maps are then used in conjunction with camera intrinsics to accurately backproject 2D facial key points into 3D space, providing absolute scale constraints.
[0025] Furthermore, the step of parsing the multi-view image sequence to obtain multiple head views includes: Camera pose estimation is performed based on the multi-view image sequence to obtain the initial pose data of the camera; The head deflection angle is calculated based on the initial posture data, and the multi-view image sequence is filtered based on the head deflection angle and a preset viewpoint classification threshold to obtain image frames that meet the viewpoint requirements. The image frames that meet the viewing angle requirements are calibrated and categorized to obtain multiple head views.
[0026] The initial attitude data refers to the camera rotation matrix and translation vector corresponding to each frame of the multi-view image sequence, which are initially calculated using camera pose estimation methods. Specifically, it can be decomposed into three rotational components: pitch angle, yaw angle, and roll angle, as well as the camera's position offset in space. This data serves as the basis for subsequent view selection and refinement, reflecting the camera's approximate orientation relative to the head during shooting.
[0027] Camera pose estimation refers to the process of using RGB image feature matching (such as ORB and SIFT features) or depth point cloud registration (such as the ICP algorithm) in a multi-view image sequence, combined with camera intrinsic parameters, to solve for the rotation and translation parameters of the camera in space when each frame was captured. This estimation can be implemented based on techniques such as visual odometry, PnP solving, or SLAM, and outputs the initial pose data corresponding to each frame.
[0028] The head yaw angle refers to the three spatial rotation components of the head relative to the camera coordinate system, calculated based on the camera's initial pose data (rotation matrix): pitch, yaw, and roll. This angle is used to measure the head's orientation in the current image frame and, by comparing it with a preset viewpoint classification threshold, determines whether the frame belongs to a front view, side view, or other viewpoint category.
[0029] The preset view classification threshold refers to a pre-defined set of angle ranges used to classify the camera's yaw angle, pitch angle, and roll angle into categories such as front view, left view, and right view. For example, a yaw angle within -15° to +15° and a pitch angle within -10° to +10° are classified as a front view, a yaw angle within +60° to +90° is classified as a left view, and a yaw angle within -90° to -60° is classified as a right view. This threshold can be adjusted according to the accuracy of the actual acquisition equipment and the modeling requirements.
[0030] Among them, image frames that meet the viewpoint requirements refer to those image frames whose yaw angle, pitch angle, and roll angle all fall within the preset viewpoint classification threshold range after initial attitude data calculation and viewpoint classification. These frames have clearly distinguishable frontal or side profile contours and are in an upright posture (without excessive tilting, lowering, or leaning), which can provide high-quality and complementary viewpoint input for subsequent modeling.
[0031] S2. Obtain two-dimensional facial key points from multiple head views, as well as the camera intrinsics of the camera. Based on multiple original depth maps and the camera intrinsics, back-project the two-dimensional facial key points to a preset three-dimensional camera coordinate system to obtain three-dimensional facial key points.
[0032] Understandably, by utilizing depth maps and camera intrinsics corresponding to multiple head views, the detected 2D facial key points (such as the corners of the eyes and the tip of the nose) in each view are back-projected into 3D space and fused to generate an accurate set of 3D key points. This solves the problems of depth blur and scale uncertainty in a single image, ensuring that the position and size ratio of the facial features in the head model are consistent with the actual height of the human face, avoiding the geometric distortion of monocular methods, and the multi-view fusion further improves the spatial robustness of the key points.
[0033] In this context, 2D facial keypoints refer to a set of semantically meaningful facial feature points detected from RGB images of multiple head views, such as the corners of the eyes, the tip of the nose, the corners of the mouth, and the auricles of the ears. Each point is represented by image pixel coordinates (x, y). These keypoints serve as geometric constraints, used for subsequent backprojection into 3D space to guide the fitting of the head model.
[0034] The camera intrinsic parameters are a parameter matrix describing the internal imaging geometry of the camera, typically including focal length (fx, fy), principal point coordinates (cx, cy), and distortion coefficients. These parameters are obtained in advance during camera calibration and are used to project three-dimensional spatial points onto a two-dimensional image plane, or conversely, to backproject two-dimensional pixels onto rays in three-dimensional space.
[0035] The 3D camera coordinate system is a three-dimensional Cartesian coordinate system with the camera's optical center as the origin and the optical axis as the Z-axis (or as defined according to different conventions). It is used to describe the true spatial position of an object relative to the camera. In this coordinate system, the coordinates of a point have a physical scale (such as millimeters), which can eliminate the depth blur problem in monocular images.
[0036] Among them, 3D facial key points refer to the spatial coordinates (X, Y, Z) of 2D facial key points obtained in the 3D camera coordinate system after backprojection of the two-dimensional facial key points through depth maps and camera intrinsic parameters. Each key point has real depth information, thus providing geometric constraints with absolute scale for subsequent FLAME model fitting and ensuring the accuracy of reconstructed head size.
[0037] S3. Based on the three-dimensional facial key points, the camera intrinsic parameters, and the preset camera extrinsic parameters, perform multi-stage distribution optimization on the preset head model to obtain a 3D head model.
[0038] Understandably, through a phased optimization strategy from coarse to fine, the shape, expression, posture, and gaze of a pre-defined head model are successively fitted, with 3D facial key points serving as geometric constraints. This ultimately generates a high-precision 3D head model with freely adjustable expression and gaze, providing an accurate geometric basis for subsequent texture mapping.
[0039] The preset camera extrinsic parameters refer to the initial values of the camera's external parameters given in advance, including the rotation matrix (describing the camera's orientation) and translation vector (describing the camera's position), which are usually initially estimated from the parsed pose data. These extrinsic parameters will be further refined in subsequent optimization processes to generate more accurate projection relationships.
[0040] The preset head model refers to a predefined parametric 3D head model, such as the FLAME model. This model controls the geometric shape of the head through shape parameters, expression parameters, pose parameters, and gaze parameters, serving as the initial template and statistical prior for optimization fitting in this method.
[0041] The 3D head model refers to a three-dimensional head mesh model with realistic facial geometric details, generated by optimizing the parameters (shape, expression, posture, and gaze) of a pre-set FLAME model through multiple stages. This head model has accurate physical dimensions and topology, making it suitable for subsequent texture mapping and glasses adaptation.
[0042] Furthermore, the step of performing multi-stage distribution optimization on the preset head model based on the three-dimensional facial key points, the camera intrinsic parameters, and the preset camera extrinsic parameters to obtain a 3D head model includes: Obtain an initial geometric function, construct an ear-specific loss term based on preset ear key points, and construct a combined geometric function based on the ear-specific loss term and the initial geometric function; The ear-specific loss term in the combined geometric function is adjusted to obtain an optimized geometric function; Based on the optimized geometric function, the 3D facial key points, the camera intrinsic parameters, and the preset camera extrinsic parameters, the preset head model is subjected to multi-stage distribution optimization to obtain a 3D head model.
[0043] The initial geometric function refers to the loss function constructed based on the geometric error between the three-dimensional facial key points and the corresponding vertices of the preset head model (such as the FLAME model). It is usually in the form of a weighted sum of squares and is used to quantify the deviation between the current model and the geometry of the real face, serving as the basis for subsequent ear-specific optimization and combined optimization.
[0044] Among them, ear key points refer to semantically meaningful two-dimensional feature points such as the auricle, earlobe, and tragus detected from the RGB image of the side view, and the coordinates of these points in three-dimensional space obtained by backprojecting the depth map into three-dimensional space. These key points serve as geometric constraints for ear reconstruction, addressing the problem of easily missing ear details in side profile contour fitting.
[0045] The ear-specific loss term refers to a loss function component constructed based on the Euclidean distance error between the 3D ear keypoints and the corresponding ear vertices in the FLAME model, typically in the form of a weighted sum of squares. This loss term acts solely on the ear region, assigning it a high optimization weight to guide the model to prioritize the accuracy of the ear contour during the fitting process, thereby improving the overall fit of the profile reconstruction.
[0046] The composite geometric function refers to a compound loss function obtained by linearly adding or weightedly fusing the initial geometric function (global loss based on all facial key points) and the ear-specific loss term (local enhancement loss for the ear) with certain weights. This function maintains the overall constraint on the shape of the entire face while enhancing the local supervision of the difficult-to-reconstruct areas of the ear, achieving the optimization goal of taking both global and local factors into account.
[0047] The optimized geometric function refers to the final loss function obtained by dynamically adjusting the ear-specific loss term in the combined geometric function (e.g., by adjusting its constraint strength through adaptive weight coefficients). Based on the initial geometric function and ear-specific constraints, this function achieves a balance between the global and local (especially ear-specific) fitting accuracy of the head model, guiding subsequent multi-stage distributed optimization.
[0048] Further, adjusting the ear-specific loss term in the combined geometric function to obtain the optimized geometric function includes: Obtain the current image frame in the multi-view image sequence, and determine the spatial location of the key points of the ear; Calculate the head deflection angle corresponding to the current image frame to obtain the current image viewpoint, and calculate the visibility of the ear region based on the current image viewpoint and the spatial position; Based on the current image viewpoint and the visibility of the ear region, calculate the adaptive weight coefficients corresponding to the ear-specific loss term; Based on the adaptive weighting coefficients, the constraint strength of the ear-specific loss term in the combined geometric function is dynamically adjusted to obtain the optimized geometric function.
[0049] The current image frame refers to an RGB-D image from a specific viewpoint being processed, typically from a frontal or side view sequence. This frame contains a color image and a corresponding depth map, used to calculate ear visibility from the current viewpoint and to dynamically adjust the weights for ear loss.
[0050] The spatial location of the ear key points refers to the three-dimensional coordinates (X, Y, Z) of points such as the auricle and earlobe obtained by back-projecting the ear key points (two-dimensional) onto the three-dimensional camera coordinate system using the original depth map and camera intrinsic parameters. This spatial location serves as a reference benchmark to determine whether the ear is occluded or visible from the current camera viewpoint.
[0051] The current image viewpoint refers to the head deflection angle calculated based on the initial pose data (rotation matrix) of the current image frame, specifically including pitch, yaw, and roll angles. This viewpoint is used to determine the camera's orientation relative to the head, and then analyze whether the ear region is observable in the image.
[0052] Ear region visibility refers to a Boolean value or continuous confidence score, based on the current image viewpoint and the ear's spatial position, to determine whether the ear is within the camera's field of view and is not obscured by other parts of the head (such as cheeks or hair). Ear keypoints are more reliable at viewpoints with high visibility and should be assigned higher loss weights; when visibility is low, the weights should be reduced to avoid misleading optimization.
[0053] The adaptive weighting coefficient is a value (usually between 0 and 1) dynamically calculated based on the current image viewpoint and the visibility of the ear region. It is used to adjust the contribution of the ear-specific loss term to the overall combined geometric function in real time. This coefficient allows the model to strengthen constraints when the ear is clearly visible and automatically weaken constraints when it is invisible or severely occluded, thereby improving the robustness of the fit.
[0054] The constraint strength of the ear-specific loss term refers to its influence on the composite geometric function, and is directly controlled by the adaptive weight coefficients. Higher strength means the model optimization tends to reduce the error between ear keypoints and model vertices; lower strength weakens the ear constraint, focusing optimization on other facial regions. By dynamically adjusting the constraint strength, the model achieves reasonable utilization of ear information from different viewpoints.
[0055] S4. Obtain the predefined texture space of the preset head model and the optimization parameters of the camera, and generate a standard head texture map based on the texture space, multiple head views, the 3D head model and the optimization parameters.
[0056] Understandably, by utilizing optimized camera parameters and a predefined texture space, color images from multiple head views are fused based on weights calculated according to the orientation of the model's facets and the angle of view from each perspective. Furthermore, areas not captured in the original view, such as the back of the ears and the top of the head, are intelligently filled in using an image diffusion algorithm. The result is a complete, seamless, and hole-free standardized head texture map, making the skin texture of the virtual head model realistic and natural.
[0057] In this context, texture space refers to the predefined two-dimensional UV parameter domain of the FLAME model (usually a rectangular area of [0, 1] × [0, 1]). Each vertex of the three-dimensional mesh corresponds to a fixed set of UV coordinates, used to map the pixels of the two-dimensional image onto the surface of the three-dimensional model. This space provides a unified coordinate framework for subsequent texture baking, enabling multi-view images to be fused into the same texture map.
[0058] The optimized parameters refer to the precise camera pose parameters obtained by iteratively refining the preset camera extrinsic parameters under the constraints of the optimized geometric function. These parameters include more accurate rotation matrices, translation vectors, and possible lens distortion compensation. These parameters are used to accurately project multi-view images onto the texture space of the 3D head model, ensuring geometric consistency in texture mapping.
[0059] The standard head texture map refers to a complete, continuous, seamless, and hole-free final texture map obtained by performing multi-view weighted fusion (calculating weights based on the angle between the facet orientation and the line of sight) and texture synthesis (filling holes in occluded areas such as behind the ears and the top of the head) on the initial texture map. When this map is applied to a 3D head model, it can make the virtual head present a realistic and natural skin texture, meeting the visual effect requirements of virtual try-on of glasses.
[0060] Further, generating a standard head texture map based on the texture space, multiple head views, the 3D head model, and the optimization parameters includes: Obtain the three-dimensional mesh points of the 3D head model, and the pixel values of each pixel point in the multiple head views; Based on the optimization parameters, the mapping relationship between grid points and pixels is found. Based on the mapping relationship, the pixel values of each pixel point are mapped to the texture space to obtain the initial texture map. The initial texture map is fused from multiple perspectives to obtain a standard head texture map.
[0061] In this context, 3D mesh points refer to the vertex coordinates of all triangular facets in the 3D head model. Each vertex has a fixed 3D position (x, y, z) in the model coordinate system, and each vertex is pre-associated with a set of UV texture coordinates (u, v). These mesh points serve as a bridge connecting 3D geometry and 2D textures, and subsequent projection and mapping operations are all based on these vertices.
[0062] The initial texture map refers to the unblended texture map obtained by projecting the RGB images of the front and side views onto the texture space of the FLAME model according to optimized parameters. Different regions in this map may come from different viewpoints and have obvious seams, overlaps, or blank areas (areas where texture was not captured due to occlusion), requiring subsequent processing to obtain a complete texture.
[0063] In this context, pixel values refer to the red, green, and blue (RGB) color components stored in each pixel in the front and side views, typically ranging from 0 to 255 or 0 to 1. These color values come from the original captured color image and represent the appearance information of a real human face at that pixel location. They will then be mapped into the texture space to generate the surface texture of the head model.
[0064] The mapping relationship between grid points and pixels refers to the association between each 3D grid point and a certain pixel coordinate (px, py) established after projecting each 3D grid point onto the pixel coordinate system of the front or side view through optimized parameters (intrinsic and extrinsic parameters). This correspondence clarifies "which grid point corresponds to which pixel in the image", thus ensuring that the color pixel value of that pixel can be correctly mapped to the UV texture coordinates of the grid point.
[0065] Further, the step of performing multi-view fusion on the initial texture map to obtain a standard head texture map includes: Obtain the normal vector of each facet in the 3D head model, and calculate the angle value corresponding to each facet based on the facet normal vector and the gaze vector of the optimization parameters; The texture fusion weight of each facet is calculated based on the included angle value corresponding to each facet, and the texture fusion weight of each facet is mapped to the texture space to generate a texture weight map with multiple perspectives. An intermediate texture map is generated based on the texture weight map from the multiple perspectives and the initial texture map; Obtain the occlusion area of the intermediate texture map, perform texture synthesis and hole filling on the occlusion area to obtain the standard head texture map.
[0066] The normal vector of a facet refers to the unit direction vector of each triangular facet in the 3D head model, perpendicular to the plane in which the facet lies, and is used to represent the orientation of the facet in space. This vector is calculated by cross product of the coordinates of the three vertices of the facet and is the basis for subsequent calculations of the angle between the facet and the viewpoint.
[0067] In the 3D head model, each facet refers to a triangular facet, which is the basic unit that makes up the 3D head mesh model. All facests are pieced together to form the complete head surface. Each facet has three vertices, one normal vector, and a corresponding texture coordinate region, and is the smallest geometric unit for texture weight calculation and fusion.
[0068] The included angle value for each facet refers to the spatial angle between the facet's normal vector and the camera's line of sight (usually ranging from 0° to 180°). A smaller angle indicates that the facet is more directly facing the camera, resulting in higher texture quality at that viewpoint; a larger angle (over 90°) indicates that the facet is facing away from the camera, and the texture is unusable.
[0069] The texture fusion weight for each facet is a value between 0 and 1 calculated using a mapping function (such as cosine or linear decay) based on the facet's angle value. This weight measures the texture reliability of the facet from the current viewpoint. The higher the weight, the greater the contribution of that viewpoint to the facet's color.
[0070] In this context, the texture weight map for multiple viewpoints refers to one or more grayscale images (one for each viewpoint) obtained by mapping the texture fusion weights of each patch under each viewpoint to the UV texture space through barycentric coordinate interpolation. The grayscale value of each pixel in the image represents the fusion weight of the corresponding texture position under that viewpoint, which is used for subsequent weighted fusion.
[0071] The intermediate texture map refers to the texture map generated by independently baking the initial texture maps for each viewpoint and then weighting and fusing them pixel-by-pixel according to the multi-view texture weight map. This map has eliminated seams and overlaps between multiple views, but there may still be holes or missing textures in areas that cannot be covered by the original viewpoints, such as behind the ear and on the top of the head.
[0072] In this context, the occluded areas of the intermediate texture map refer to the texture gaps that cannot be covered by the original shooting perspective (such as the back of the ear, the top of the head, and the bottom of the jaw), which are represented by black, transparent, or invalid pixels. These areas need to be semantically completed using texture synthesis and hole-filling algorithms based on the surrounding known textures to ultimately generate a complete standard head texture map.
[0073] S5. Identify the scalp region of the 3D head model, register the preset hairstyle point cloud based on the scalp region to obtain a registered hairstyle model, and construct an integrated head hairstyle model based on the registered hairstyle model, the 3D head model and the standard head texture map.
[0074] Understandably, based on the identified scalp area, the preset hairstyle point cloud is precisely registered to avoid the hairstyle floating or misaligning, and the registered hairstyle model is then fused with the 3D head model and standard head texture map. The final result is a unified head and hairstyle model that can be directly used for collision detection and rendering in virtual glasses try-on.
[0075] The scalp region refers to the mesh area extracted from the 3D head model, corresponding to the top of the head and the surrounding scalp surface. It typically includes all triangular facets within the hair root distribution area. This region serves as the target geometry for hairstyle registration, providing an alignment reference for the subsequent hairstyle point cloud and ensuring a seamless fit between the hair strand starting position and the head model surface.
[0076] The preset hairstyle point cloud refers to the 3D point cloud set of hair root positions extracted from the NPZ format hairstyle data pre-generated by the DiffLocks diffusion model after coordinate system transformation and alignment. Each point represents the starting point coordinates of a hair strand. This point cloud serves as the source data to be registered and needs to be rigidly aligned with the scalp area to obtain accurate hairstyle positions.
[0077] The registered hairstyle model refers to the transformed hairstyle point cloud (or preliminary hairstyle geometry) obtained by rigidly registering the hairstyle point cloud with the scalp region using an improved ICP algorithm with a global scaling factor, and applying optimal rotation, translation, and fine scaling parameters. This model accurately fits the scalp surface, avoiding hairstyle floating or misalignment issues, and providing an accurate hairstyle foundation for subsequent solidification.
[0078] The integrated head hairstyle model refers to a complete 3D head model formed by converting the registered hairstyle model into solid hair pieces (adjustable width meshes) through geometric nodes, assigning standard materials, and then merging it with the 3D head model and standard head texture map. This model is usually exported in GLB format and can be directly used for collision detection, rendering, and interaction in virtual glasses try-on, achieving an organic integration of hairstyle and head model.
[0079] Further, the step of registering the preset hairstyle point cloud based on the scalp region to obtain a registered hairstyle model includes: Obtain the scalp region point cloud of the scalp area, and calculate the root mean square radius ratio of the preset hairstyle point cloud and the scalp region point cloud; Based on the root mean square radius ratio, the preset hairstyle point cloud is initially scaled to obtain an initial scaled hairstyle point cloud; A rigid transformation is performed on the initial scaled hairstyle point cloud to obtain the registered hairstyle model.
[0080] The scalp region point cloud refers to a set of three-dimensional points sampled from the scalp region surface of the 3D head model. Each point has spatial coordinates (X, Y, Z) and is used for registration with the hairstyle point cloud. This point cloud represents the geometry of the scalp surface and serves as a target reference for global scaling and rigid transformation of the hairstyle.
[0081] The root mean square radius ratio refers to the ratio of the root mean square distance of each point in the hairstyle point cloud to its own center to the root mean square distance of each point in the scalp region point cloud to its own center. This ratio reflects the difference in overall size between the two point clouds and serves as a benchmark for subsequent global scaling.
[0082] The initial scaled hairstyle point cloud refers to the set of point clouds obtained by multiplying the original hairstyle point cloud by the root mean square radius ratio. Its overall size is comparable to that of the scalp region, but its position and pose are not yet aligned. This point cloud serves as the input for rigid transformation, providing a size-consistent starting point for subsequent fine registration.
[0083] This invention relates to a method, system, electronic device, and computer-readable storage medium for head modeling based on multi-view image sequence analysis of head views and corresponding depth maps; extraction of 2D facial key points, combined with depth maps and camera intrinsic parameters for back-projection to obtain 3D key points; multi-stage optimization of a preset head model based on 3D key points and intrinsic / extrinsic parameters to obtain a 3D head model; generation of standard texture maps using texture space, optimized camera parameters, and various views; identification of the scalp region, registration of the hairstyle point cloud, and construction of an integrated hairstyle model with the head model and texture map. Therefore, this invention proposes a multi-view image-based method for facial modeling in eyeglass fitting, which, through the construction of a multi-view analysis and step-by-step optimization architecture for the preset head model, enables high-fidelity head geometry and texture reconstruction using a regular camera without relying on dedicated scanning equipment. This solves the problems of inaccurate dimensions in monocular reconstruction, gaps in multi-view texture seams, and poor hairstyle fit, improving the geometric realism and fitting accuracy of virtual try-on in intelligent eyeglass fitting scenarios.
[0084] like Figure 2 The diagram shown is a schematic diagram of the lens head modeling system based on multi-view images of the present invention.
[0085] The multi-view image-based lens head modeling system 100 of this invention can be installed in an electronic device. Depending on the functions implemented, the multi-view image-based lens head modeling system may include an image processing module 101, a head model calculation module 102, a texture synthesis module 103, and a model unification module 104. The module described in this invention can also be called a unit, which refers to a series of computer program segments that can be executed by the processor of an electronic device and can perform a fixed function, and which are stored in the memory of the electronic device.
[0086] In this embodiment, the functions of each module / unit are as follows: The image processing module 101 is used to acquire a multi-view image sequence of the head based on the camera, parse the multi-view image sequence to obtain multiple head views, and generate multiple original depth maps of the multiple head views, with one head view corresponding to one original depth map. Two-dimensional facial key points are obtained from multiple head views, as well as the camera intrinsics of the camera. Based on multiple original depth maps and the camera intrinsics, the two-dimensional facial key points are back-projected onto a preset three-dimensional camera coordinate system to obtain three-dimensional facial key points. The head model calculation module 102 is used to perform multi-stage distribution optimization on the preset head model based on the three-dimensional facial key points, the camera intrinsic parameters and the preset camera extrinsic parameters to obtain a 3D head model. The texture synthesis module 103 is used to obtain the predefined texture space of the preset head model and the optimization parameters of the camera, and generate a standard head texture map based on the texture space, multiple head views, the 3D head model and the optimization parameters; The model unification module 104 is used to identify the scalp region of the 3D head model, register the preset hairstyle point cloud based on the scalp region to obtain the registered hairstyle model, and construct an integrated head hairstyle model based on the registered hairstyle model, the 3D head model and the standard head texture map.
[0087] In detail, the modules in the lens part modeling system 100 based on multi-view images described in this embodiment of the invention employ the same methods as described above. Figure 1 The method used is the same as the multi-view image-based lens facial modeling method and can produce the same technical effect, so it will not be described in detail here.
[0088] like Figure 3 The diagram shown is a structural schematic of an electronic device that implements the lens part modeling method based on multi-view images according to the present invention.
[0089] The electronic device may include a processor 10, a memory 11, a communication bus 12 and a communication interface 13, and may also include a computer program stored in the memory 11 and capable of running on the processor 10, such as a lens part modeling program based on multi-view images.
[0090] In some embodiments, the processor 10 may be composed of integrated circuits, such as a single packaged integrated circuit or multiple integrated circuits with the same or different functions, including combinations of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control unit of the electronic device, connecting various components of the entire electronic device through various interfaces and lines. It executes programs or modules stored in the memory 11 (e.g., executing a lens modeling program based on multi-view images) and calls data stored in the memory 11 to perform various functions of the electronic device and process data.
[0091] The memory 11 includes at least one type of readable storage medium, including flash memory, portable hard drive, multimedia card, card-type memory (e.g., SD or DX memory), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 11 can be an internal storage unit of an electronic device, such as a portable hard drive. In other embodiments, the memory 11 can be an external storage device of the electronic device, such as a plug-in portable hard drive, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. Furthermore, the memory 11 can include both internal and external storage units of the electronic device. The memory 11 can be used not only to store application software and various types of data installed on the electronic device, such as code for lens detail modeling programs based on multi-view images, but also to temporarily store data that has been output or will be output.
[0092] The communication bus 12 can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This bus can be divided into an address bus, a data bus, a control bus, etc. The bus is configured to enable communication between the memory 11 and at least one processor 10, etc.
[0093] The communication interface 13 is used for communication between the aforementioned electronic device and other devices, including a network interface and a user interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a Wi-Fi interface, Bluetooth interface, etc.), typically used to establish communication connections between the electronic device and other electronic devices. The user interface may be a display, an input unit (such as a keyboard), or optionally, a standard wired or wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen, etc. The display may also be appropriately referred to as a screen or display unit, used to display information processed in the electronic device and to display a visual user interface.
[0094] Figure 3 Only electronic devices with components are shown; it will be understood by those skilled in the art that... Figure 3 The structure shown does not constitute a limitation on the electronic device and may include fewer or more components than shown, or combine certain components, or have different component arrangements.
[0095] For example, although not shown, the electronic device may also include a power supply (such as a battery) to power the various components. Preferably, the power supply can be logically connected to the at least one processor 10 through a power management device, thereby enabling functions such as charging management, discharging management, and power consumption management. The power supply may also include one or more DC or AC power supplies, recharging devices, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The electronic device may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be described in detail here.
[0096] It should be understood that the embodiments described are for illustrative purposes only and are not limited to this structure in the scope of the patent application.
[0097] The lens head modeling program based on multi-view images stored in the memory 11 of the electronic device is a combination of multiple computer programs. When run in the processor 10, it can achieve the following: The head is acquired by a camera with a multi-view image sequence. The multi-view image sequence is parsed to obtain multiple head views and generate multiple original depth maps of the multiple head views. Each head view corresponds to one original depth map. Two-dimensional facial key points are obtained from multiple head views, as well as the camera intrinsics of the camera. Based on multiple original depth maps and the camera intrinsics, the two-dimensional facial key points are back-projected onto a preset three-dimensional camera coordinate system to obtain three-dimensional facial key points. Based on the three-dimensional facial key points, the camera intrinsic parameters, and the preset camera extrinsic parameters, a multi-stage distribution optimization is performed on the preset head model to obtain a 3D head model. Obtain the predefined texture space of the preset head model and the optimization parameters of the camera, and generate a standard head texture map based on the texture space, multiple head views, the 3D head model and the optimization parameters; Identify the scalp region of the 3D head model, register the preset hairstyle point cloud based on the scalp region to obtain a registered hairstyle model, and construct an integrated head hairstyle model based on the registered hairstyle model, the 3D head model and the standard head texture map.
[0098] Specifically, the processor 10's implementation method of the above-mentioned computer program can be found in [reference needed]. Figure 1 The descriptions of the relevant steps in the corresponding embodiments are not repeated here.
[0099] Furthermore, if the modules / units integrated into the electronic device are implemented as software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium. The computer-readable storage medium can be volatile or non-volatile. For example, the computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, or a read-only memory (ROM).
[0100] The present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor of an electronic device, can perform the following: The head is acquired by a camera with a multi-view image sequence. The multi-view image sequence is parsed to obtain multiple head views and generate multiple original depth maps of the multiple head views. Each head view corresponds to one original depth map. Two-dimensional facial key points are obtained from multiple head views, as well as the camera intrinsics of the camera. Based on multiple original depth maps and the camera intrinsics, the two-dimensional facial key points are back-projected onto a preset three-dimensional camera coordinate system to obtain three-dimensional facial key points. Based on the three-dimensional facial key points, the camera intrinsic parameters, and the preset camera extrinsic parameters, a multi-stage distribution optimization is performed on the preset head model to obtain a 3D head model. Obtain the predefined texture space of the preset head model and the optimization parameters of the camera, and generate a standard head texture map based on the texture space, multiple head views, the 3D head model and the optimization parameters; Identify the scalp region of the 3D head model, register the preset hairstyle point cloud based on the scalp region to obtain a registered hairstyle model, and construct an integrated head hairstyle model based on the registered hairstyle model, the 3D head model and the standard head texture map.
[0101] In the several embodiments provided by this invention, it should be understood that the disclosed devices, systems, and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and other division methods may be used in actual implementation.
[0102] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0103] Furthermore, the functional modules in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or in the form of hardware plus software functional modules.
[0104] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0105] Therefore, the embodiments should be considered exemplary and non-limiting in all respects, and the scope of the invention is defined by the appended claims rather than the foregoing description. Thus, all variations falling within the meaning and scope of equivalents of the claims are intended to be embraced within the invention. No appended diagram markings in the claims should be construed as limiting the scope of the claims.
[0106] The blockchain referred to in this invention is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.
[0107] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0108] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or systems stated in a system claim may also be implemented by a single unit or system through software or hardware. The term "second class" is used to indicate names and does not indicate any specific order.
[0109] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A method for modeling lens features based on multi-view images, characterized in that, The method includes: The head is captured by a camera with multiple viewpoint image sequences. The multiple viewpoint image sequences are parsed to obtain multiple head views. Multiple original depth maps of the multiple head views are obtained, wherein each head view corresponds to one original depth map. Two-dimensional facial key points are obtained from multiple head views, as well as the camera intrinsics of the camera. Based on multiple original depth maps and the camera intrinsics, the two-dimensional facial key points are back-projected onto a preset three-dimensional camera coordinate system to obtain three-dimensional facial key points. Based on the three-dimensional facial key points, the camera intrinsic parameters, and the preset camera extrinsic parameters, a multi-stage distribution optimization is performed on the preset head model to obtain a 3D head model. Obtain the predefined texture space of the preset head model and the optimization parameters of the camera, and generate a standard head texture map based on the texture space, multiple head views, the 3D head model and the optimization parameters; Identify the scalp region of the 3D head model, register the preset hairstyle point cloud based on the scalp region to obtain a registered hairstyle model, and construct an integrated head hairstyle model based on the registered hairstyle model, the 3D head model and the standard head texture map.
2. The lens appearance modeling method based on multi-view images as described in claim 1, characterized in that, The process of parsing the multi-view image sequence yields multiple head views, including: Camera pose estimation is performed based on the multi-view image sequence to obtain the initial pose data of the camera; The head deflection angle is calculated based on the initial posture data, and the multi-view image sequence is filtered based on the head deflection angle and a preset viewpoint classification threshold to obtain image frames that meet the viewpoint requirements. The image frames that meet the viewing angle requirements are calibrated and categorized to obtain multiple head views.
3. The lens appearance modeling method based on multi-view images as described in claim 1, characterized in that, The process involves multi-stage distribution optimization of a preset head model based on the 3D facial key points, camera intrinsic parameters, and preset camera extrinsic parameters to obtain a 3D head model, including: Obtain an initial geometric function, construct an ear-specific loss term based on preset ear key points, and construct a combined geometric function based on the ear-specific loss term and the initial geometric function; The ear-specific loss term in the combined geometric function is adjusted to obtain an optimized geometric function; Based on the optimized geometric function, the 3D facial key points, the camera intrinsic parameters, and the preset camera extrinsic parameters, the preset head model is subjected to multi-stage distribution optimization to obtain a 3D head model.
4. The lens appearance modeling method based on multi-view images as described in claim 3, characterized in that, The step of adjusting the ear-specific loss term in the combined geometric function to obtain the optimized geometric function includes: Obtain the current image frame in the multi-view image sequence, and determine the spatial location of the key points of the ear; Calculate the head deflection angle corresponding to the current image frame to obtain the current image viewpoint, and calculate the visibility of the ear region based on the current image viewpoint and the spatial position; Based on the current image viewpoint and the visibility of the ear region, calculate the adaptive weight coefficients corresponding to the ear-specific loss term; Based on the adaptive weighting coefficients, the constraint strength of the ear-specific loss term in the combined geometric function is dynamically adjusted to obtain the optimized geometric function.
5. The lens appearance modeling method based on multi-view images as described in claim 1, characterized in that, The process of generating a standard head texture map based on the texture space, multiple head views, the 3D head model, and the optimization parameters includes: Obtain the three-dimensional mesh points of the 3D head model, and the pixel values of each pixel point in the multiple head views; Based on the optimization parameters, the mapping relationship between grid points and pixels is found. Based on the mapping relationship, the pixel values of each pixel point are mapped to the texture space to obtain the initial texture map. The initial texture map is fused from multiple perspectives to obtain a standard head texture map.
6. The lens appearance modeling method based on multi-view images as described in claim 5, characterized in that, The process of multi-view fusion of the initial texture map to obtain a standard head texture map includes: Obtain the normal vector of each facet in the 3D head model, and calculate the angle value corresponding to each facet based on the facet normal vector and the gaze vector of the optimization parameters; The texture fusion weight of each facet is calculated based on the included angle value corresponding to each facet, and the texture fusion weight of each facet is mapped to the texture space to generate a texture weight map with multiple perspectives. An intermediate texture map is generated based on the texture weight map from the multiple perspectives and the initial texture map; Obtain the occlusion area of the intermediate texture map, perform texture synthesis and hole filling on the occlusion area to obtain the standard head texture map.
7. The lens appearance modeling method based on multi-view images as described in claim 1, characterized in that, The process of registering the preset hairstyle point cloud based on the scalp region to obtain the registered hairstyle model includes: Obtain the scalp region point cloud of the scalp area, and calculate the root mean square radius ratio of the preset hairstyle point cloud and the scalp region point cloud; Based on the root mean square radius ratio, the preset hairstyle point cloud is initially scaled to obtain an initial scaled hairstyle point cloud; A rigid transformation is performed on the initial scaled hairstyle point cloud to obtain the registered hairstyle model.
8. A lens appearance modeling system based on multi-view images, characterized in that, The system includes: The image processing module is used to acquire a multi-view image sequence of the head based on the camera, parse the multi-view image sequence to obtain multiple head views, and generate multiple original depth maps of the multiple head views, with one head view corresponding to one original depth map. Two-dimensional facial key points are obtained from multiple head views, as well as the camera intrinsics of the camera. Based on multiple original depth maps and the camera intrinsics, the two-dimensional facial key points are back-projected onto a preset three-dimensional camera coordinate system to obtain three-dimensional facial key points. The head model calculation module is used to perform multi-stage distribution optimization on the preset head model based on the three-dimensional facial key points, the camera intrinsic parameters and the preset camera extrinsic parameters to obtain a 3D head model. The texture synthesis module is used to obtain the predefined texture space of the preset head model and the optimization parameters of the camera, and generate a standard head texture map based on the texture space, multiple head views, the 3D head model and the optimization parameters; The model unification module is used to identify the scalp region of the 3D head model, register the preset hairstyle point cloud based on the scalp region to obtain the registered hairstyle model, and construct an integrated head hairstyle model based on the registered hairstyle model, the 3D head model and the standard head texture map.
9. An electronic device, characterized in that, The electronic device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the lens part modeling method based on multi-view images as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, It includes a data storage area and a program storage area. The data storage area stores the created data, and the program storage area stores the computer program. When the computer program is executed by the processor, it implements the lens head modeling method based on multi-view images as described in any one of claims 1 to 7.