Face multi-camera reconstruction system and method

By combining semantic segmentation networks and multi-face 3D deformable models, dynamic facial regions are identified in real time and weighted fusion is performed, which solves the problem of insufficient reconstruction accuracy in dynamic changes in 3D face modeling and achieves high-precision 3D face reconstruction results.

CN121505046AInactive Publication Date: 2026-02-10TIANSHUN (WUHAN) TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511672837.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-14
Publication Date
2026-02-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies struggle to handle large-angle head turns, occlusion changes, or rapid facial expression shifts in 3D face modeling, leading to reconstruction distortion or loss of key details. Furthermore, multi-camera systems lack sufficient reconstruction accuracy in dynamically changing areas and fail to perceive and adapt to the dynamics of facial regions.

Method used

A semantic segmentation network is used to identify dynamic facial regions in real time, and a multi-camera system is controlled to acquire images synchronously. The system is then calibrated and optimized by combining multiple 3D deformable facial models. Point cloud weighted fusion is performed based on semantic information, and a Poisson surface reconstruction algorithm is used to generate a high-precision 3D facial model.

Benefits of technology

It achieves high-precision 3D face reconstruction in complex dynamic scenes, improves the system's adaptability and robustness to complex dynamic changes, and ensures structural continuity and detail preservation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121505046A_ABST
    Figure CN121505046A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image reconstruction, in particular to a face multi-camera reconstruction system and method, and the method comprises the steps: S1, recognizing a face dynamic region in an input video stream in real time through a semantic segmentation network, generating a collection instruction, and controlling a multi-camera system to synchronously collect a multi-view image sequence; s2, performing calibration parameter optimization on the multi-camera system by using a multi-face three-dimensional deformable model as a semantic calibration target, and performing three-dimensional reconstruction on the multi-view image sequence based on the optimized parameters to generate an initial three-dimensional face point cloud; and S3, according to the semantic information of the face dynamic region, carrying out feature weighted fusion on the generated initial three-dimensional face point cloud, inhibiting the noise of the dynamic region, enhancing the details of the static region, and finally generating a high-precision three-dimensional face model. According to the method, the calculation efficiency is maintained, the reconstruction quality and the reality sense are considered, and the high-precision three-dimensional modeling capability for the complex dynamic face scene is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image reconstruction, in particular to a face multi-camera reconstruction system and method. BACKGROUND

[0002] With the rapid development of virtual reality, smart security, digital human modeling and other applications, high-precision three-dimensional face modeling technology has become an important research direction in the field of computer vision. Traditional single-camera or dual-camera stereo imaging schemes are limited by the number of viewing angles and synchronization, and often cannot stably obtain complete and accurate three-dimensional face models when facing large-angle head turning, occlusion changes or rapid expression changes. Especially in actual dynamic scenes, a single static modeling process lacks adaptability to the motion state of the face region, which easily leads to dynamic region deformation, reconstruction distortion or key detail loss, seriously restricting its practicality in high-precision applications.

[0003] On the other hand, although multi-camera systems have natural advantages in acquiring multi-view information, they still face challenges such as difficulty in unified calibration of camera parameters, accumulation of errors in dynamic changing regions, and noise interference in point cloud fusion process in actual deployment. Current methods generally lack perception and adaptation mechanisms for the dynamic nature of the face region, and fail to fully utilize semantic partition priors for differential weighting processing during point cloud fusion and surface modeling, resulting in insufficient reconstruction accuracy in dynamic regions and poor detail preservation in static regions. In addition, existing surface reconstruction methods mostly rely on regular grids or surface interpolation, making it difficult to balance structure continuity and local precision at the same time. Therefore, there is an urgent need for a three-dimensional face reconstruction method that combines dynamic semantic perception, multi-view high-precision acquisition, intelligent point cloud fusion and high-fidelity surface modeling capabilities to meet the modeling needs in complex real-world scenarios. SUMMARY

[0004] The present application provides a face multi-camera reconstruction system and method.

[0005] The face multi-camera reconstruction method comprises the following steps: S1: Real-time recognition of the face dynamic region in the input video stream through a semantic segmentation network, generation of acquisition instructions, and control of the multi-camera system to synchronously acquire multi-view image sequences; S2: Using a multi-face three-dimensional deformable model as a semantic calibration target, optimizing the calibration parameters of the multi-camera system, and based on the optimized parameters, performing three-dimensional reconstruction on the multi-view image sequences to generate an initial three-dimensional face point cloud; S3: According to the semantic information of the face dynamic region, feature-weighted fusion of the generated initial three-dimensional face point cloud is performed to suppress dynamic region noise and enhance static region details, and finally a high-precision three-dimensional face model is generated.

[0006] Optionally, the S1 comprises: S11: Deploy a lightweight semantic segmentation network in the processing unit of the multi-camera system. The lightweight semantic segmentation network is configured to perform real-time pixel-level segmentation on the input video stream and identify dynamic facial regions including the eyes, mouth and surrounding muscle groups. S12: Based on the dynamic region identification results output by the lightweight semantic segmentation network, calculate the motion amplitude of pixels in each region and the overall texture change frequency of the region. S13: Based on the motion amplitude and texture change frequency, construct a dynamic change scoring mechanism and generate acquisition instructions corresponding to the score.

[0007] Optionally, S13 includes: S131: When the score exceeds the preset threshold, it indicates that the facial expression changes drastically. The multi-camera system automatically switches to high-speed synchronous acquisition mode to improve the frame rate. S132: When the score is lower than the preset threshold, it indicates that the face is relatively still. Then, the acquisition frame rate is maintained or reduced to achieve dynamic balance and optimization between image quality and system resources.

[0008] Optionally, S2 includes: S21: Acquire face images of multiple individuals simultaneously captured from different perspectives by the multi-camera system, and construct a multi-face calibration image set; S22: Fit the three-dimensional deformable face model to each face in the multi-face calibration image set, initialize the three-dimensional shape parameters and pose parameters of each face, and initialize the intrinsic and extrinsic parameters of the multi-camera system. S23: Construct a joint optimization objective function, and simultaneously optimize the intrinsic and extrinsic parameters of the multi-camera system and the model parameters of each face by minimizing the joint optimization objective function to obtain the optimized camera parameters; S24: Based on the optimized camera parameters, perform 3D reconstruction on the acquired multi-view image sequence to generate an initial 3D face point cloud.

[0009] Optionally, the joint optimization objective function includes a geometric reprojection error term and a semantic consistency error term.

[0010] Optionally, S24 includes: S241: Based on the optimized camera intrinsic and extrinsic parameters, perform disparity stereo matching processing on the acquired multi-view image sequence, extract the pixel correspondence of the same facial feature point in different view images, and generate a dense disparity map. S242: Based on the dense disparity map results and combined with the intrinsic and extrinsic parameters of the multi-camera system, the spatial coordinates of the matched pixels are solved using the multi-view triangulation method to reconstruct the corresponding three-dimensional point cloud data and generate an initial three-dimensional face point cloud model.

[0011] Optionally, S3 includes: S31: The semantic labels of the dynamic facial regions identified by the semantic segmentation network are projected from the two-dimensional image space onto the initial three-dimensional facial point cloud, and a semantic partition label is assigned to each three-dimensional point. The semantic partition includes at least dynamic regions and static regions. S32: Assign fusion weights to each 3D point based on the semantic partition labels; S33: Fuse point cloud data generated from multiple camera perspectives. During the fusion process, for overlapping 3D points at the same location in space, perform a weighted average based on their respective fusion weights to obtain more stable 3D point coordinates. S34: Perform surface reconstruction and meshing on the fused point cloud to generate a high-precision 3D face model.

[0012] Optionally, S32 includes: S321: The weights of the three-dimensional points in the dynamic region are assigned based on the motion amplitude of the region to which they belong; the greater the motion amplitude, the lower the weight. S322: Weighting of 3D points in static regions is based on the stability of local surface curvature and the consistency of color and texture.

[0013] Optionally, the surface reconstruction is processed using the Poisson surface reconstruction algorithm, which treats the normal information of each three-dimensional point in the fused point cloud as the gradient field of an implicit function, and generates a continuous and smooth three-dimensional face surface by solving the corresponding Poisson equation.

[0014] A multi-camera face reconstruction system, used to implement the above-mentioned multi-camera face reconstruction method, includes the following modules: Dynamic Region Recognition and Acquisition Control Module: Deploys a semantic segmentation network to perform real-time pixel-level segmentation of the input video stream, recognizes the dynamic regions of the face, and generates acquisition commands accordingly to control the multi-camera system to perform multi-view synchronous image acquisition; Semantic calibration and initial reconstruction module: Using a multi-face 3D deformable model as a calibration target, the multi-camera system is jointly calibrated and optimized, and 3D reconstruction is performed on the acquired images based on the optimized camera parameters to generate an initial 3D face point cloud; Weighted fusion and fine modeling module: It fuses the initial 3D face point cloud with the semantic information of the dynamic face region, performs feature weighting processing according to the region characteristics, suppresses dynamic noise, enhances static details, and then completes the fusion of 3D point cloud and surface modeling, outputting a high-precision 3D face model.

[0015] The beneficial effects of this invention are: This invention introduces a semantic segmentation network to perform real-time recognition of dynamic facial regions in the input video stream, thereby dynamically adjusting the acquisition frame rate and mode of the multi-camera system. This achieves an intelligent control strategy that automatically switches to high-speed acquisition mode in areas of intense motion and maintains or reduces the frame rate in static areas. This mechanism not only improves the efficiency of keyframe capture but also effectively avoids the generation of redundant data, providing a high-quality, time-aligned image foundation for subsequent 3D modeling and enhancing the system's adaptability and robustness to complex dynamic changes in real-world scenes.

[0016] This invention constructs a point cloud weighted fusion mechanism based on dynamic region semantic information as a priori, dividing the face region into dynamic and static sub-regions, and adaptively assigning weights to each 3D point according to factors such as motion amplitude, surface curvature stability, and texture consistency. In the multi-view point cloud fusion process, high-weight points contribute more to the final model, effectively suppressing noise interference in dynamic regions. At the same time, the Poisson surface reconstruction algorithm is subsequently used to perform high-fidelity modeling of the fused point cloud, ensuring the excellent performance of the final 3D face model in terms of structural continuity, detail preservation, and topological integrity. While maintaining computational efficiency, it also takes into account reconstruction quality and realism, realizing high-precision 3D modeling capabilities for complex dynamic face scenes. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only for this invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart of a method according to an embodiment of the present invention; Figure 2 This is a system block diagram of an embodiment of the present invention. Detailed Implementation

[0019] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments. Those skilled in the art may employ other alternative methods to implement some well-known technologies; moreover, the accompanying drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.

[0020] like Figure 1 As shown, the multi-camera face reconstruction method includes the following steps: S1: Real-time identification of dynamic facial regions in the input video stream via a semantic segmentation network, generating acquisition instructions to control the multi-camera system to synchronously acquire multi-view image sequences; S1 specifically includes: S11: Deploy a lightweight semantic segmentation network (an improved MobileNet-DeepLab architecture) on the edge processing unit of a multi-camera system, with continuous video stream frames as input. The output is the corresponding pixel-level semantic segmentation map. , is represented as: ; in, This represents a semantic segmentation network. Indicates time Input video frames, Indicates the segmentation result. , The image height is 480-1080. The image width, ranging from 640 to 1920. This represents the number of semantic categories, with a value ranging from 4 to 8. This approach aims to achieve real-time recognition and annotation of dynamic facial regions in an input video stream. Its core design involves deploying a lightweight semantic segmentation network within the edge processing unit of a multi-camera system. This allows each frame to be quickly and accurately segmented into specific facial functional regions, such as the eyes, mouth, and cheek muscles. This pixel-level segmentation provides refined regional guidance for subsequent acquisition strategy adjustments and 3D reconstruction.

[0021] The design is based on the differentiated dynamic characteristics of different regions of the face during facial expression changes. For example, the mouth undergoes high-frequency deformation during speech, and the eyes exhibit rapid localized movements during blinking or gaze shifts. Therefore, relying solely on the overall facial movement amplitude to determine the acquisition rhythm can easily lead to information loss or redundant acquisition. By deploying a semantic segmentation network, different regions can be accurately distinguished and their dynamic features extracted individually, facilitating more granular dynamic detection and response.

[0022] The lightweight network (an improved MobileNet-DeepLab) was chosen to balance real-time performance with resource consumption, enabling the entire segmentation process to run locally on terminal devices (such as edge nodes or multi-camera acquisition controllers) without relying on remote computing resources, thus ensuring system response speed and deployment flexibility. Deploying the segmentation network locally also facilitates deep coupling with camera control logic, achieving a closed-loop acquisition mechanism integrating recognition, response, and acquisition. This solution lays a solid foundation for subsequent dynamic region-oriented acquisition, 3D reconstruction optimization, and semantic fusion.

[0023] S12: Based on regions with consistent semantic labels in adjacent frames Calculate the average motion amplitude of pixels in each region. With local texture change frequency , represented as: ; ; in, Indicates the pixel at time [time]. The spatial position or optical flow displacement vector, with values ​​ranging from... , Indicates the image at point The grayscale or Gabor response texture intensity at that location, with a value ranging from 0 to 255. Indicates the region Number of pixels Indicates the region The average motion amplitude, ranging from 0 to 10, indicates that muscle movement in facial expressions is typically within 5-10 pixels, with values ​​exceeding 10 pixels being rare. A value that is too small may indicate that the area is stationary. Indicates the region The average texture change frequency ranges from 0 to 30. Changes in grayscale or Gabor response are usually affected by lighting, facial expressions, jitter, etc. The intensity of the change mostly falls within this range. If it exceeds 30, it may be due to drastic changes in lighting or incorrect matching.

[0024] This step involves analyzing the movement and changes of different semantic regions of the face in the video stream between consecutive frames during the multi-camera face reconstruction process to quantify the activity level of the current facial expression, thereby providing a basis for subsequent dynamic acquisition and scheduling.

[0025] Specifically, the semantic segmentation results of two adjacent frames are first aligned to extract the overlapping parts of the same semantic tags (such as mouth, eyes, etc.) in the two frames. Then, the average motion amplitude of pixels within each region is calculated to measure the degree of muscle displacement within that region. Simultaneously, the intensity of grayscale or texture response changes in each region is extracted to reflect the dynamic changes in local illumination and morphology. These two metrics together describe the temporal dynamic activity of the region. By introducing texture change metrics instead of relying solely on displacement amplitude, robustness can be further improved, avoiding misjudgments of facial expressions due to camera shake or lighting fluctuations, thereby enhancing the system's judgment accuracy and response precision. This dynamic region evaluation mechanism based on dual-metric fusion is one of the foundations for achieving intelligent acquisition scheduling and high-quality face reconstruction.

[0026] S13: Define the global data acquisition scheduling function, expressed as: ; when At that time, the acquisition command triggers the high-speed synchronous acquisition mode; when At that time, the sampling frequency will be reduced to the standard mode; when At the same time, maintain the current sampling frequency.

[0027] in, This represents the set of all identified facial dynamic regions. This is a weighting coefficient used to adjust the contribution of motion amplitude and texture changes to the final scheduling index, with a value ranging from 0.3 to 0.7. The upper threshold for frame rate switching. The lower threshold for frame rate switching. The system scores overall dynamic changes. Through this dynamic change perception mechanism, the synchronous acquisition mode of the multi-camera system can be adjusted in real time. This improves the capture frequency and image quality when facial expressions change drastically, while reducing redundant acquisition burden and optimizing resource utilization and data quality in a stable state.

[0028] In 3D reconstruction, the required temporal resolution of images varies significantly depending on the facial expression state. Insufficient acquisition frequency during rapid changes can lead to motion blur and reconstruction distortion; conversely, continuous high frame rate acquisition during static facial phases wastes computation and storage resources. Therefore, introducing this dynamic scheduling mechanism significantly improves the system's energy efficiency while maintaining reconstruction accuracy. Furthermore, the parameters in the weighting mechanism allow for flexible adjustment of the response sensitivity to different types of changes (such as muscle movement and texture changes), enhancing the system's adaptability and controllability in various application scenarios. The entire mechanism forms a closed-loop acquisition and control process based on perception-evaluation-response, providing stable and efficient data support for high-quality dynamic face reconstruction.

[0029] S2: Using a multi-face 3D deformable model as a semantic calibration target, the calibration parameters of the multi-camera system are optimized, and based on the optimized parameters, the multi-view image sequence is reconstructed in 3D to generate an initial 3D face point cloud. S2 specifically includes: S21: Acquire an image sequence simultaneously captured from different perspectives by a multi-camera system. The images include the faces of multiple individuals, forming a multi-face calibration image set, represented as: ; in, For multi-face image sets, Indicates the first The first camera captured the first Individual facial images, This represents the total number of faces participating in the calibration, ranging from 5 to 20. It is recommended to keep the number of samples within this range to balance training generalization and computational cost. Fewer than 5 samples may lead to insufficient model fitting, while more than 20 samples result in diminishing returns during the calibration phase. This indicates the number of cameras, ranging from 4 to 12. The number of cameras determines the field of view coverage and ensures the accuracy of multi-angle, occlusion compensation, and depth reconstruction. The accuracy of 3D reconstruction and camera calibration highly depends on the sufficiency of viewpoint coverage and the richness of geometric features. Using multiple cameras to simultaneously acquire facial data from different individuals from multiple angles not only effectively enhances the diversity of structural information in the calibration image set but also improves the robustness and generalization ability of the fitting through multiple independent samples. Furthermore, acquiring multiple faces instead of a single face allows the 3D deformable model used to better model identity differences and facial expression changes during the fitting process, thus avoiding the model getting trapped in local optima or overfitting individual features. In addition, employing a synchronous acquisition mechanism ensures that all images are acquired at the same time, which facilitates subsequent rigorous spatial alignment and temporal consistency processing, thereby further improving the reconstruction quality.

[0030] S22: 3D deformable human face model The model is fitted to each face in the image set, and the model parameters are initialized. Simultaneously, the intrinsic parameters of each camera are initialized. With extrinsic parameters (attitude matrix) , is represented as: ; in, For the first A 3D shape model of a person's face, represented as a set of vertex coordinates. It is a mean face model, which is the average 3D shape of all face samples, used as an initial reference shape. This is the identity shape basis, a principal component basis matrix used to describe the differences in appearance between individuals. The identity dimension determines the modeling ability for diverse facial shapes, with a value ranging from 80 to 150. Using a higher identity dimension is to better fit differences in gender, race, and facial proportions. It is a shape basis for facial expressions, used to describe the deformation basis of facial expression changes. The expression dimension is generally less than the identity dimension, mainly describing mouth shape, eyebrows, cheek movements, etc., with a value range of 20-60. The low number of expression dimensions is sufficient to capture common dynamic features (such as opening the mouth, raising eyebrows, etc.). The identity coefficient vector is a set of coefficients that control the shift of the face shape towards different identity features, with values ​​ranging from [value range missing]. To prevent abnormal deformation The expression coefficient vector is a set of coefficients that control the degree of facial distortion in the expression space, with values ​​ranging from 1 to 2. The facial expression distortion remains within the range of real human facial expression distribution; It is the first The intrinsic parameter matrix of a camera, including camera intrinsic parameters such as focal length, principal point coordinates, and radial distortion. It is an extrinsic rotation matrix, the first... The orientation of each camera relative to the world coordinate system. Let be the extrinsic translation vector, the first... The position offset of each camera in the world coordinate system, with a value range of [value missing]. Multi-camera systems are typically set up around the face, with the relative distance controlled within 1-2 meters to ensure stereo accuracy and parallax redundancy; S23: After completing the initial fitting of the multi-face 3D model and the initialization of the calibration parameters of the multi-camera system, a joint optimization objective function is further constructed to improve the overall reconstruction accuracy of the system. This is achieved by minimizing the objective function to simultaneously optimize all camera parameters and face model parameters, specifically including: (1) Geometric reprojection error term: This term measures the geometric deviation between the 3D face model projected onto the 2D image plane under the current camera parameters and the actual detected 2D face feature points. It reflects the consistency between the current model and the image observation and is expressed as: ; in, It is the first Zhang Renfa (the first one) Three-dimensional feature points, It is the first The 3D feature point at the th ... Two-dimensional position in a camera image This is a 3D to 2D projection function. It is the number of feature points used for labeling each face; (2) Semantic consistency error term: To prevent the shape parameters of the 3D deformable face model from deviating from the statistical prior space, a regularization constraint term, namely the semantic consistency error term, is introduced to limit the norms of the identity parameters and expression parameters to a reasonable range, and is expressed as: ; This term is used to constrain the model parameters to not deviate from the statistical prior of the 3DMM face shape space; (3) Joint optimization objective function: In order to balance geometric reprojection error and semantic prior consistency, a joint loss function is adopted in the overall optimization, which is expressed as: ; in, The regularization weight coefficient is usually taken as... By minimizing the above objective function, the LM algorithm or SGD optimizer is used to simultaneously optimize all cameras. And each person's face The final output is the optimized camera parameter set; S24: Based on the optimized intrinsic and extrinsic parameters, perform 3D reconstruction on the acquired multi-view image sequence. Employ the semi-global matching algorithm and multi-view triangulation to generate an initial 3D face point cloud, specifically including: (1) The disparity stereo matching algorithm (Semi-Global Matching, SGM) is used to perform pixel-level matching between the main view and each auxiliary view, calculate the positional correspondence of each pixel in the image from different perspectives, and generate a dense disparity map to describe the spatial displacement characteristics of pixels in multiple perspectives. Let the front view image be The auxiliary view image is Define each pixel Assuming parallax The cost function is as follows: ; To enhance global consistency, each pixel is processed in multiple directions. Path cost aggregation is performed on (e.g., horizontal, vertical, diagonal) lines to obtain the path cost, which is represented as: ; Finally, the costs from multiple directions are aggregated to obtain a disparity map, and the disparity corresponding to the minimum cost is taken as the optimal match, as shown below: ; ; in, For knowledge graph images, For the first Auxiliary view image, For pixel coordinates, This is the parallax value, the horizontal displacement between corresponding pixels in the main view and the auxiliary view. Its value ranges from 0 to 128, covering most actual distance and depth differences. The initial cost function is the similarity error between a pixel in the main view and a pixel in the auxiliary view under the assumed disparity. For the direction of aggregation, For the set of aggregation directions, For single-path cost function, This is a penalty term for small disparity changes, representing the penalty weight when the disparity change between adjacent pixels is ±1. Its value ranges from 5 to 20. Too small a value results in excessive disparity discontinuity, while too large a value suppresses natural depth undulations. This is a large disparity jump penalty term, with a larger penalty weight when the disparity between adjacent pixels changes abruptly (not ±1). Its value ranges from 50 to 100, and it is used to strongly suppress abrupt disparity changes, effectively reducing noisy matching and erroneous jumps. It is suitable for curved but locally smooth structures such as faces. The aggregated cost function is the weighted sum of the costs of all paths in all directions. The optimal disparity result is the disparity value that minimizes the total cost, which is used as the final disparity estimate for that pixel. (2) Based on the dense disparity map information and the optimized camera parameters, perform multi-view triangulation to recover the point cloud structure of the face in three-dimensional space, as shown below: ; ; in, It is the first one obtained from the reconstruction A three-dimensional point, It is the first The pixel positions of a 3D point in images from different cameras. The multi-view triangulation function of a three-dimensional point. It is the initial 3D face point cloud composed of all 3D points; In the process of multi-camera 3D face reconstruction, it is necessary to first establish dense pixel-level correspondences between multiple camera images, that is, to complete the matching of the same 3D point in different images. To achieve high-precision pixel matching, this invention adopts a semi-global stereo matching algorithm (SGM). The core idea of ​​SGM is to accumulate the matching cost of pixels in multiple path directions (such as horizontal direction, diagonal direction, etc.) to form a globally consistent disparity estimation map. In practice, the system first calculates the matching cost of each pixel under a series of assumed disparity values ​​(such as based on gray-level difference, Census transform, etc.) to construct an initial cost volume; then, it performs path cost aggregation in multiple directions, so that the disparity estimation considers both local image content and global constraints from a larger range; finally, it selects the disparity value with the minimum total cost for each pixel as the final result. This algorithm has high robustness and matching accuracy, and is particularly suitable for situations where the facial region has complex texture changes, uneven lighting, or weak texture areas; the high-quality disparity map obtained by SGM provides a reliable point pair basis for subsequent triangulation, which is a prerequisite for generating accurate 3D face point clouds. After completing pixel-level matching between multiple camera images, this invention further employs multi-view triangulation to reconstruct the spatial coordinates of the matching points to obtain the initial point cloud of the 3D face. This method, based on the known intrinsic parameters (focal length, principal point position, etc.) and extrinsic parameters (spatial position and orientation) of each camera in the multi-camera system, back-projects the matched 2D pixels into spatial rays. Theoretically, rays from different cameras should intersect at the same actual point in 3D space. However, due to noise and matching errors, the actual rays may not intersect perfectly. Therefore, the system uses a least-squares optimization strategy to find the spatial intersection point that minimizes the reprojection error, which serves as the 3D coordinate of that pixel. This process is performed on all matching points in batches, ultimately forming a complete 3D face point cloud structure. Compared to binocular methods, multi-view triangulation has stronger geometric constraints and occlusion robustness, especially in areas with drastic changes in facial depth (such as the tip of the nose, eye sockets, and jawline), exhibiting higher accuracy and continuity, effectively improving the structural integrity and morphological accuracy of 3D face reconstruction.

[0031] This step, based on the correspondence between multi-view images, uses stereo matching and triangulation to recover the true structure of a human face in three-dimensional space from a two-dimensional image. Specifically, the system utilizes the pixel positions of the same feature point identified in multiple camera images across different images, combined with the intrinsic parameters (such as focal length and principal point) and extrinsic parameters (i.e., the camera's spatial position and orientation) of each camera, to calculate the precise coordinates of that point in three-dimensional space using a triangulation algorithm. This process is performed on all feature points one by one, ultimately forming a point cloud composed of three-dimensional points, representing the three-dimensional geometry of the observed human face.

[0032] S3: Based on the semantic information of the dynamic regions of the face, the generated initial 3D face point cloud is subjected to feature weighted fusion to suppress noise in the dynamic regions and enhance the details in the static regions, ultimately generating a high-precision 3D face model. S3 specifically includes: S31: Obtain the two-dimensional image semantic labels through the semantic segmentation network. The model is projected onto a 3D point cloud via camera projection. Every three-dimensional point on Assign a semantic partition label to each point ; in, Semantic labels are categorized into two types: dynamic regions (such as the mouth and eyes) and static regions (such as the forehead and cheeks). Therefore, the label value is either 0 or 1, simplifying the classification process. This indicates that the point is in a dynamic region (such as a highly deformable region like the mouth or eyes). This indicates that the point is located in a static region (such as a stable structure like the forehead, cheek, or bridge of the nose), and the mapping rule is as follows: , Indicates the first Projection function of each camera; In the process of 3D reconstruction, mapping semantic labels to 3D point clouds is to combine the dynamic and static regional information of a face with the actual 3D structure. Dynamic regions usually represent changes in facial expressions, while static regions represent unchanging facial structures. Using the camera's projection function to map the semantic labels of 2D images to 3D point clouds is a standard method in computer vision. Each camera's projection model, based on the camera's intrinsic and extrinsic parameters, can map points in 3D space to the 2D image plane, thus ensuring an accurate correspondence between image space and 3D space. Through this mapping process, we can effectively combine the semantic information of the image (such as dynamic or static regions) with the 3D model, providing a semantic foundation for subsequent point cloud weighted fusion.

[0033] S32: For each three-dimensional point Assign fusion weights The partitioning logic includes: (1) For (Dynamic area): ; (2) For (Static area): ; in, This is the corresponding point in three-dimensional space, belonging to the first point in the initial three-dimensional face point cloud. One point, This represents the range of motion within the region to which the point belongs, with a value ranging from 0 to 20. It typically reflects changes in facial expressions, and the maximum motion range generally does not exceed 20 pixels. Larger values ​​will result in a lower weight for that point. This index represents the local surface curvature stability at that point, with a value ranging from 0 to 1. Local curvature stability is a measure of surface smoothness; a higher value indicates a smoother and more stable surface. This represents the color or texture consistency score for that point, ranging from 0 to 1. The color or texture consistency score measures the texture fidelity of a static area; a higher score indicates more consistent texture at that point. This is the dynamic weight decay coefficient, which controls the degree of weight decay of points in the dynamic region. Its value ranges from 1.0 to 3.0. The dynamic region weight decay coefficient is used to adjust the influence of points in the dynamic region; the higher the value, the lower the weight of the dynamic region. This is the static region fusion coefficient, which controls the balance between static regions and texture stability. Its value ranges from 0.5 to 0.8. It controls the balance between structure and texture. The static region fusion coefficient controls the balance between static regions and texture. A higher value (such as 0.7-0.8) indicates that the static region has a larger weight in the fusion. S33: A collection of point clouds generated from multiple camera views Weighted fusion is performed. For overlapping points at the same location in space (through spatial clustering or voxel hash matching), a weighted average is used to generate fusion points, represented as: ; in, Indicates from the The first point cloud A spatial location, For the first The fusion weights of each point cloud in the corresponding point cloud, weighted fusion has significant advantages in suppressing dynamic region drift and preserving static region details, and the value range is [value range missing]. , To determine the location of a point after fusion, the weighted fusion result of multiple viewpoints at the same location in space is used as the coordinates of a point in the final point cloud set. Number of cameras / views: This represents the total number of cameras in the multi-camera system participating in the fusion, ranging from 2 to 16. Multi-camera systems typically consist of two or more cameras to achieve multi-view acquisition. It is the first Point clouds from various perspectives; S34: Weighted fusion of point cloud sets Surface reconstruction and meshing were performed on the dense point cloud data. The Poisson Surface Reconstruction algorithm was used to reconstruct the surface, generating a high-precision 3D face mesh model with continuous topology and complete structure, represented as follows: ; in, To achieve the final high-precision 3D face model, For merging point cloud collections; The Poisson surface reconstruction algorithm can effectively fill in some missing data regions, improve the topological integrity of the model, has strong robustness to noise, and can generate smooth, closed facial surface structures, making it suitable for detailed facial modeling. The core idea of ​​the Poisson surface reconstruction algorithm is to reconstruct the surface of each point in the point cloud. The corresponding normal vector is used to reconstruct the implicit function by solving the following Poisson equation. Specifically, it is expressed as: ; in, Represents the Laplace operator. It is a vector field composed of normal vectors in the point cloud. The iso-surface represents the surface of the final reconstructed 3D model. The Poisson reconstruction function from a 3D reconstruction library (such as Open3D or MeshLab) is used to reconstruct the surface. The process is repeated to obtain the final high-precision 3D face mesh model. Its structure is complete and can be used in various application scenarios such as subsequent recognition, animation, and medical reconstruction.

[0034] like Figure 2 As shown, the multi-camera face reconstruction system, used to implement the above-mentioned multi-camera face reconstruction method, includes the following modules: Dynamic Region Recognition and Acquisition Control Module: Deploys a semantic segmentation network to perform real-time pixel-level segmentation of the input video stream, recognizes the dynamic regions of the face, and generates acquisition commands accordingly to control the multi-camera system to perform multi-view synchronous image acquisition; Semantic calibration and initial reconstruction module: Using a multi-face 3D deformable model as a calibration target, the multi-camera system is jointly calibrated and optimized, and 3D reconstruction is performed on the acquired images based on the optimized camera parameters to generate an initial 3D face point cloud; Weighted Fusion and Fine Modeling Module: This module fuses the initial 3D face point cloud with the semantic information of the dynamic face region, performs feature weighting based on region characteristics, suppresses dynamic noise, enhances static details, and thus completes the fusion of the 3D point cloud and surface modeling, outputting a high-precision 3D face model.

[0035] This invention encompasses any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of this invention. To provide the public with a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments; however, those skilled in the art will fully understand the invention even without these details. Furthermore, to avoid unnecessary misunderstanding of the essence of this invention, well-known methods, processes, procedures, components, and circuits are not described in detail.

[0036] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A multi-camera face reconstruction method, characterized in that, Includes the following steps: S1: Real-time identification of dynamic facial regions in the input video stream via a semantic segmentation network, generating acquisition instructions to control the multi-camera system to synchronously acquire multi-view image sequences; S2: Using a multi-face 3D deformable model as a semantic calibration target, the calibration parameters of the multi-camera system are optimized, and based on the optimized parameters, the multi-view image sequence is reconstructed in 3D to generate an initial 3D face point cloud. S3: Based on the semantic information of the dynamic facial region, the generated initial 3D facial point cloud is subjected to feature weighted fusion to suppress noise in the dynamic region and enhance details in the static region, ultimately generating a high-precision 3D facial model.

2. The face multi-camera reconstruction method according to claim 1, characterized in that, S1 includes: S11: Deploy a lightweight semantic segmentation network in the processing unit of the multi-camera system. The lightweight semantic segmentation network is configured to perform real-time pixel-level segmentation on the input video stream and identify dynamic facial regions including the eyes, mouth and surrounding muscle groups. S12: Based on the dynamic region identification results output by the lightweight semantic segmentation network, calculate the motion amplitude of pixels in each region and the overall texture change frequency of the region. S13: Based on the motion amplitude and texture change frequency, construct a dynamic change scoring mechanism and generate acquisition instructions corresponding to the score.

3. The face multi-camera reconstruction method according to claim 2, characterized in that, S13 includes: S131: When the score exceeds the preset threshold, it indicates that the facial expression changes drastically. The multi-camera system automatically switches to high-speed synchronous acquisition mode to improve the frame rate. S132: When the score is lower than the preset threshold, it indicates that the face is relatively still. Then, the acquisition frame rate is maintained or reduced to achieve dynamic balance and optimization between image quality and system resources.

4. The face multi-camera reconstruction method according to claim 1, characterized in that, S2 includes: S21: Acquire face images of multiple individuals simultaneously captured from different perspectives by the multi-camera system, and construct a multi-face calibration image set; S22: Fit the three-dimensional deformable face model to each face in the multi-face calibration image set, initialize the three-dimensional shape parameters and pose parameters of each face, and initialize the intrinsic and extrinsic parameters of the multi-camera system. S23: Construct a joint optimization objective function, and simultaneously optimize the intrinsic and extrinsic parameters of the multi-camera system and the model parameters of each face by minimizing the joint optimization objective function to obtain the optimized camera parameters; S24: Based on the optimized camera parameters, perform 3D reconstruction on the acquired multi-view image sequence to generate an initial 3D face point cloud.

5. The face multi-camera reconstruction method according to claim 4, characterized in that, The joint optimization objective function includes a geometric reprojection error term and a semantic consistency error term.

6. The face multi-camera reconstruction method according to claim 4, characterized in that, S24 includes: S241: Based on the optimized camera intrinsic and extrinsic parameters, perform disparity stereo matching processing on the acquired multi-view image sequence, extract the pixel correspondence of the same facial feature point in different view images, and generate a dense disparity map. S242: Based on the dense disparity map results and combined with the intrinsic and extrinsic parameters of the multi-camera system, the spatial coordinates of the matched pixels are solved using the multi-view triangulation method to reconstruct the corresponding three-dimensional point cloud data and generate an initial three-dimensional face point cloud model.

7. The face multi-camera reconstruction method according to claim 2, characterized in that, S3 includes: S31: The semantic labels of the dynamic facial regions identified by the semantic segmentation network are projected from the two-dimensional image space onto the initial three-dimensional facial point cloud, and a semantic partition label is assigned to each three-dimensional point. The semantic partition includes at least dynamic regions and static regions. S32: Assign fusion weights to each 3D point based on the semantic partition labels; S33: Fuse point cloud data generated from multiple camera perspectives. During the fusion process, for overlapping 3D points at the same location in space, perform a weighted average based on their respective fusion weights to obtain more stable 3D point coordinates. S34: Perform surface reconstruction and meshing on the fused point cloud to generate a high-precision 3D face model.

8. The face multi-camera reconstruction method according to claim 7, characterized in that, S32 includes: S321: The weights of the three-dimensional points in the dynamic region are assigned based on the motion amplitude of the region to which they belong; the greater the motion amplitude, the lower the weight. S322: Weighting of 3D points in static regions is based on the stability of local surface curvature and the consistency of color and texture.

9. The face multi-camera reconstruction method according to claim 7, characterized in that, The surface reconstruction is processed using the Poisson surface reconstruction algorithm, which treats the normal information of each 3D point in the fused point cloud as the gradient field of an implicit function, and generates a continuous and smooth 3D face surface by solving the corresponding Poisson equation.

10. A face multi-camera reconstruction system, used to implement the face multi-camera reconstruction method as described in any one of claims 1-9, characterized in that, Includes the following modules: Dynamic Region Recognition and Acquisition Control Module: Deploys a semantic segmentation network to perform real-time pixel-level segmentation of the input video stream, recognizes the dynamic regions of the face, and generates acquisition commands accordingly to control the multi-camera system to perform multi-view synchronous image acquisition; Semantic calibration and initial reconstruction module: Using a multi-face 3D deformable model as a calibration target, the multi-camera system is jointly calibrated and optimized, and 3D reconstruction is performed on the acquired images based on the optimized camera parameters to generate an initial 3D face point cloud; Weighted fusion and fine modeling module: It fuses the initial 3D face point cloud with the semantic information of the dynamic face region, performs feature weighting processing according to the region characteristics, suppresses dynamic noise, enhances static details, and then completes the fusion of 3D point cloud and surface modeling, outputting a high-precision 3D face model.

Citation Information

Cited By

  • Dynamic face reconstruction method and device based on partial differential equation

    CN121982224A