A method for real-time rendering and dynamic interaction of virtual masks based on visual feedback

CN122574334APending Publication Date: 2026-08-14JIANGXI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-20
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0003]在图像驱动过程中,硬件采集设备与图像捕获精度存在局限,软件控制方法与网格重构机制同样面临技术瓶颈,例如,公开号为CN113674385A的中国发明专利申请公开了一种虚拟表情生成方法、装置、电子设备及存储介质,利用神经网络模型处理中立表情网格拓扑边特征,预测待生成虚拟表情的关键点数据偏移量,这类纯数据驱动控制方法应对高频大尺度动态交互场景时存在原理性局限,过度依赖离线训练静态映射关系,缺乏面向实时视觉反馈的几何刚度约束机制,无法在异构拓扑映射导致的非等位同构冲突下动态完成切空间法向量原位重建,将算法预测脱离三维实体物理反射特性,在动态形变极限工况下极易在网格边界产生多边形面片自交与拉伸畸变,破坏表面法线场切空间连续性,由于图形渲染管线依赖网格曲面切空间几何连续性,三维模型表面法向量直接决定光照着色计算中的漫反射与镜面反射分配,在具体工程实践中,当面部驱动源拓扑与目标异型虚拟模型网格存在非等位同构冲突时,直接采用特征点线性位移变化率作为权重去计算不规则三角网顶点坐标,必然引发局部多边形面片交叉自交,此空间非线性拓扑畸变破坏表面法线平滑度,产生突变的离散法向量数据,导致光影传导路径在渲染级联中发生偏折,最终在图像帧的褶皱与边缘区域引起明暗断层

Benefits of technology

1、在虚拟面具实时渲染与动态交互中,通过建立面部特征标记点与不规则三角网顶点之间的空间射影重心场传导机制,系统将人类面部的平滑表情变化传导至具有异型几何结构的模型网格;该机制通过初始几何曲率标定的局部几何曲率保持刚度矩阵,针对面部高隆起与深凹陷区域配置差异化刚度系数,并利用稀疏矩阵乘法寻优计算对抗非线性形变产生的网格自交损耗;这种曲率自适应缩放制约方式使不规则三角网各顶点在形变时维持原位曲率,避免多边形面片交叉重叠与局部纹理拉伸撕裂,确立网格曲面的几何完整性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122574334A_ABST
    Figure CN122574334A_ABST
Patent Text Reader

Abstract

This invention relates to the field of image data processing technology and discloses a method for real-time rendering and dynamic interaction of virtual masks based on visual feedback. The method includes: extracting the two-dimensional coordinates of key facial features in a continuous image; interpolating the mesh vertex position vectors by combining the key point coordinates with the local projective centroid coordinates; scaling the position vectors using a local geometric curvature-preserving stiffness matrix to generate deformable mask vertex coordinates; calculating the cross product of the edge vectors of the deformable mesh patches to determine the dynamic surface normal vectors; and reconstructing and outputting the surface normal vector field by combining the weighted average of adjacent patch areas. This invention utilizes a stiffness matrix to configure stiffness coefficients for surfaces with differentiated curvatures to counteract mesh self-intersection, maintain mesh geometric integrity, and balances the light and shadow transmission path based on the surface normal vector field to eliminate edge optical artifacts and maintain the physical light and shadow reflection of the material surface.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method for real-time rendering and dynamic interaction of virtual masks based on visual feedback, belonging to the field of image data processing technology. Background Technology

[0002] Currently, the mainstream technical approach is to use image sensors to capture dynamic image sequences of faces and drive the virtual space three-dimensional geometric model to generate synchronous surface deformation. This approach uses facial key point fitting algorithms to extract the two-dimensional coordinates of facial feature points in the image stream, and establishes a geometric mapping between feature points and target mesh vertices through geometric space region control projection, thereby driving the digital character to generate mesh deformation accompanied by the displacement of feature points.

[0003] In image-driven processes, hardware acquisition devices and image capture accuracy have limitations, and software control methods and mesh reconstruction mechanisms also face technical bottlenecks. For example, Chinese invention patent application CN113674385A discloses a virtual expression generation method, device, electronic device, and storage medium. It uses a neural network model to process the topological edge features of a neutral expression mesh and predicts the key point data offsets of the virtual expression to be generated. This type of pure data-driven control method has fundamental limitations when dealing with high-frequency, large-scale dynamic interaction scenarios. It relies excessively on offline training of static mapping relationships and lacks a geometric stiffness constraint mechanism for real-time visual feedback. It cannot dynamically complete the in-situ reconstruction of the tangent space normal vector under the non-equipotential isomorphic conflict caused by heterogeneous topological mapping, thus detaching the algorithm prediction from the physical reflection of three-dimensional entities. Under dynamic deformation extreme conditions, polygonal facets are prone to self-intersection and stretching distortion at mesh boundaries, disrupting the continuity of the surface normal field tangent space. Since the graphics rendering pipeline relies on the geometric continuity of the mesh surface tangent space, the surface normal vector of the 3D model directly determines the allocation of diffuse and specular reflection in the lighting shading calculation. In specific engineering practice, when there is a non-equipotential isomorphic conflict between the topology of the face driving source and the target irregular virtual model mesh, directly using the linear displacement change rate of feature points as the weight to calculate the vertex coordinates of the irregular triangular mesh will inevitably cause local polygonal facets to intersect. This nonlinear topological distortion in space disrupts the smoothness of the surface normal, producing abrupt discrete normal vector data, causing the light and shadow transmission path to deflect in the rendering cascade, and ultimately causing light and shadow discontinuities in the wrinkles and edge areas of the image frame.

[0004] Therefore, the technical problem to be solved by this invention is how to avoid polygonal facet intersections during irregular triangular mesh deformation by using curvature adaptive stiffness constraints, and how to complete the in-situ reconstruction of tangent space normal vectors in geometric deformation to maintain physical light and shadow transmission calculations. Summary of the Invention

[0005] To address the problems in the background art, the technical solution of the present invention is as follows: A method for real-time rendering and dynamic interaction of virtual masks based on visual feedback, comprising the following steps: Step S1: Receive continuous image and video stream data, and load a cascaded classifier model to detect rectangular face regions in the continuous image and video stream data; Step S2: Use the key point detector to extract the two-dimensional coordinates of the facial key points within the rectangular area of ​​the face; Step S3: Obtain the initial three-dimensional mesh of the mask model, calculate the real-time two-dimensional displacement of the control polygon vertices based on the two-dimensional coordinates of the key points of the facial features, and use the reference local projective centroid coordinates of the control polygon vertices as interpolation weights to linearly interpolate and calculate the preliminary three-dimensional position vectors of the vertices of the initial three-dimensional mesh. Step S4: Retrieve the local geometric curvature preservation stiffness matrix calibrated according to the initial surface geometric curvature of the mask model, input the initial three-dimensional position vector into the local geometric curvature preservation stiffness matrix to perform matrix multiplication, scale the displacement of the initial three-dimensional mesh vertices, and generate the three-dimensional coordinates of the deformed mask vertices. Step S5: Traverse the irregular triangular facets in the mesh constructed from the three-dimensional coordinates of the vertices after deformation, and calculate the cross product of the two edge vectors that share common mesh vertices in the irregular triangular facets to determine the dynamic surface normal vector of the irregular triangular facets. Step S6: Retrieve all adjacent irregular triangular faces that share the same grid vertex, calculate the average value of the dynamic face normal vector of each adjacent irregular triangular face based on the face area weight, reconstruct the surface normal vector field, and output the surface normal vector field and the three-dimensional coordinates of the deformed face vertex to the display terminal.

[0006] Preferably, step S4, which retrieves the local geometric curvature preservation stiffness matrix calibrated based on the initial surface geometric curvature of the mask model, includes the following sub-steps: Step S41, obtaining the geometric curvature of the surface region of the mask model; when the surface bulge is greater than 2.5mm, configuring a first stiffness coefficient for the surface region; Step S42, when the surface bulge is not greater than 2.5mm, configuring a second stiffness coefficient for the surface region, wherein the first stiffness coefficient is greater than the second stiffness coefficient; Step S43, integrating the first stiffness coefficient and the second stiffness coefficient to construct the local geometric curvature preservation stiffness matrix.

[0007] Preferably, the following sub-steps are included before step S1: step S11, converting the received continuous image video stream data into a grayscale image; step S12, denoising the grayscale image using Gaussian filtering, and inputting the denoised grayscale image as continuous image video stream data into the cascaded classifier model.

[0008] Preferably, step S2, which calls the keypoint detector to extract the two-dimensional coordinates of facial key points within the rectangular area of ​​the face, includes the following sub-steps: Step S21, initialize the shape residual vector within the rectangular area of ​​the face; Step S22, call the keypoint detector to perform iterative approximation fitting on the shape residual vector based on local binary features, and stop iterating and output the two-dimensional coordinates of the facial key points when the coordinate increment between adjacent iteration steps is less than 0.05 pixels.

[0009] Preferably, before using the reference local projective barycentric coordinates of the control polygon vertices as interpolation weights in step S3, the following sub-steps are included: Step S31, marking the control polygon composed of multiple core control points on the initial 3D mesh of the mask model; Step S32, calculating the projective invariant of each non-control point mesh vertex in the initial 3D mesh with respect to the neighboring control polygon vertices, generating the reference local projective barycentric coordinates, and using the reference local projective barycentric coordinates as the basis for the spatial topology mapping of subsequent linear interpolation.

[0010] Preferably, the displacement of the initial three-dimensional mesh vertex in step S4 includes the following limitation: the three-dimensional displacement vector of the initial three-dimensional mesh vertex is equal to the product of the curvature adaptive scaling coefficient and the corresponding displacement component in the initial three-dimensional position vector; wherein, the curvature adaptive scaling coefficient is determined based on the displacement ratio change after matrix multiplication by the local geometric curvature preserving stiffness matrix.

[0011] Preferably, step S5, which involves calculating the cross product of two edge vectors sharing a common mesh vertex in the irregular triangular facet and determining the dynamic surface normal vector of the irregular triangular facet, includes the following sub-steps: Step S51, extracting the first vertex, the second vertex, and the third vertex of the irregular triangular facet; Step S52, constructing a first edge vector by pointing from the first vertex to the second vertex, and constructing a second edge vector by pointing from the first vertex to the third vertex; Step S53, calculating the cross product of the first edge vector and the second edge vector and performing normalization to generate the dynamic surface normal vector of the irregular triangular facet.

[0012] Preferably, the calculation of the dynamic surface normal vector of each adjacent irregular triangular facet in step S6, based on the weighted average of the facet areas, includes the following sub-steps: Step S61, count the total number of all adjacent irregular triangular facets sharing the same mesh vertex; Step S62, obtain the surface geometric area of ​​each adjacent irregular triangular facet, and use the proportion of each surface geometric area to the total area as a weight coefficient; Step S63, multiply the dynamic surface normal vector of each adjacent irregular triangular facet by the corresponding weight coefficient and sum them to complete the in-situ reconstruction of the surface normal vector field of the mesh surface.

[0013] Preferably, step S6, which outputs the surface normal field and the three-dimensional coordinates of the deformed mask vertices to the display terminal, includes the following sub-steps: Step S64, inputting the in-situ reconstructed surface normal field and the three-dimensional coordinates of the deformed mask vertices to the real-time rendering pipeline; Step S65, providing tangent space parameters with continuous geometric features for the bidirectional reflection distribution function of the mask model's material based on the surface normal field, maintaining the physical light and shadow reflection of the material surface during mesh dynamic deformation, and finally outputting it to the display terminal.

[0014] Preferably, in step S6, when outputting the surface normal field and the three-dimensional coordinates of the deformed face vertex to the display terminal, the output frame rate is controlled to be no less than 30 frames / second, so that the output graphic data is stabilized within the preset refresh rate range during the dynamic deformation of the mesh.

[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. In real-time rendering and dynamic interaction of virtual masks, by establishing a spatial projective centroid field transmission mechanism between facial feature markers and irregular triangular mesh vertices, the system transmits smooth facial expression changes to the model mesh with irregular geometric structures. This mechanism maintains the stiffness matrix of local geometric curvature calibrated by the initial geometric curvature, configures differentiated stiffness coefficients for high-protrusion and deep-concave areas of the face, and uses sparse matrix multiplication to optimize the calculation to counteract mesh self-crossing loss caused by nonlinear deformation. This curvature adaptive scaling constraint method enables each vertex of the irregular triangular mesh to maintain its in-situ curvature during deformation, avoids polygonal facet intersection and local texture stretching and tearing, and establishes the geometric integrity of the mesh surface.

[0016] 2. By traversing the deformed irregular triangular mesh facets and performing cross product operations on the connected edge vectors, the system calculates the dynamic surface normal vector of each triangular facet. This mechanism retrieves adjacent triangular facets that share the same mesh vertex and performs a mean calculation based on the facet area weighting of the dynamic surface normal vectors of adjacent triangular facets to reconstruct the surface normal vector field of the mesh surface in situ. This surface normal vector field directly balances the light and shadow transmission path of the material surface, providing tangent space parameters with continuous geometric characteristics for the subsequent bidirectional reflection distribution function, eliminating edge optical artifacts and light and dark discontinuities caused by abrupt changes in the normal direction, and maintaining the true physical light and shadow reflection of the material surface during the dynamic deformation of the mesh.

[0017] 3. The system constructs a closed-loop processing pipeline including image acquisition, feature fitting, spatial mapping, and lighting rendering to decouple feature data extraction from local mesh deformation. Furthermore, due to the use of a grayscale image-based classifier for detection and iterative fitting of local binary features, the computational overhead of the pre-processing image is reduced. On this basis, in conjunction with the stiffness matrix maintained by the initial surface geometric curvature calibration, the computational load of each vertex of the irregular triangular mesh is greatly reduced. This allows the updated geometric vertex data and the reconstructed surface normal vectors to be continuously output to the display terminal at a rate of no less than 30 frames per second, eliminating screen tearing, discontinuous flickering of brightness and darkness, and texture stretching distortion during large-scale dynamic interaction. Attached Figure Description

[0018] Figure 1 This is a flowchart of the global method for real-time rendering and dynamic interaction of virtual masks with visual feedback according to the present invention. Figure 2 This is a sequence diagram showing the data flow and execution timing of each module during the dynamic interaction of the virtual mask in this invention.

[0019] The objectives, features, and advantages of this invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of this application are within the scope of protection of this application.

[0021] A method for real-time rendering and dynamic interaction of virtual masks based on visual feedback includes the following steps: Step S1: Receive continuous image and video stream data, and load a cascaded classifier model to detect rectangular face regions in the continuous image and video stream data; Step S2: Use the key point detector to extract the two-dimensional coordinates of the facial key points within the rectangular area of ​​the face; Step S3: Obtain the initial three-dimensional mesh of the mask model, calculate the real-time two-dimensional displacement of the control polygon vertices based on the two-dimensional coordinates of the key points of the facial features, and use the reference local projective centroid coordinates of the control polygon vertices as interpolation weights to linearly interpolate and calculate the preliminary three-dimensional position vectors of the vertices of the initial three-dimensional mesh. Step S4: Retrieve the local geometric curvature preservation stiffness matrix calibrated according to the initial surface geometric curvature of the mask model, input the initial three-dimensional position vector into the local geometric curvature preservation stiffness matrix to perform matrix multiplication, scale the displacement of the initial three-dimensional mesh vertices, and generate the three-dimensional coordinates of the deformed mask vertices. Step S5: Traverse the irregular triangular facets in the mesh constructed from the three-dimensional coordinates of the vertices after deformation, and calculate the cross product of the two edge vectors that share common mesh vertices in the irregular triangular facets to determine the dynamic surface normal vector of the irregular triangular facets. Step S6: Retrieve all adjacent irregular triangular faces that share the same grid vertex, calculate the average value of the dynamic face normal vector of each adjacent irregular triangular face based on the face area weight, reconstruct the surface normal vector field, and output the surface normal vector field and the three-dimensional coordinates of the deformed face vertex to the display terminal.

[0022] Preferably, step S4, which retrieves the local geometric curvature preservation stiffness matrix calibrated based on the initial surface geometric curvature of the mask model, includes the following sub-steps: Step S41, obtaining the geometric curvature of the surface region of the mask model; when the surface bulge is greater than 2.5mm, configuring a first stiffness coefficient for the surface region; Step S42, when the surface bulge is not greater than 2.5mm, configuring a second stiffness coefficient for the surface region, wherein the first stiffness coefficient is greater than the second stiffness coefficient; Step S43, integrating the first stiffness coefficient and the second stiffness coefficient to construct the local geometric curvature preservation stiffness matrix.

[0023] Preferably, the following sub-steps are included before step S1: step S11, converting the received continuous image video stream data into a grayscale image; step S12, denoising the grayscale image using Gaussian filtering, and inputting the denoised grayscale image as continuous image video stream data into the cascaded classifier model.

[0024] Preferably, step S2, which calls the keypoint detector to extract the two-dimensional coordinates of facial key points within the rectangular area of ​​the face, includes the following sub-steps: Step S21, initialize the shape residual vector within the rectangular area of ​​the face; Step S22, call the keypoint detector to perform iterative approximation fitting on the shape residual vector based on local binary features, and stop iterating and output the two-dimensional coordinates of the facial key points when the coordinate increment between adjacent iteration steps is less than 0.05 pixels.

[0025] Preferably, before using the reference local projective barycentric coordinates of the control polygon vertices as interpolation weights in step S3, the following sub-steps are included: Step S31, marking the control polygon composed of multiple core control points on the initial 3D mesh of the mask model; Step S32, calculating the projective invariant of each non-control point mesh vertex in the initial 3D mesh with respect to the neighboring control polygon vertices, generating the reference local projective barycentric coordinates, and using the reference local projective barycentric coordinates as the basis for the spatial topology mapping of subsequent linear interpolation.

[0026] Preferably, the displacement of the initial three-dimensional mesh vertex in step S4 includes the following limitation: the three-dimensional displacement vector of the initial three-dimensional mesh vertex is equal to the product of the curvature adaptive scaling coefficient and the corresponding displacement component in the initial three-dimensional position vector; wherein, the curvature adaptive scaling coefficient is determined based on the displacement ratio change after matrix multiplication by the local geometric curvature preserving stiffness matrix.

[0027] Preferably, step S5, which involves calculating the cross product of two edge vectors sharing a common mesh vertex in the irregular triangular facet and determining the dynamic surface normal vector of the irregular triangular facet, includes the following sub-steps: Step S51, extracting the first vertex, the second vertex, and the third vertex of the irregular triangular facet; Step S52, constructing a first edge vector by pointing from the first vertex to the second vertex, and constructing a second edge vector by pointing from the first vertex to the third vertex; Step S53, calculating the cross product of the first edge vector and the second edge vector and performing normalization to generate the dynamic surface normal vector of the irregular triangular facet.

[0028] Preferably, the calculation of the dynamic surface normal vector of each adjacent irregular triangular facet in step S6, based on the weighted average of the facet areas, includes the following sub-steps: Step S61, count the total number of all adjacent irregular triangular facets sharing the same mesh vertex; Step S62, obtain the surface geometric area of ​​each adjacent irregular triangular facet, and use the proportion of each surface geometric area to the total area as a weight coefficient; Step S63, multiply the dynamic surface normal vector of each adjacent irregular triangular facet by the corresponding weight coefficient and sum them to complete the in-situ reconstruction of the surface normal vector field of the mesh surface.

[0029] Preferably, step S6, which outputs the surface normal field and the three-dimensional coordinates of the deformed mask vertices to the display terminal, includes the following sub-steps: Step S64, inputting the in-situ reconstructed surface normal field and the three-dimensional coordinates of the deformed mask vertices to the real-time rendering pipeline; Step S65, providing tangent space parameters with continuous geometric features for the bidirectional reflection distribution function of the mask model's material based on the surface normal field, maintaining the physical light and shadow reflection of the material surface during mesh dynamic deformation, and finally outputting it to the display terminal.

[0030] Preferably, in step S6, when outputting the surface normal field and the three-dimensional coordinates of the deformed face vertex to the display terminal, the output frame rate is controlled to be no less than 30 frames / second, so that the output graphic data is stabilized within the preset refresh rate range during the dynamic deformation of the mesh.

[0031] Example 1: In an image processing scenario involving the virtualized display of the intangible cultural heritage Nuo opera, the image processing system receives a continuous image and video stream containing real-time facial dynamics of the user. A cascaded classifier model is loaded to extract a rectangular region of the face from the continuous image and video stream. The system calls a keypoint detector to locate multiple facial keypoints within the rectangular region of the face, calculates the two-dimensional coordinate set of each keypoint, and uses this as the geometric driving input for the subsequent dynamic deformation of the virtual mask. To achieve accurate alignment between the two-dimensional facial features and the heterogeneous three-dimensional mask control polygons, the system establishes an explicit index topological mapping table during the initialization phase. This mapping table aligns the extracted 68 facial keypoints according to anatomical... The features are divided into five independent feature regions: eyebrows, eyes, nose, mouth, and facial contours. A one-to-one index association is established between the 2D coordinates of each key point and the corresponding control vertex of the feature region on the mask model's control polygon. During runtime, the system dynamically captures the user's facial midline and performs normalized rigid alignment to eliminate interference from overall head rotation and displacement. Then, the normalized 2D coordinates of any given facial key point are directly assigned to its associated mask control vertex, thus losslessly transferring real-time facial deformation features to the horizontal and vertical coordinates of the control polygon vertices. This achieves precise geometric driving across topologies when constructing the virtual mask's rendering architecture. The system pre-defines a 3D mask model, which consists of multiple triangular facets with a pre-defined initial surface geometric curvature distribution. The system calibrates the core control points of the mask model according to pre-defined rules, establishes a feature skeleton composed of multiple control polygons, calculates and stores the reference local projective barycentric coordinates of each mesh vertex relative to the control polygon vertices, and updates the 2D coordinates of the feature marker points when facial expressions change in the video stream. It calculates the real-time 2D displacement of the control polygon vertices through geometric projection mapping, and uses the reference local projective barycentric coordinates as interpolation weights to linearly interpolate and generate the preliminary 3D position vectors of each vertex of the 3D mask model. In this step of cross-dimensional calculation, in order to accurately convert the real-time two-dimensional displacement belonging to the two-dimensional image plane into a position vector in three-dimensional space, the system obtains the orthogonal projection matrix of the initial three-dimensional mesh under the two-dimensional camera plane. The displacement increment in the two-dimensional plane is mapped to the tangent plane of the tangent space of the initial three-dimensional mesh surface using the orthogonal projection matrix. The reference local projective centroid coordinates corresponding to each non-control point mesh vertex are used as discrete weight coefficients and weighted summed with the mapped displacement of multiple control polygon vertices adjacent to the vertex in the tangent plane. Thus, the displacement components of the non-control point mesh vertex in the horizontal and vertical directions are superimposed and calculated in the linear space.Simultaneously, to recover the missing depth direction displacement, this invention projects the calculated two-dimensional displacement in the tangent plane along the corresponding normal direction based on the normal direction vectors of each vertex of the initial surface mesh of the mask model, calculating the real-time displacement increment in the depth direction. Finally, the three-dimensional displacement increments along the horizontal axis, vertical axis, and depth direction are vector-superimposed with the reference three-dimensional coordinates of the initial three-dimensional mesh vertices, thereby generating a preliminary three-dimensional position vector that possesses both three-dimensional spatial scale attributes.

[0032] To prevent mesh self-intersections in local areas during mesh deformation, the system retrieves a local geometric curvature preservation stiffness matrix pre-calibrated based on the initial surface geometric curvature of the 3D mask model. This matrix assigns differentiated stiffness coefficients based on the surface bulge of the mask model, with higher stiffness coefficients for areas with bulges greater than 2.5 mm and lower stiffness coefficients for flat areas. In actual construction, the local geometric curvature preservation stiffness matrix is ​​a sparse square matrix, with both the number of rows and columns equal to the total number of vertices in the initial 3D mesh. The diagonal elements of this stiffness matrix are filled with either the first or second stiffness coefficient, while the off-diagonal elements are filled based on the spacing between mesh vertices. The topological connectivity is determined. If two mesh vertices share the same mesh edge, their corresponding off-diagonal elements are set to a negative mean stiffness coefficient; otherwise, they are set to 0. The system converts the initial 3D position vector into a one-dimensional column vector and performs standard matrix multiplication with this matrix. The calculated output column vector represents the corrected coordinates after structural stiffness constraints. Then, the corrected coordinates are divided point-by-point by the corresponding components of the initial 3D position vector to calculate the coordinate change ratio of each vertex along each axis in 3D space. This coordinate change ratio is the displacement ratio change. The system sets a baseline smoothing mapping function to limit the displacement ratio change to a safe deformation between 0 and 1. Within the specified interval, when the displacement ratio exceeds the safe range, the system directly truncates and takes the boundary value. The value determined here is the curvature adaptive scaling coefficient. The system inputs the initial three-dimensional position vector into the local geometric curvature preservation stiffness matrix and performs sparse matrix multiplication optimization calculation. The displacement of each mesh vertex is constrained by curvature adaptive scaling to obtain the three-dimensional coordinates of the deformed vertex after deformation, avoiding mesh self-intersection. In this operation, in order to eliminate the mechanical inconsistency of directly multiplying the structural stiffness matrix with the absolute geometric coordinates, the system introduces a reference position transformation step before performing matrix multiplication, that is, subtracting the initial three-dimensional mesh vertex static reference position coordinates from the input initial three-dimensional position vector. The absolute position coordinates are transformed into relative displacement vectors of mesh vertices representing the degree of pure geometric deformation. This relative displacement vector is then input as a column vector of nodal loads into the stiffness matrix equation and multiplied with the local geometric curvature-preserving stiffness matrix. Essentially, this utilizes the curvature characteristic operator in the stiffness matrix as a spatial constraint boundary to solve for the energy minimization constraint of local shear and tensile deformations of the discrete mesh, calculating the equilibrium displacement increment that can resist local self-intersection. Finally, this equilibrium displacement increment is superimposed back onto the static reference position coordinates of the initial three-dimensional mesh vertices, thereby synthesizing the output three-dimensional coordinates of the deformed vertices that avoid mesh self-intersection and satisfy physical stiffness continuity.

[0033] The system traverses the deformed mesh surface, performs cross product operations on adjacent edge vectors of irregular triangular facets to determine the dynamic surface normal vectors of each facet. The system retrieves all adjacent triangular facets sharing the same mesh vertex, performs weighted average calculation based on the area of ​​each adjacent facet, and dynamically balances and reconstructs the surface normal vector field. The rendering engine uses the updated vertex 3D coordinates and the reconstructed surface normal vector field to calculate the bidirectional reflection distribution function. While maintaining the geometric continuity of the mask material, the processed virtual mask rendering image is output to the display terminal at a rate of 30 frames per second. During the bidirectional reflection distribution function calculation, in order to fully construct the tangent space parameters required for the rendering pipeline using the surface normal vector field, the system employs an in-situ basis orthogonalization algorithm to solve for the missing tangent and bitangent vectors. Specifically, the system extracts the initial state of the mask model... The preset 2D material texture coordinate axis is used as the reference initial tangent vector, with the horizontal growth direction of the 2D texture coordinate as the reference. Using the Schmidt orthogonalization method, the reference initial tangent vector is projected onto the plane perpendicular to the dynamically reconstructed surface normal vector, and its component in the normal vector direction is subtracted. After normalization, the real-time unit tangent vector that is absolutely orthogonal to the dynamic normal vector is obtained. Finally, the real-time unit normal vector and the real-time unit tangent vector are cross-producted to directly calculate the mutually perpendicular third axis vector, i.e., the unit secondary tangent vector. The three-dimensional orthogonal coordinate system composed of the real-time unit normal vector, the real-time unit tangent vector, and the unit secondary tangent vector is the continuous geometric tangent space. The parameters of this tangent space are directly input into the shader of the bidirectional reflection distribution function to accurately correct the micro-geometric diffuse reflection and specular reflection energy distribution of light on complex deformed surfaces.

[0034] Example 2: In the virtualized display of Nuo opera, an intangible cultural heritage, which requires real-time rendering based on user expressions, the image processing system acquires continuous image and video streams at a frequency of 30Hz. A pre-trained cascaded classifier model is used to detect and select rectangular areas of the face in real time. The system calls a keypoint detector to locate 68 key facial features within the rectangular area, extracts and stores the real-time two-dimensional coordinates of these key facial features. To verify the response effect of the method to different expression intensities and the feasibility of curvature adaptive deformation, three controlled experimental conditions were designed. The first experimental group was a conventional rendering experiment based on moderate-amplitude expression changes; the second experimental group was an extreme condition experiment based on large-scale expression changes and high-frequency expression switching; and the third control group was a comparative experiment involving removing local geometric curvature while maintaining the stiffness matrix constraint. In the first experimental group, when the system detected a 15-pixel displacement of the keypoints caused by the user's mouth opening action, the mesh vertex vector was calculated by linear interpolation of the reference local projective centroid coordinates. Experimental monitoring data showed that the maximum displacement deviation of the mesh vertices was 0.2mm at this time, and the surface mesh did not self-intersect.

[0035] In the second set of extreme working condition test samples, the user's large-scale facial expression switching caused the two-dimensional displacement of key points to increase instantaneously to 45 pixels. To avoid nonlinear deformation causing facet intersections, the system used a local geometric curvature maintaining stiffness matrix for sparse matrix multiplication optimization calculation. Test monitoring showed that when the local bulge in the nose bridge area of ​​the mask model reached 3.8mm, the high stiffness coefficient configured in this area effectively suppressed the displacement redundancy of the vertices, keeping the minimum spatial gap between mesh faces above 0.15mm, thus establishing geometric integrity. The aforementioned 2.5mm bulge threshold is an engineering critical value determined based on high-precision facial 3D scan statistics and finite element deformation simulation limits. In multiple preliminary experimental measurements of this invention, when the local surface of the mesh model... When the bulge height is less than or equal to 2.5mm, the normal component of the curvature of the irregular triangular facets is insufficient to cause facet intersection when driven by large-scale displacement within 45 pixels. The mesh itself has sufficient geometric tolerance space. At this time, configuring a smaller second stiffness coefficient can ensure that the mesh maintains good flexibility and facial detail. However, once the local bulge height exceeds the critical point of 2.5mm, the nonlinear spatial curvature of the mesh will be amplified sharply. Under large-scale displacement, the edges of adjacent triangular facets are very likely to penetrate each other and be severely stretched, resulting in the self-intersection rate soaring to over 18.5%. Therefore, a larger first stiffness coefficient is forcibly switched to forcibly block the excessive displacement of vertices by artificially increasing the rigidity constraint, ensuring the integrity of the mesh topology under extreme conditions.

[0036] In the third control group, the system removed the constraint of the local geometric curvature preserving stiffness matrix. Under the same 45-pixel displacement input, the mask model underwent geometric self-intersection in the nose bridge and corners of the mouth, and the polygonal patches overlapped in space, resulting in obvious optical artifacts and black spots at the edge of the nose bridge in the rendered image. Comparing the experimental data, the average normal vector reconstruction error of the first and second test groups was 0.02 units, while the normal vector error of the control group surged to 0.45 units due to the failure of normal vector field reconstruction calculation caused by mesh self-intersection. When the surface bulge of the mask was higher than 2.5 mm, the stiffness adaptive matrix was enabled to reduce the patch self-intersection rate from 18.5% in the control group to below 0.3%. The experimental results confirmed that the system avoided mesh self-intersection through curvature adaptive stiffness constraint and balanced the light and shadow transmission path based on the surface normal vector field, so that the processed virtual mask maintained stable physical light and shadow reflection under different expression amplitudes. After completing the image data processing, the system outputs the image data stream at a rate of 30 frames per second.

[0037] Example 3: In digital rendering scenarios for high-precision facial expression capture and real-time resetting, the system suffers from visual jitter caused by instantaneous high-frequency noise generated by the coordinates of user facial feature markers, which causes the displacement fluctuation of virtual mask mesh vertices to exceed the preset stability threshold. To eliminate this noise source from the underlying logic and achieve anti-jitter rendering, after obtaining the real-time two-dimensional coordinate set of key facial features, the system does not directly perform the preliminary three-dimensional position vector calculation of mesh vertices. Instead, it inputs the extracted two-dimensional coordinate sequence into a weighted sliding window buffer with transient error filtering capabilities.

[0038] The buffer maintains a sliding time window with a length of 5 sampling periods. In each sampling period, the system normalizes the coordinates of all marked points within the window. When the movement increment of any point within the window relative to the previous sampling period is greater than a preset noise threshold, the system activates error suppression logic. This logic calculates the mean Euclidean distance between the current coordinate and the coordinates of the other 4 sampling points within the window, removes abnormal sampling points with a deviation exceeding 3 times the mean distance, and calculates a smoothed two-dimensional coordinate input value by using the weighted arithmetic mean of the remaining valid sampling point coordinates. This smoothed two-dimensional coordinate input value serves as the sole basis for calculating the two-dimensional displacement of the control polygon vertices, thereby ensuring that the input data stream entering the subsequent interpolation solution stage has stable statistical characteristics.

[0039] To further verify the stability of the anti-shake mechanism under different expression switching speeds, the system set the input signal source to a synthetic noise sequence containing jitter of different frequencies. The experiment showed that when the video input signal generated random noise interference with a frequency of 10Hz, the above-mentioned error suppression logic based on the sliding window reduced the displacement jitter amplitude of the grid vertices from 1.2mm to 0.08mm. When the user performed a high-frequency head nodding action, the system did not experience inter-frame rendering jumps because the effective sampling point elimination mechanism avoided misjudgment. By embedding the noise suppression logic in the preprocessing stage after data acquisition, the system ensured that the real-time two-dimensional displacement of the control polygon vertices was within the preset spatial continuity constraint range. Based on the above processing path, when the rendering engine received the smoothed three-dimensional vertex position vector and the updated surface normal vector field, it was able to achieve visually stable output of the virtual mask edge contour while maintaining high-fidelity texture mapping, eliminating the grid vertex displacement fluctuations caused by high-frequency noise in the input signal.

[0040] Example 4: In the digital display of cultural heritage involving high-frequency dynamic facial expression capture and virtual mask rendering, the image processing system faces the technical challenge of rendering image flickering caused by fluctuations in ambient light intensity and jitter in facial key point detection. After loading a cascaded classifier model to detect the rectangular region of the face in a continuous image video stream and extracting the two-dimensional coordinates of the facial key points, the system inputs the set of two-dimensional coordinates of the facial key points into a weighted sliding window buffer with transient error filtering function. The weighted sliding window buffer is set as a sliding time window with a length of 5 sampling periods. In each sampling period, the system performs normalization on the coordinates of all marked points in the window. In the processing, when the system detects that the movement increment of any point coordinate within the window relative to the previous sampling period exceeds a preset noise amplitude threshold, error suppression logic is triggered. This logic calculates the mean Euclidean distance between the current sampling point and the coordinates of the other four sampling points within the window, eliminates abnormal sampling points with a deviation exceeding three times the mean distance, and calculates a smoothed two-dimensional coordinate input value by using the weighted arithmetic mean of the remaining valid sampling point coordinates. This smoothed two-dimensional coordinate input value serves as the sole basis for calculating the two-dimensional displacement of the control polygon vertices, thereby ensuring that the evolution trajectory of each vertex position vector in the three-dimensional mask model has stable statistical characteristics during mesh deformation.

[0041] The system inputs the smoothed 3D vertex position vectors into a sparse matrix multiplication optimization calculation using a local geometric curvature-preserving stiffness matrix pre-calibrated based on the initial surface geometric curvature of the mask model. The system then uses this calculation result to adaptively scale the mesh vertex displacements. It establishes dynamic surface normal vectors by performing cross product operations on the face vectors of the deformed mesh and reconstructs the surface normal vector field based on the weighted average of adjacent face areas. The rendering engine receives the 3D coordinates of the deformed vertices and the reconstructed surface normal vector field at a rate of 30 frames per second, performs bidirectional reflection distribution function calculations on the virtual mask, and outputs continuously rendered image frames. Experimental data shows that, under Gaussian white noise interference with a signal-to-noise ratio of 20dB superimposed on the input video stream, this data processing path suppresses the displacement jitter amplitude of the mesh vertices from 1.2mm to 0.08mm, effectively eliminating mesh vertex displacement fluctuations caused by high-frequency noise in the input signal.

[0042] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0043] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A method for real-time rendering and dynamic interaction of virtual masks based on visual feedback, characterized in that, Includes the following steps: Step S1: Receive continuous image and video stream data, and load a cascaded classifier model to detect rectangular face regions in the continuous image and video stream data; Step S2: Use the key point detector to extract the two-dimensional coordinates of the facial key points within the rectangular area of ​​the face; Step S3: Obtain the initial three-dimensional mesh of the mask model, calculate the real-time two-dimensional displacement of the control polygon vertices based on the two-dimensional coordinates of the key points of the facial features, and use the reference local projective centroid coordinates of the control polygon vertices as interpolation weights to linearly interpolate and calculate the preliminary three-dimensional position vectors of the vertices of the initial three-dimensional mesh. Step S4: Retrieve the local geometric curvature preservation stiffness matrix calibrated according to the initial surface geometric curvature of the mask model, input the initial three-dimensional position vector into the local geometric curvature preservation stiffness matrix to perform matrix multiplication, scale the displacement of the initial three-dimensional mesh vertices, and generate the three-dimensional coordinates of the deformed mask vertices. Step S5: Traverse the irregular triangular facets in the mesh constructed from the three-dimensional coordinates of the vertices after deformation, and calculate the cross product of the two edge vectors that share common mesh vertices in the irregular triangular facets to determine the dynamic surface normal vector of the irregular triangular facets. Step S6: Retrieve all adjacent irregular triangular faces that share the same grid vertex, calculate the average value of the dynamic face normal vector of each adjacent irregular triangular face based on the face area weight, reconstruct the surface normal vector field, and output the surface normal vector field and the three-dimensional coordinates of the deformed face vertex to the display terminal.

2. The method for real-time rendering and dynamic interaction of virtual masks based on visual feedback according to claim 1, characterized in that, Step S4, which retrieves the local geometric curvature preservation stiffness matrix calibrated based on the initial surface geometric curvature of the mask model, includes the following sub-steps: Step S41, obtain the geometric curvature of the surface region of the mask model; when the surface bulge is greater than 2.5mm, configure a first stiffness coefficient for the surface region; Step S42, when the surface bulge is not greater than 2.5mm, configure a second stiffness coefficient for the surface region, wherein the first stiffness coefficient is greater than the second stiffness coefficient; Step S43, integrate and construct the local geometric curvature preservation stiffness matrix based on the first stiffness coefficient and the second stiffness coefficient.

3. The method for real-time rendering and dynamic interaction of virtual masks based on visual feedback according to claim 1, characterized in that, The following sub-steps are included before step S1: Step S11, converting the received continuous image video stream data into a grayscale image; Step S12, denoising the grayscale image using Gaussian filtering, and inputting the denoised grayscale image as continuous image video stream data into the cascaded classifier model.

4. The method for real-time rendering and dynamic interaction of virtual masks based on visual feedback according to claim 1, characterized in that, Step S2 involves calling the keypoint detector to extract the two-dimensional coordinates of facial key points within the rectangular area of ​​the face, which includes the following sub-steps: Step S21, initialize the shape residual vector within the rectangular area of ​​the face; Step S22, call the keypoint detector to perform iterative approximation fitting on the shape residual vector based on local binary features. When the coordinate increment between adjacent iteration steps is less than 0.05 pixels, stop the iteration and output the two-dimensional coordinates of the facial key points.

5. The method for real-time rendering and dynamic interaction of virtual masks based on visual feedback according to claim 1, characterized in that, Before using the reference local projective barycentric coordinates of the control polygon vertices as interpolation weights in step S3, the following sub-steps are included: Step S31, marking the control polygon composed of multiple core control points on the initial 3D mesh of the mask model; Step S32, calculating the projective invariant of each non-control point mesh vertex in the initial 3D mesh with respect to the neighboring control polygon vertices, generating the reference local projective barycentric coordinates, and using the reference local projective barycentric coordinates as the basis for the spatial topology mapping of subsequent linear interpolation.

6. The method for real-time rendering and dynamic interaction of virtual masks based on visual feedback according to claim 1, characterized in that, In step S4, the displacement of the initial three-dimensional mesh vertices is scaled with the following limitations: the three-dimensional displacement vector of the initial three-dimensional mesh vertex is equal to the product of the curvature adaptive scaling factor and the corresponding displacement component in the initial three-dimensional position vector; wherein, the curvature adaptive scaling factor is determined based on the displacement ratio change after matrix multiplication of the local geometric curvature preserving stiffness matrix.

7. The method for real-time rendering and dynamic interaction of virtual masks based on visual feedback according to claim 1, characterized in that, Step S5 involves calculating the cross product of two edge vectors sharing a common mesh vertex in the irregular triangular facet and determining the dynamic face normal vector of the irregular triangular facet. This includes the following sub-steps: Step S51: Extract the first vertex, second vertex, and third vertex of the irregular triangular facet; Step S52: Construct the first edge vector by pointing from the first vertex to the second vertex, and construct the second edge vector by pointing from the first vertex to the third vertex. Step S53: Calculate the cross product of the first and second side vectors and normalize them to generate the dynamic surface normal vector of the irregular triangular facet.

8. The method for real-time rendering and dynamic interaction of virtual masks based on visual feedback according to claim 1, characterized in that, Step S6 involves calculating the dynamic surface normal vector of each adjacent irregular triangular facet based on the facet area weighted average, which includes the following sub-steps: Step S61, count the total number of all adjacent irregular triangular facets sharing the same mesh vertex; Step S62, obtain the surface geometric area of ​​each adjacent irregular triangular facet, and use the proportion of each surface geometric area to the total area as a weighting coefficient; Step S63, multiply the dynamic surface normal vector of each adjacent irregular triangular facet by the corresponding weighting coefficient and sum them to complete the in-situ reconstruction of the surface normal vector field of the mesh surface.

9. The method for real-time rendering and dynamic interaction of virtual masks based on visual feedback according to claim 1, characterized in that, Step S6, which outputs the surface normal field and the three-dimensional coordinates of the deformed mask vertices to the display terminal, includes the following sub-steps: Step S64, inputting the in-situ reconstructed surface normal field and the three-dimensional coordinates of the deformed mask vertices to the real-time rendering pipeline; Step S65, providing tangent space parameters with continuous geometric features for the bidirectional reflection distribution function of the mask model's material based on the surface normal field, maintaining the physical light and shadow reflection of the material surface during dynamic mesh deformation and finally outputting it to the display terminal.

10. The method for real-time rendering and dynamic interaction of virtual masks based on visual feedback according to claim 1, characterized in that, In step S6, when the surface normal field and the three-dimensional coordinates of the deformed face vertex are output to the display terminal, the output frame rate is controlled to be no less than 30 frames / second, so that the output graphic data is stabilized within the preset refresh rate range during the dynamic deformation of the mesh.

Citation Information

Patent Citations

  • Virtual expression generation method and device, electronic equipment and storage medium

    CN113674385A