Human three-dimensional modeling method based on human image data scanning
By jointly optimizing the quantized frame quality weights and the adaptive deformation field, the geometric inconsistency problem caused by human body micro-movements in multi-view 3D reconstruction is solved, achieving high-precision and consistent 3D human body modeling.
Patent Information
- Application Number
- CN202511367824.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-24
- Publication Date
- 2026-02-06
- Estimated Expiration
- 2045-09-24
AI Technical Summary
Existing technologies in human 3D modeling based on multi-view image sequences fail to effectively utilize frame quality assessment data generated during frame screening, resulting in ghosting, blurring, and geometric distortion in the reconstructed model. Furthermore, the optimization process lacks a differentiated trust mechanism for data of different quality levels, which limits the improvement of model accuracy.
By generating two sets of quantized frame quality weights in parallel, and calculating the geometric consistency error of the pixel flow variation pattern and the 3D point cloud, a set of high-quality and highly consistent key frames is selected. An adaptive deformation field is constructed, and the energy function is jointly optimized by spatial consistency and temporal smoothness constraints to optimize the 3D geometric model.
It significantly improves the accuracy and consistency of 3D reconstruction, enhances the adaptability to dynamic human body scenes, outputs high-quality and highly reliable 3D models, and solves the geometric inconsistency problem caused by human body micro-movements.
Smart Images

Figure CN120876740B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of image analysis, in particular to a human three-dimensional modeling method based on human image data scanning BACKGROUND
[0002] When performing human three-dimensional modeling based on multi-view image sequences, the small involuntary motion of the scanned object is a long-standing technical challenge. Traditional static reconstruction algorithms ignore such temporal changes and directly treat all views as being collected from the same static moment, resulting in serious ghosting, blurring and geometric distortion in the reconstructed three-dimensional model. Some existing dynamic reconstruction methods attempt to address this problem by estimating global motion or segmenting static assumptions, but they usually treat frame screening and model optimization as two independent stages. That is, first, the frames with good quality are selected, and then these selected frames are equally used in the subsequent optimization, while the valuable data generated in the screening process, which quantifies the quality level of each frame, is discarded. This processing method fails to fully utilize the deep information extracted in the screening stage, resulting in a lack of differentiated trust mechanism for different quality data in the optimization process, ultimately limiting the further improvement of model accuracy. SUMMARY
[0003] The present application aims to provide a human three-dimensional modeling method based on human image data scanning to solve the problems raised in the background. The specific technical problems include how to convert the frame quality evaluation data generated in the key frame screening process into quantifiable weight factors, i.e., frame quality weights, and embed them into the subsequent joint optimization model to guide the algorithm to give higher confidence to high-quality and high-consistency data, thereby solving the geometric inconsistency problem of multi-view three-dimensional reconstruction caused by human micro-motion.
[0004] To achieve the above-mentioned purpose, the present application provides a human three-dimensional modeling method based on human image data scanning, which includes the following method steps:
[0005] S1, first, synchronously collect image sequences from multiple different angles around the human body; the purpose of this step is to obtain basic raw data covering 360 degrees of the human body surface and containing time dimension information, providing sufficient data basis for subsequent three-dimensional reconstruction; the effect is to ensure the completeness and time sequence of the reconstructed object information, which is a prerequisite for high-precision dynamic modeling.
[0006] S2, screen the collected human image sequences, the core of which is not only to select a key frame set of multiple views representing the relatively stable posture of the human body, but more importantly, to generate two sets of quantified frame quality weights (first frame quality weight and second frame quality weight) in parallel. The process of generating the key frame set specifically includes:
[0007] Two independent quantitative evaluation strategies are implemented in parallel, one of which is to analyze the pixel flow change pattern in the human image sequence, and the other is to calculate the geometric consistency error between the three-dimensional point cloud data generated by the human image sequence;
[0008] According to the quantitative results generated by the evaluation, a corresponding frame quality weight is assigned to each frame of the human image sequence; according to the size of the frame quality weight, the candidate image frames with the smallest motion amplitude and the highest geometric consistency are selected from the human image sequence of each view;
[0009] Based on the proximity of the timestamps and the coordination of the posture parameters between the candidate frames, a global optimization selection is performed to combine and construct an optimal multi-view key frame set; the process of global optimization selection specifically includes:
[0010] A filtering condition based on timestamp proximity is established, and candidate frame groups with timestamp differences exceeding the preset threshold are excluded from the consideration range;
[0011] The candidate frame groups filtered by time are evaluated for posture parameter coordination, and the similarity between the three-dimensional posture parameters corresponding to the candidate frames of each view is calculated to quantitatively evaluate the multi-view posture consistency;
[0012] Based on the dual indicators of timestamp proximity and posture parameter coordination, a multi-objective optimization function is constructed, and the optimal solution of the function is solved to determine the best key frame combination scheme as the optimal multi-view key frame set.
[0013] In addition, in the process of analyzing the pixel flow change pattern in the human image sequence, a first frame quality weight is also calculated based on the amplitude of the pixel flow change pattern, specifically including:
[0014] The human image sequence arranged in time sequence is processed frame by frame, and the motion vector field of dense pixels between adjacent image frames is calculated to obtain the pixel flow change pattern representing the motion state of each point in the picture; the calculation process of the motion vector field uses the optical flow method to solve the displacement vector of the pixel between adjacent frames;
[0015] The amplitude of the pixel flow change pattern is globally calculated to obtain a scalar indicator representing the global motion intensity of the entire image frame;
[0016] Based on the scalar indicator, the amplitude is mapped to a weight value using an inverse proportional function relationship to calculate the first frame quality weight.
[0017] The geometric consistency error between the three-dimensional point cloud data generated by the human image sequence is calculated, and the second frame quality weight is calculated based on the geometric consistency error value, specifically including:
[0018] performing a preliminary three-dimensional reconstruction operation on the input human image sequence to generate a time-series three-dimensional point cloud dataset corresponding to different time points; wherein the preliminary three-dimensional reconstruction operation recovers a three-dimensional point cloud from two-dimensional images using a multi-view stereo vision algorithm;
[0019] In the time dimension, each time-series three-dimensional point cloud data is registered with the point cloud data of the adjacent time point, and the residual error existing after registration is accurately calculated, which is defined as a geometric consistency error; wherein the registration operation uses an iterative closest point algorithm to solve the spatial transformation parameters between the point clouds;
[0020] Based on the geometric consistency error value, an inverse proportional function relationship is used to map the error value to a weight value, and a second frame quality weight is calculated.
[0021] The core role of step S2 is to convert the human micro-motion, a fuzzy problem, into a quantifiable weight factor. By parallel computing the pixel flow motion amplitude (to generate the first frame quality weight) and the three-dimensional point cloud geometric consistency error (to generate the second frame quality weight), and combining global optimization of time and posture to screen the key frame set, the effect is to output two types of key outputs, one is to identify the quantified weight of each frame image quality and three-dimensional consistency, and the other is to form a high-quality frame set with high self-consistency and good time synchronization; This provides an accurate basis for subsequent optimization algorithms and directly reduces geometric inconsistency from the data source.
[0022] S3, based on the key frame set, calculating an adaptive deformation field for compensating the deformation of the human body due to micro-motion, wherein the calculation process of the adaptive deformation field specifically includes:
[0023] Input the key frame set into an elastic deformation calculation model, which regards the human body as a continuous elastic medium, and describes the micro-motion deformation by establishing the displacement mapping relationship of the points on the surface of the human body;
[0024] Using a physics-based deformation model, a continuous adaptive deformation field is constructed by solving the displacement vector field of the control points; the adaptive deformation field is a three-dimensional vector field that defines a spatial mapping function from the actual posture at the acquisition time to the reference posture;
[0025] The calculation process of the deformation field is performed by iterative optimization, and each iteration adjusts the parameters of the deformation field to make the deformed three-dimensional model satisfy the multi-view consistency constraint.
[0026] An energy function is established for joint optimization, which simultaneously contains spatial consistency constraints and temporal smoothness constraints, and the optimization target of the spatial consistency constraint term is modulated by the first frame quality weight and the second frame quality weight, wherein the construction process of the joint optimization energy function specifically includes:
[0027] The spatial consistency constraint term ensures the consistency between the three-dimensional model and the two-dimensional image observation value by calculating the multi-view reprojection error;
[0028] The temporal smoothness constraint term ensures the spatio-temporal continuity of the deformation process by limiting the variation gradient of adjacent time sequence deformation fields;
[0029] The optimization target of the spatial consistency constraint term is modulated by the first frame quality weight and the second frame quality weight, and the modulation process is to combine the two weights into a comprehensive weight coefficient through weighted fusion, and the comprehensive weight coefficient is used as a weighted factor in the reprojection error calculation process in the spatial consistency constraint term;
[0030] The step S3 is to build a mathematical optimization framework that can be modulated by the aforementioned weights; by establishing a physically driven adaptive deformation field to model human micro-motions, and designing a joint optimization energy function containing spatial consistency constraints and temporal smoothness constraints; the key effect lies in that the two frame quality weights generated in step S2 are fused into a comprehensive weight coefficient, which is directly used as a modulation factor in the calculation of the spatial consistency constraint term (reprojection error); this makes the optimization process can be accurately guided, and the algorithm will preferentially meet the constraints from the high-quality and high-consistency key frames, thereby laying the optimization foundation for solving the geometric inconsistency problem.
[0031] S4, deforming and compensating the three-dimensional point cloud data by using the adaptive deformation field, synchronously optimizing the three-dimensional geometric model and the adaptive deformation field by minimizing the joint optimization energy function, and outputting the human three-dimensional model after optimization, specifically comprising:
[0032] The adaptive deformation field is used to deform and compensate the three-dimensional point cloud data, and the compensation process is to apply the spatial mapping function provided by the deformation field to the three-dimensional point cloud data at the acquisition time, and transform it to a unified reference posture space;
[0033] The three-dimensional geometric model and the adaptive deformation field are synchronously optimized by minimizing the joint optimization energy function, and the optimization process updates the vertex coordinates of the three-dimensional geometric model and the parameter vector of the deformation field at the same time by using an iterative optimization algorithm;
[0034] When the joint optimization energy function converges to a minimum value, the final human three-dimensional model after optimization is output.
[0035] Step S4 is the final execution and problem solving link, which is to use the deformation field to compensate data, and to synchronize the optimization of three-dimensional model and deformation field parameters by minimizing the modulated joint optimization energy function; its effect is that under the continuous guidance of the weight factor, the optimization algorithm finally converges to an optimal solution, and the output three-dimensional model not only shows high reprojection accuracy (spatial consistency) on the key frames identified by high weight, but also has a smooth and reasonable deformation process in time; thereby solving the geometric inconsistency problem of multi-view three-dimensional reconstruction caused by human body micro-motion, and outputting a final model with high precision and high consistency.
[0036] Compared with the prior art, the beneficial effects of the present application are:
[0037] The present application effectively solves the geometric inconsistency problem caused by human body micro-motion by converting the frame quality evaluation data generated in the key frame screening process into quantifiable frame quality weights and embedding them into the joint optimization model, significantly improves the accuracy of three-dimensional reconstruction and the consistency of the model, enhances the adaptability to dynamic human body scenes, and at the same time ensures the spatio-temporal smoothness of the deformation process, finally outputs a high-quality and high-reliability three-dimensional human body model. BRIEF DESCRIPTION OF DRAWINGS
[0038] Figure 1 The figure is a schematic diagram of the overall method steps of the present application;
[0039] Figure 2 The figure is a flowchart of step S2 of the present application;
[0040] Figure 3 The figure is a flowchart of step S3 of the present application. DETAILED DESCRIPTION
[0041] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0042] Next, please refer to Figure 1 One of the purposes of the present embodiment is a human body three-dimensional modeling method based on human body image data scanning, which includes the following steps:
[0043] S1, in the data acquisition step, the human body surface optical information is captured synchronously or near-synchronously from different angles by multiple image acquisition devices arranged around the measured human body; the image acquisition devices adopt uniform lighting to ensure the consistency of illumination, and multiple image sequences are continuously acquired at a fixed frame rate; each acquisition node includes a color imaging unit, and a synchronous trigger control unit ensures the time consistency of image acquisition from different angles; the measured human body is kept in a natural standing posture during the acquisition process, and finally a plurality of human body image sequences from different angles containing human body surface texture and geometric information are obtained, providing complete original data for subsequent processing.
[0044] S2, the process aims to select a group of image frames, i.e. a key frame set, which is least affected by the human body's slight physiological movement and has the highest relative posture consistency between different angles from the plurality of human body image sequences collected asynchronously; the core is to comprehensively evaluate the stability of the human body posture captured by each image frame, which specifically includes:
[0045] Two independent quantitative evaluation strategies are implemented in parallel, one of which is to analyze the pixel flow change pattern in the human body image sequence (the calculation process is realized by calculating the motion vector of the pixel points between adjacent image frames), and the other is to calculate the geometric consistency error between the three-dimensional point cloud data generated by the human body image sequence (the calculation process is realized by calculating the distance residual between the corresponding points of the registered three-dimensional point cloud); the above strategies measure the stability of the image frame from two dimensions of two-dimensional apparent motion and three-dimensional geometric deformation;
[0046] According to the quantitative results of the evaluation, each frame of the human body image sequence is assigned a corresponding frame quality weight; according to the size of the frame quality weight, the candidate image frames with the smallest motion amplitude and the highest geometric consistency are selected from the human body image sequences from different angles; finally, based on the proximity of the time stamps and the coordination of the posture parameters between the candidate frames, a global optimization selection is performed to combine and construct an optimal multi-angle key frame set; the set represents the relatively stable spatial posture of the human body at a certain moment, providing high-quality and time-space synchronous data input for the subsequent three-dimensional reconstruction process; the process of global optimization selection specifically includes:
[0047] A screening condition based on time stamp proximity is established to exclude candidate frames with time stamp differences exceeding a preset threshold from the consideration range, ensuring the time synchronization of the selected key frame set;
[0048] Subsequently, the candidate frame group filtered by time is evaluated for pose parameter consistency, and the consistency of multi-view poses is quantitatively evaluated by calculating the similarity measure between the three-dimensional pose parameters corresponding to each view candidate frame; wherein the three-dimensional pose parameters refer to the three-dimensional spatial position coordinates of the main body joint points extracted by the standard human body skeleton key point detection algorithm (such as shoulders, elbows, wrists, hips, knees, ankles, etc.), and the translation vector calculated from these joint points describing the overall orientation and position of the human body; the similarity measure adopts a quantitative evaluation method based on weighted distance, first calculates the three-dimensional Euclidean distance of the corresponding joint points under different views, then gives different weights according to the preset importance of each joint point, and finally averages all the weighted distances and converts them into a consistency score (the smaller the distance, the higher the consistency score), to quantitatively evaluate the degree of coordination and consistency of multi-view poses.
[0049] Based on the dual indicators of timestamp proximity and pose parameter consistency, a multi-objective optimization function is constructed, and the optimal key frame combination scheme is determined by solving the optimal solution of the function; the final selected multi-view key frame set not only maintains high synchronization in the time dimension, but also presents optimal pose consistency in the spatial dimension, providing high-quality data input for subsequent three-dimensional reconstruction. The design of the multi-objective optimization function considers the two key targets of timestamp proximity (synchronization) and pose parameter consistency (consistency), and by setting the relative importance weight of the two targets (for example, both are set to be equally important), a comprehensive cost function is constructed for optimization solution; wherein the timestamp proximity is quantified by the sum of the squares of the time differences between the candidate frames, and the pose parameter consistency is calculated by the weighted sum of the Euclidean distances of the joint point coordinates; the multi-objective function is constructed as a linear weighted combination of the two metrics, and the weight is set empirically according to the actual data distribution.
[0050] In addition, in the process of analyzing the pixel flow change pattern in the human image sequence, a first frame quality weight is generated based on the amplitude of the pixel flow change pattern, specifically including:
[0051] This process processes the human image sequence arranged in time sequence frame by frame, and obtains the pixel flow change pattern representing the motion state of each point in the picture by calculating the motion vector field of the dense pixel points between adjacent image frames (its calculation process uses an optical flow method to solve the displacement vector of the pixel points between adjacent frames); wherein the calculation process of the motion vector field is realized by an optical flow method, which analyzes the brightness change and space-time gradient of the corresponding pixel points between adjacent image frames to solve and obtain the motion direction and displacement size of each pixel point between consecutive frames, thereby generating a vector field describing the motion state of the entire picture pixels;
[0052] Subsequently, a global statistical calculation is performed on the amplitude of the pixel flow change pattern (the calculation process is to statistically average the magnitudes of all vectors in the motion vector field), so as to obtain a scalar index representing the global motion intensity of the entire frame image; based on the scalar index, the calculation of the first frame quality weight is performed (the calculation process uses an inverse proportional function relationship to map the amplitude to the weight value); the calculation follows the strict inverse proportion principle, that is, the smaller the amplitude of the pixel flow change pattern analyzed and obtained, the weaker the human motion when the frame image is captured, the more stable the posture represented, and therefore the greater the first frame quality weight value calculated and generated; the first frame quality weight is output as the primary quantitative evaluation basis for the reliability of the frame image data, and is used for subsequent key frame screening and optimization modulation. The inverse proportional function relationship specifically adopts the following mathematical form:
[0053] wherein is the input value (motion amplitude or geometric consistency error), is a scaling coefficient greater than 0, used to control the sensitivity of the weight to the input value; the coefficient is pre-set according to the magnitude of the input data, for example, when the motion amplitude range is [0, 10] pixels, k = 1.0 can be set; this function ensures that the smaller the input value , the closer the output weight to 1 (high quality); the larger the input value , the closer the output weight to 0 (low quality).
[0054] In the process of calculating the geometric consistency error between the three-dimensional point cloud data generated by the human image sequence, the second frame quality weight is also calculated based on the geometric consistency error value, specifically including:
[0055] The process first performs a preliminary three-dimensional reconstruction operation on the input human image sequence (the calculation process uses a multi-view stereo vision algorithm to recover three-dimensional point cloud from two-dimensional images), to generate a time-series three-dimensional point cloud data set corresponding to different time points; wherein the preliminary three-dimensional reconstruction operation is realized by a multi-view stereo vision calculation technology, which comprehensively utilizes the two-dimensional image information of multiple different views, calculates and recovers the three-dimensional space coordinates of the human body surface according to the parallax principle and triangulation method, thereby generating three-dimensional point cloud data;
[0056] Then, in the time dimension, the registration operation is performed between the time-series three-dimensional point cloud data and the point cloud data of the adjacent time point (the calculation process uses the iterative closest point algorithm to solve the spatial transformation parameters between the point clouds), and the residual error existing after registration is accurately calculated (the calculation process is to solve the average Euclidean distance between the corresponding points in the two transformed point clouds), which is defined as the geometric consistency error; this error value quantifies the degree of surface deformation of the three-dimensional model caused by human micro-movement; the registration operation uses the iterative closest point algorithm, which continuously iterates and optimizes the rotation and translation parameters between the two point clouds to find their best spatial correspondence, thereby achieving accurate alignment of the point clouds.
[0057] Subsequently, the calculation of the second frame quality weight is performed based on the geometric consistency error value (the calculation process uses an inverse proportional function relationship to map the error value to the weight value); this calculation also follows the strict inverse proportion principle, that is, the smaller the geometric consistency error value calculated, the more consistent the three-dimensional geometry at this moment with the adjacent moment, the smaller the deformation, the higher the data consistency, and therefore the larger the second frame quality weight value calculated; this weight value is output as the core indicator for evaluating data reliability from the three-dimensional geometric dimension. Wherein:
[0058] The inverse proportional function relationship specifically refers to a strict monotonic decreasing mapping rule, the smaller the input motion amplitude or geometric consistency error, the larger the weight value calculated, and the closer to the maximum value 1 (representing the highest quality / reliability); on the contrary, the larger the input value, the smaller the output weight value, and the closer to 0 (representing the lowest quality / reliability); the sensitivity of the mapping (i.e. the degree of change in the weight value caused by the change in the input value) is controlled by an adjustable scaling coefficient, which is pre-set according to the approximate range of the input data (for example, if the motion amplitude is between 0 and 1, the coefficient can be set to 10).
[0059] Referring to Figure 2 , the core goal of step S2 is to select a key frame set representing a stable posture from a multi-view human image sequence, and to generate a quantified frame quality weight to solve the geometric inconsistency problem caused by human micro-movement; this process starts with implementing two independent quantitative evaluation strategies in parallel, specifically including:
[0060] First, the pixel flow change pattern between adjacent frames is analyzed by the optical flow method, the global motion intensity scalar indicator is calculated, and the first frame quality weight is generated based on the inverse proportional mapping (the smaller the motion amplitude, the higher the weight); secondly, the time-series point cloud data is generated by performing preliminary three-dimensional reconstruction, the point cloud registration residual error (geometric consistency error) is calculated using the iterative closest point algorithm, and the second frame quality weight is generated by inverse proportional mapping (the smaller the error, the higher the weight)
[0061] Subsequently, a comprehensive weight is assigned to each frame of image, and a candidate frame with the smallest motion amplitude and the highest geometric consistency is selected; finally, global optimization selection is performed, including timestamp proximity screening (excluding a frame group exceeding a threshold) and posture parameter coordination evaluation (similarity measurement), and a multi-objective optimization function is constructed to determine an optimal key frame set.
[0062] S3, based on the key frame set, an adaptive deformation field for compensating for the deformation of the human body due to micro-movement is calculated, specifically comprising:
[0063] Based on the key frame set obtained by screening, an adaptive deformation field for compensating for the deformation of the human body due to micro-movement is calculated; first, the key frame set is input into an elastic deformation calculation model, which regards the human body as a continuous elastic medium, and describes the micro-movement deformation by establishing the displacement mapping relationship of the surface points of the human body; the calculation process adopts a physical-based deformation model, and a continuous adaptive deformation field is constructed by solving the displacement vector field of the control points; the adaptive deformation field is a three-dimensional vector field, which defines a spatial mapping function from the actual posture at the time of acquisition to the reference posture; the calculation process of the deformation field is performed by iterative optimization, and the parameters of the deformation field are adjusted each time to make the three-dimensional model after deformation better satisfy the multi-view consistency constraint; the finally obtained adaptive deformation field can accurately describe the local deformation of the human body due to micro-movement, and provides a mathematical basis for subsequent deformation compensation. Wherein:
[0064] The control points refer to sparse control grid points in the deformation model, which are initialized as grid points uniformly sampled on the surface of the human body, and the displacement thereof is solved by iterative optimization of the parameters of the deformation field;
[0065] The reference posture is determined by the average of the posture parameters in all key frames, that is, the reference posture is determined by calculating the average of the joint positions of all candidate frames, to ensure that it represents the central tendency of the overall posture;
[0066] The elastic deformation calculation model specifically adopts a linear mixed deformation model based on bones and skin weights; the model regards the surface of the human body as a grid driven by a group of predefined bones, and the final deformation position of each grid vertex is obtained by weighted mixing of its initial position according to a group of preset weights (skin weights) representing the degree of influence of different bones on the vertex, and then applying the rotation and translation transformation of the corresponding bone. The optimization process continuously adjusts the rotation and translation transformation parameters of all bones through iterative algorithms (such as gradient descent) to minimize the error of the three-dimensional model after deformation by these bone transformation on the projection to each view image.
[0067] In addition, step S3 also constructs a joint optimization energy function, which simultaneously contains spatial consistency constraint and temporal smoothness constraint, and the optimization target of the spatial consistency constraint term is modulated by the first frame quality weight and the second frame quality weight, specifically comprising:
[0068] The spatial consistency constraint term ensures the consistency of the three-dimensional model with the two-dimensional image observation values by calculating the multi-view reprojection error; the temporal smoothness constraint term ensures the spatio-temporal continuity of the deformation process by limiting the variation gradient of the adjacent time sequence deformation field; in particular, the optimization target of the spatial consistency constraint term is jointly modulated by the first frame quality weight and the second frame quality weight, and the modulation process first combines the two weights into a comprehensive weight coefficient through weighted fusion, and then the comprehensive weight coefficient is used as a weighted factor in the reprojection error calculation process in the spatial consistency constraint term; this modulation mechanism makes the higher the quality weight of the image frame, the greater the influence on the spatial consistency constraint, thereby giving the high-quality frame data greater weight in the optimization process and improving the accuracy and reliability of the final three-dimensional reconstruction model; wherein:
[0069] The weighted fusion method adopts linear weighted combination, which combines the first frame quality weight (based on pixel flow motion) and the second frame quality weight (based on geometric consistency error) into a single comprehensive weight coefficient according to a preset fusion ratio (for example, initially set as fifty-fifty), and the fusion ratio can be determined by grid search on a typical data set to determine the optimal value. The expression formula of the joint optimization energy function is as follows:
[0070] , wherein represents the total energy function about the three-dimensional point cloud position and the deformation field parameter ; , represents the position set of all three-dimensional points at all time points , which is an optimization variable; , represents the deformation field parameter set of all time points , which is also an optimization variable; represents the feature point index; represents the time frame index; represents the comprehensive weight coefficient of the th point on the th frame, which is generated by the fusion of the first weight and the second weight of the frame (for example , is the fusion coefficient, the initial default value is set to 0.5, and can be fine-tuned according to the characteristics of the data set in actual application); represents a robust kernel function for reducing the influence of outliers; represents the weight coefficient of the temporal smoothness term; represents the L2 norm of the difference value of the deformation field parameters of adjacent time points, which constrains the smooth change of the deformation field with time; represents the camera projection model function, which projects the three-dimensional points to the two-dimensional image plane; denotes an adaptive deformation field function, which is parameterized by a set of deformation field parameters is defined to deform a three-dimensional point from its initial position to a reference pose; denotes camera extrinsic parameters (rotation and translation); denotes two-dimensional pixel coordinates of corresponding feature points observed in images; denotes a spatial consistency constraint term; denotes a temporal smoothness constraint term.
[0071] Referring to Figure 3 , the core of step S3 is to model human micro-motions based on the key frame set and construct a modulatable optimization framework to output high-precision three-dimensional models synchronously; the process starts with inputting the key frame into the elastic deformation calculation model, regarding the human body as a continuous elastic medium, and generating an adaptive deformation field (a three-dimensional vector field, defining a spatial mapping function from the actual pose to the reference pose) by establishing a surface point displacement mapping relationship; then, a joint optimization energy function is constructed, in which the spatial consistency constraint term calculates the multi-view re-projection error, and the first frame quality weight and the second frame quality weight generated by S2 are weighted and fused into a comprehensive weight coefficient for modulation (high-quality frames are given higher weights).
[0072] S4, deforming and compensating the three-dimensional point cloud data by using the adaptive deformation field, synchronously optimizing the three-dimensional geometric model and the adaptive deformation field by minimizing the joint optimization energy function, and outputting the optimized human three-dimensional model, specifically including:
[0073] First, the adaptive deformation field calculated is used to deform and compensate the three-dimensional point cloud data, and the compensation process is performed by applying the spatial mapping function provided by the deformation field to the three-dimensional point cloud data at the acquisition time to transform it to a unified reference pose space;
[0074] Subsequently, the three-dimensional geometric model and the adaptive deformation field are synchronously optimized by minimizing the joint optimization energy function, and an iterative optimization algorithm is used to update the vertex coordinates of the three-dimensional geometric model and the parameter vector of the deformation field simultaneously; in the optimization process, the spatial consistency constraint term ensures that the optimized three-dimensional model is consistent with the observation values of the images of each view, the temporal smoothness constraint term ensures the smooth continuity of the change of adjacent time sequence deformation fields, and the calculation of the re-projection error in the spatial consistency constraint term is jointly modulated by the first frame quality weight and the second frame quality weight, so that the high-quality image frame has a greater weight influence in the optimization process; the iterative optimization algorithm specifically uses the Levenberg-Marquardt algorithm, which is a numerical optimization method widely used to solve nonlinear least squares problems, and combines the advantages of the Gauss-Newton method and the gradient descent method; in each iteration, the vertex coordinate parameters of the three-dimensional geometric model and the parameters (such as the rotation and translation of the skeleton) describing the deformation field are calculated and updated simultaneously, and the core is to intelligently determine the direction and step of parameter update according to the gradient information of the current objective function (joint optimization energy function) and an adaptive damping factor; this mechanism can achieve a good balance between convergence speed and solution stability when dealing with problems such as re-projection error.
[0075] After multiple iterations and optimization, when the joint optimization energy function converges to a minimum value, the final optimized three-dimensional human model is output, which has high geometric accuracy and visual reality, and effectively eliminates the deformation artifacts caused by human micro-movement.
[0076] The above shows and describes the basic principles, main features and advantages of the present application. Those skilled in the art should understand that the present application is not limited by the above examples, and the above examples and descriptions in the specification are only preferred examples of the present application and are not intended to limit the present application. Without departing from the spirit and scope of the present application, various changes and improvements can be made to the present application, and these changes and improvements all fall within the scope of the claimed present application. The scope of protection of the present application is defined by the appended claims and their equivalents.
Claims
1. A method of three-dimensional modeling of a human body based on human body image data scanning, characterized by, The method comprises the following steps: S1, collecting a human body image sequence from multiple perspectives; S2, screening a key frame set representing multiple perspectives of a human body in a relatively stable posture from the human body image sequence, and generating corresponding frame quality weights in the screening process, specifically including: analyzing the pixel flow change pattern in the human body image sequence, and generating a first frame quality weight based on the amplitude of the pixel flow change pattern; calculating the geometric consistency error between the three-dimensional point cloud data generated by the human body image sequence, and generating a second frame quality weight based on the geometric consistency error value; S3, based on the key frame set, calculating an adaptive deformation field for compensating for the deformation of the human body due to micro-movement, and constructing a joint optimization energy function; wherein the joint optimization energy function includes a spatial consistency constraint and a temporal smoothness constraint; wherein the optimization target of the spatial consistency constraint is jointly modulated by the first frame quality weight and the second frame quality weight; S4, using the adaptive deformation field to compensate for the deformation of the three-dimensional point cloud data, synchronously optimizing the three-dimensional geometric model and the adaptive deformation field by minimizing the joint optimization energy function, and outputting the optimized human three-dimensional model; The construction process of the joint optimization energy function specifically includes: The spatial consistency constraint term ensures the consistency of the three-dimensional model and the two-dimensional image observation value by calculating the multi-perspective reprojection error; The temporal smoothness constraint term ensures the spatio-temporal continuity of the deformation process by limiting the change gradient of the adjacent time sequence deformation field; The optimization target of the spatial consistency constraint term is jointly modulated by the first frame quality weight and the second frame quality weight, and the modulation process combines the two weights into a comprehensive weight coefficient through weighted fusion, and the comprehensive weight coefficient is used as a weighted factor in the reprojection error calculation process in the spatial consistency constraint term; The screening process of the key frame set specifically includes: Two independent quantitative evaluation strategies are implemented in parallel, one of which is to analyze the pixel flow change pattern in the human body image sequence, and the other of which is to calculate the geometric consistency error between the three-dimensional point cloud data generated by the human body image sequence; According to the quantitative results generated by the evaluation, a corresponding frame quality weight is assigned to each frame of the human body image sequence; according to the size of the frame quality weight, the candidate image frames with the smallest motion amplitude and the highest geometric consistency are selected from the human body image sequence of each perspective; Based on the proximity of the timestamps and the coordination of the posture parameters between the candidate frames, a global optimization selection is performed to combine and construct an optimal multi-perspective key frame set; The process of global optimization selection specifically includes: Establish a screening condition based on the proximity of the timestamps, and exclude candidate frame groups with a timestamp difference exceeding a preset threshold from the consideration range; Perform posture parameter coordination evaluation on the candidate frame groups screened by time, and quantify the multi-perspective posture consistency by calculating the similarity measure between the three-dimensional posture parameters corresponding to the candidate frames of each perspective; Based on the dual indicators of timestamp proximity and pose parameter consistency, a multi-objective optimization function is constructed, and the optimal key frame combination scheme is determined by solving the optimal solution of the function as the optimal multi-view key frame set.
2. The method of claim 1, wherein the method further comprises: The generation process of the first frame quality weight specifically includes: The human body image sequence arranged in time sequence is processed frame by frame, the motion vector field of the dense pixel points between adjacent image frames is calculated, and the pixel flow change mode representing the motion state of each point in the picture is obtained; The amplitude of the pixel flow change mode is globally statistically calculated to obtain a scalar indicator representing the global motion intensity of the whole image frame; Based on the scalar indicator, the amplitude is mapped to a weight value by using an inverse proportional function relationship, and the first frame quality weight is calculated and generated.
3. The method of claim 2, wherein the method further comprises: The calculation process of the motion vector field solves the displacement vector of the pixel points between adjacent frames by using an optical flow method.
4. The method of claim 1, wherein the method further comprises: The generation process of the second frame quality weight specifically includes: A preliminary three-dimensional reconstruction operation is performed on the input human body image sequence to generate a time sequence three-dimensional point cloud data set corresponding to different time points; In the time dimension, the registration operation is performed on each time sequence three-dimensional point cloud data and the point cloud data of the adjacent time point, and the residual error existing after registration is accurately calculated, which is defined as a geometric consistency error; Based on the geometric consistency error value, the error value is mapped to a weight value by using an inverse proportional function relationship, and the second frame quality weight is calculated and generated.
5. The method of claim 4, wherein the method further comprises: The preliminary three-dimensional reconstruction operation recovers the three-dimensional point cloud from the two-dimensional image by using a multi-view stereo vision algorithm; and the registration operation solves the space transformation parameters between the point clouds by using an iterative closest point algorithm.
6. The method of claim 1, wherein the method further comprises: The generation process of the adaptive deformation field specifically includes: The key frame set is input into an elastic deformation calculation model, the model regards the human body as a continuous elastic medium, and describes the micro-motion deformation by establishing the displacement mapping relationship of the human body surface points; An adaptive deformation field is constructed by solving the displacement vector field of the control points based on the physical deformation model; the adaptive deformation field is a three-dimensional vector field, and defines a space mapping function from the actual pose at the collection time to the reference pose; The calculation process of the deformation field is performed by iteration optimization, and the parameters of the deformation field are adjusted at each iteration to make the deformed three-dimensional model meet the multi-view consistency constraint.
7. The method of claim 1, wherein the method further comprises: determining a plurality of body image data; and determining a plurality of body image data. The output process of the human body three-dimensional model specifically includes: The adaptive deformation field is calculated, the three-dimensional point cloud data is deformed and compensated by using the adaptive deformation field, the compensation process applies the space mapping function provided by the deformation field to the three-dimensional point cloud data at the collection time, and the three-dimensional point cloud data is transformed to a unified reference pose space; The three-dimensional geometric model and the adaptive deformation field are simultaneously optimized by minimizing the joint optimization energy function, and the optimization process simultaneously updates the vertex coordinates of the three-dimensional geometric model and the parameter vector of the deformation field by using an iterative optimization algorithm; When the joint optimization energy function converges to a minimum value, the final optimized human body three-dimensional model is output.
Citation Information
Patent Citations
Three-dimensional digital modeling method based on ball screen video stream
CN108830925A
Dense motion tracking mechanism
US20200234455A1