A foot parameter measurement extraction method and system based on a parameterized model
By using multi-view visual acquisition and multi-head attention deep learning technology, a parametric model of adolescent feet is generated, which solves the problems of poor model adaptability, low measurement accuracy, complex operation and high cost in the existing technology, and realizes accurate measurement and efficient assessment of adolescent feet.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-01-05
- Publication Date
- 2026-04-07
AI Technical Summary
Existing 3D foot modeling technology cannot accurately depict the growth pattern of adolescent feet, nor can it effectively capture the local geometric distortion features of unhealthy foot types. Furthermore, existing measurement methods are complex to operate, costly, and have low accuracy, making it difficult to meet the batch measurement needs of adolescents.
Employing multi-view visual acquisition, mesh registration optimization, and multi-head attention deep learning techniques, depth and color images are acquired through an industrial camera to generate multi-view initial point clouds and convert them into triangular mesh models. Combined with principal component analysis and graph structure features, a parametric model is constructed, and foot anthropometric indicators are extracted through a multi-head attention neural network.
It achieves accurate characterization of adolescent foot shape, improves the accuracy of capturing geometric distortions of unhealthy foot types, reduces measurement errors, improves batch measurement efficiency, reduces equipment costs, and is suitable for batch screening scenarios in schools and communities.
Smart Images

Figure CN121457334B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of three-dimensional parametric modeling and precise measurement technology of the foot, and in particular to a method and system for measuring and extracting foot parameters based on a parametric model. Background Technology
[0002] Currently, the mainstream 3D foot modeling technologies in the industry are mainly divided into two categories, but both have significant shortcomings in adapting to the foot characteristics of adolescents and unhealthy foot types:
[0003] 1. Parametric modeling technology based on principal component analysis (PCA): This technology relies on the dimensionality reduction capability of PCA to represent the positional changes of various points on the foot surface with a few key dimensions, thereby achieving a parametric description of foot shape. It is one of the core technical paths for existing foot modeling. Its advantage lies in its ability to quickly associate foot shape features through low-dimensional parameters, making it suitable for batch data processing scenarios. However, its core shortcomings are: existing parametric models are mostly built based on foot data of people of all ages, without specifically optimizing for the developmental characteristics of adolescent foot bones and soft tissues (such as the dynamic changes in arch height with age, rapid growth rate of foot length, and significant differences in foot fat distribution compared to adults). This makes it impossible to accurately depict the growth pattern of adolescent feet, resulting in insufficient accuracy in the model's representation of adolescent foot types. The principal components extracted by PCA mostly focus on the common shape changes of healthy foot types, making it difficult to cover the local geometric distortion features of unhealthy foot types such as flat feet (collapsed arches) and high arches (excessively convex arches), thus failing to provide effective support for subsequent measurement and health assessment of unhealthy foot types.
[0004] 2. Implicit Field-Based Encoder-Decoder Network Modeling Technology: This technology constructs an encoder-decoder neural network targeting foot features, mapping foot parameters to implicit fields to generate foot models containing fine mesh structures, texture information, and multiple poses. Its advantage lies in generating highly detailed 3D foot models, adapting to the modeling needs of diverse foot poses; however, it has significant limitations: the implicit field output is a continuous foot geometric surface, requiring additional complex post-processing steps such as point cloud sampling and local extremum detection to extract anthropometric indicators such as foot length and arch height. This prevents direct modeling and measurement, resulting in low measurement efficiency, and the post-processing process is prone to introducing errors. Furthermore, model training requires collecting massive amounts of foot data of different foot types and poses, and the data annotation process is complex. For special groups such as adolescents, the difficulty and cost of data collection are even higher, making large-scale application difficult.
[0005] Furthermore, current methods for obtaining foot anthropometry indicators mainly fall into three categories, but each has shortcomings in balancing efficiency, cost, and accuracy, and is particularly unsuitable for the needs of batch measurement and precise assessment of adolescents:
[0006] 1. Measurement technology based on 3D scanning equipment: This technology acquires 3D point clouds of the foot using specialized equipment such as laser scanning and structured light scanning, and then processes and extracts measurement indicators using dedicated software. It is currently the main method for high-precision measurement. Its advantages lie in the high accuracy of point cloud data and the ability to obtain comprehensive geometric information of the foot; however, it has significant drawbacks: the unit price of industrial-grade 3D scanning equipment usually exceeds 100,000 yuan, and professional personnel are required to debug the equipment and control the measurement environment, making it difficult to promote on a large scale in schools, communities, and other scenarios for adolescent foot health screening; a single scan and data processing usually takes more than 5 minutes, which cannot meet the batch measurement needs of adolescents; at the same time, environmental interference can easily lead to incomplete point clouds, thus affecting measurement accuracy.
[0007] 2. Manual Measurement Techniques; This technique involves manually measuring foot length, width, and arch height using tools such as measuring tapes and protractors, and is the primary measurement method for traditional custom footwear. Its advantages lie in its low cost and flexibility; however, it has fatal flaws: the measurement results rely heavily on the operator's experience, resulting in a large error rate, and it cannot meet the stringent accuracy requirements of adolescent foot health assessments; it can only obtain two-dimensional indicators such as linearity and angles, failing to capture the three-dimensional geometric features of the foot, making it difficult to support the three-dimensional design of personalized footwear.
[0008] 3. Deep learning measurement techniques based on traditional neural networks; To address the efficiency and cost issues of the aforementioned techniques, measurement schemes based on traditional neural networks (such as CNN and ResNet) have emerged in recent years. The core of these schemes is to directly regress measurement indicators from foot images. However, they suffer from two major drawbacks: First, local features of adolescent feet (such as arches and toe joints) are prone to slight deformation due to posture and stress. Traditional neural networks lack the ability to capture local geometric relationships, resulting in low prediction accuracy for indicators such as arch height and foot width that depend on local features. Second, the network directly establishes a mapping relationship between image pixels and measurement values without incorporating prior information from the three-dimensional geometric model of the foot. This leads to a disconnect between the output measurement values and the actual geometric features of the foot, failing to meet the needs of health assessment.
[0009] In summary, existing technologies for precise measurement and parametric modeling of adolescent feet have the following problems: existing parametric models are not optimized for adolescent foot development characteristics and unhealthy foot types, and cannot accurately represent the foot type patterns of the target group; high-precision 3D scanning technology is complex and costly to operate, while convenient manual and traditional deep learning measurement technologies are not accurate enough; existing modeling technologies require additional post-processing to extract measurement indicators, which cannot achieve a closed loop of modeling and measurement, resulting in low efficiency and easy introduction of errors; traditional deep learning technologies lack foot geometric priors, cannot effectively capture the local deformation characteristics of adolescent feet, and affect the reliability of measurement. Summary of the Invention
[0010] The purpose of this invention is to provide a method and system for measuring and extracting foot parameters based on a parametric model, so as to solve the above-mentioned technical problems in the prior art.
[0011] According to a first aspect of the present invention, a method for measuring and extracting foot parameters based on a parametric model is provided.
[0012] The foot parameter measurement and extraction method based on a parametric model includes:
[0013] An industrial camera is set up according to a preset layout, and the camera's internal and external parameters in a unified coordinate system are obtained. Based on the camera's internal and external parameters, the subject's foot depth image and foot color image are acquired synchronously.
[0014] Based on camera intrinsic and extrinsic parameters and foot depth images, a multi-view initial point cloud is generated, and a complete foot point cloud is generated based on the multi-view initial point cloud; the multi-view initial point cloud is then converted into a foot triangular mesh model.
[0015] Based on the foot triangular mesh model, standard mesh templates are selected, and the foot triangular mesh model is aligned with the standard mesh templates at the vertex level to obtain a standardized foot mesh. Based on the complete foot point cloud, anthropometry indicators are calculated, and a linear relationship between the vertex coordinates of the standardized foot mesh and the anthropometry indicators is constructed to generate a measurement mapping matrix.
[0016] The vertex displacement field of the standardized foot mesh relative to the standard mesh template is calculated, the shape principal component matrix is extracted by principal component analysis, a parametric model of the foot is established by combining the standard mesh template, and graph structure features are constructed based on the point cloud topology of the foot parametric model.
[0017] A multi-head attention mechanism neural network with fused graph structure features is constructed. The multi-head attention mechanism neural network is trained with a foot color image as input and a foot triangular mesh model as labeled samples. After the multi-head attention mechanism neural network outputs the target foot mesh coordinates, the corresponding vertex coordinates are indexed by the measurement mapping matrix to extract foot anthropometric indicators.
[0018] The industrial cameras consist of four units, arranged in a rectangular layout with uniform height from the ground.
[0019] The process of obtaining camera intrinsic and extrinsic parameters in a unified coordinate system includes: placing a calibration plate at the center of the camera array, using the Zhang Zhengyou calibration method to jointly calibrate the intrinsic and extrinsic parameters of multiple cameras, and obtaining the camera intrinsic and extrinsic parameters in a unified coordinate system.
[0020] The process involves generating multi-view initial point clouds based on camera intrinsic and extrinsic parameters and foot depth images, and then generating complete foot point clouds based on these initial point clouds. The transformation of the multi-view initial point clouds into a foot triangular mesh model includes: calculating the 3D coordinates of corresponding pixel points using a stereo matching algorithm based on the obtained camera intrinsic and extrinsic parameters and foot depth images to generate multi-view initial point clouds; converting the multi-view initial point clouds to the same coordinate system using the spatial relationship of the camera intrinsic and extrinsic parameters; fusing them into a complete foot point cloud after deduplication and hole filling; and then using a Poisson surface reconstruction algorithm on the complete foot point cloud to extract isosurfaces through normal vector field construction and Poisson equation solving, transforming it into a closed foot triangular mesh model.
[0021] The process of aligning the foot triangular mesh model with the standard mesh template at the vertex level to obtain a standardized foot mesh includes: using a registration algorithm and a deformation algorithm to align a batch of the foot triangular mesh models with the standard mesh template at the vertex level to obtain a standardized foot mesh; wherein the registration algorithm is the nearest iteration point algorithm; and the deformation algorithm is an area-preserving deformation algorithm.
[0022] The calculation of anthropometry indicators based on complete foot point clouds includes: calculating foot anthropometry indicators based on complete foot point clouds using a local extremum detection algorithm; wherein, the local extremum detection algorithm calculates foot length by indexing the maximum and minimum differences of the coordinates of the complete foot point cloud in the foot length direction and calculates arch height by detecting the height of the apex of the arch region, thereby obtaining foot anthropometry indicators; the foot anthropometry indicators include at least: foot length and arch height.
[0023] The calculation of the vertex displacement field of the standardized foot mesh relative to the standard mesh template, and the extraction of the shape principal component matrix by principal component analysis, includes: calculating the vertex displacement field of the standardized foot mesh relative to the standard mesh template, flattening the vertex displacement field into a high-dimensional vector and combining it into a displacement matrix, and extracting the shape principal component matrix after centering, calculating the covariance matrix and filtering the eigenvalues of the displacement matrix.
[0024] Wherein, the measurement mapping matrix is a complete foot point cloud sequence index matrix corresponding to foot anthropometric indicators; the graph structure feature is a tree diagram constructed based on the point cloud topology relationship of the foot parametric model.
[0025] The subjects were adolescents, and the foot parametric model was constructed based on standardized foot mesh data of several adolescents to characterize the growth characteristics of adolescent feet and the shape changes of unhealthy foot types.
[0026] According to a second aspect of the present invention, a foot parameter measurement and extraction system based on a parametric model is provided.
[0027] The foot parameter measurement and extraction system based on a parametric model includes:
[0028] The data acquisition module is used to set up an industrial camera according to a preset layout and acquire the camera's internal and external parameters in a unified coordinate system. Based on the camera's internal and external parameters, the module synchronously acquires the subject's foot depth image and foot color image.
[0029] The mesh generation module is used to generate multi-view initial point clouds based on camera intrinsic and extrinsic parameters and foot depth images, and to generate complete foot point clouds based on multi-view initial point clouds; and to convert multi-view initial point clouds into foot triangular mesh models.
[0030] The mesh processing module is used to select standard mesh templates based on the foot triangular mesh model, align the foot triangular mesh model with the standard mesh templates at the vertex level to obtain a standardized foot mesh; and calculate anthropometric indicators based on the complete foot point cloud, construct a linear relationship between the vertex coordinates of the standardized foot mesh and the anthropometric indicators, and generate a measurement mapping matrix.
[0031] The parameterization module is used to calculate the vertex displacement field of the standardized foot mesh relative to the standard mesh template, extract the shape principal component matrix through principal component analysis, establish a foot parameterization model in combination with the standard mesh template, and construct graph structure features based on the point cloud topology of the foot parameterization model.
[0032] The network extraction module is used to construct a multi-head attention mechanism neural network that integrates graph structure features. The multi-head attention mechanism neural network is trained with a foot color image as input and a foot triangular mesh model as labeled samples. After the multi-head attention mechanism neural network outputs the target foot mesh coordinates, the corresponding vertex coordinates are indexed by the measurement mapping matrix to extract foot anthropometric indicators.
[0033] The industrial cameras consist of four units, arranged in a rectangular layout with uniform height from the ground.
[0034] When acquiring camera intrinsic and extrinsic parameters in a unified coordinate system, the data acquisition module places a calibration plate at the center of the camera array and uses the Zhang Zhengyou calibration method to jointly calibrate the intrinsic and extrinsic parameters of multiple cameras to obtain the camera intrinsic and extrinsic parameters in a unified coordinate system.
[0035] The mesh generation module generates multi-view initial point clouds based on camera intrinsic and extrinsic parameters and foot depth images, and then generates complete foot point clouds based on these initial point clouds. When converting the multi-view initial point clouds into a foot triangular mesh model, the module calculates the three-dimensional coordinates of corresponding points of pixels using a stereo matching algorithm based on the obtained camera intrinsic and extrinsic parameters and foot depth images to generate multi-view initial point clouds. The spatial relationship of the camera intrinsic and extrinsic parameters is used to convert the multi-view initial point clouds to the same coordinate system. After deduplication and hole filling, the points are fused into a complete foot point cloud. Then, the Poisson surface reconstruction algorithm is applied to the complete foot point cloud. The isosurface is extracted by constructing a normal vector field and solving the Poisson equation, and the point cloud is converted into a closed foot triangular mesh model.
[0036] Specifically, when the mesh processing module aligns the foot triangular mesh model with the standard mesh template at the vertex level to obtain a standardized foot mesh, it employs a registration algorithm and a deformation algorithm to align a batch of foot triangular mesh models with the standard mesh template at the vertex level to obtain a standardized foot mesh. The registration algorithm is the nearest iteration point algorithm, and the deformation algorithm is an area-preserving deformation algorithm.
[0037] Specifically, when the mesh processing module calculates anthropometric indicators based on a complete foot point cloud, it calculates the foot anthropometric indicators using a local extremum detection algorithm. The local extremum detection algorithm calculates foot length by indexing the maximum / minimum difference of the complete foot point cloud coordinates along the foot length direction and calculates arch height by detecting the height of the arch region apex, thus obtaining the foot anthropometric indicators. These foot anthropometric indicators include at least: foot length and arch height.
[0038] Specifically, when the parameterization module calculates the vertex displacement field of the standardized foot mesh relative to the standard mesh template and extracts the shape principal component matrix through principal component analysis, it calculates the vertex displacement field of the standardized foot mesh relative to the standard mesh template, flattens the vertex displacement field into a high-dimensional vector and combines it into a displacement matrix, and extracts the shape principal component matrix after centering, calculating the covariance matrix and filtering the eigenvalues of the displacement matrix.
[0039] Wherein, the measurement mapping matrix is a complete foot point cloud sequence index matrix corresponding to foot anthropometric indicators; the graph structure feature is a tree diagram constructed based on the point cloud topology relationship of the foot parametric model.
[0040] The subjects were adolescents, and the foot parametric model was constructed based on standardized foot mesh data of several adolescents to characterize the growth characteristics of adolescent feet and the shape changes of unhealthy foot types.
[0041] The technical solution provided by this invention may include the following beneficial effects:
[0042] This invention addresses the core need for parametric modeling and accurate measurement of adolescent feet. By combining multi-view visual acquisition, grid registration optimization, principal component analysis parameterization, and multi-head attention deep learning technology, it solves the problems of poor model adaptability, low measurement accuracy, complex operation, and high cost of existing technologies.
[0043] This invention trains a parametric model using multi-view point cloud data of adolescent feet, and extracts principal component vectors specific to adolescent foot types using principal component analysis. This reduces the model's shape representation error of normal adolescent foot types and improves the accuracy of capturing geometric distortions of flat feet and high arches. It is superior to models for all age groups and can effectively support adolescent foot development monitoring and health assessment.
[0044] This invention achieves vertex-level alignment between batch foot meshes and standard templates through the nearest iteration point algorithm and the area-preserving deformation algorithm, ensuring a unified mesh topology and avoiding measurement deviations caused by structural inconsistencies. The multi-head attention network can accurately capture minute deformations of local foot features, reducing measurement errors of indicators such as arch height and foot width that depend on local features. Its accuracy is superior to traditional neural networks and manual measurements, meeting the precision requirements of accurate health assessment and personalized footwear customization.
[0045] This invention only requires the subject to stand in the center of the camera array, and simultaneously acquire color images of the feet. After inference through a multi-head attention network, indicators such as foot length and arch height can be directly output without additional post-processing, which greatly improves the screening efficiency when measuring in batches. It can be adapted to batch screening scenarios for adolescents in schools, communities and other places.
[0046] This invention requires only 4 ordinary industrial cameras and terminals, significantly reducing the equipment investment threshold; at the same time, the camera layout does not require complex environmental modifications and can be deployed in ordinary scenarios such as classrooms and shoe stores, solving the problem of high environmental requirements and difficulty in promotion of existing technologies.
[0047] This invention can generate 3D meshes for different adolescent foot types by adjusting the principal component weights, seamlessly adapting to downstream scenarios such as animation production and personalized shoe last design. Compared with existing single-function modeling / measurement technologies, the reusability of this model is greatly improved, expanding the application boundaries of the technology.
[0048] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit the invention. Attached Figure Description
[0049] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with the invention and, together with the description, serve to explain the principles of the invention.
[0050] Figure 1This is a flowchart illustrating a method for measuring and extracting foot parameters based on a parametric model, according to an exemplary embodiment.
[0051] Figure 2 This is a structural block diagram of a foot parameter measurement and extraction system based on a parametric model, according to an exemplary embodiment.
[0052] Figure 3 This is a schematic diagram of the structure of a computer device according to an exemplary embodiment. Detailed Implementation
[0053] The following description and accompanying drawings fully illustrate specific embodiments described herein to enable those skilled in the art to practice them. Some embodiments may include or substitute parts and features of other embodiments. The scope of the embodiments herein encompasses the entire scope of the claims and all available equivalents thereof. Throughout this document, the terms “first,” “second,” etc., are used only to distinguish one element from another without requiring or implying any actual relationship or order between the elements. Indeed, a first element can also be referred to as a second element, and vice versa. Furthermore, the terms “comprising,” “including,” or any other variations thereof are intended to cover non-exclusive inclusion, such that a structure, apparatus, or device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a structure, apparatus, or device. Without further limitation, an element defined by the phrase “comprising one…” does not exclude the presence of other identical elements in the structure, apparatus, or device that includes said element. The various embodiments described herein are presented in a progressive manner, with each embodiment focusing on its differences from other embodiments; similar or identical parts between embodiments can be referred to interchangeably.
[0054] The terms "longitudinal," "lateral," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer" used in this document to indicate orientations or positional relationships are based on the orientations or positional relationships shown in the accompanying drawings. They are used solely for the convenience of describing the document and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In the description herein, unless otherwise specified and limited, the terms "installed," "connected," and "linked" should be interpreted broadly. For example, they can refer to mechanical or electrical connections, or internal connections between two elements; they can be direct connections or indirect connections through an intermediate medium. Those skilled in the art can understand the specific meaning of the above terms according to the specific circumstances.
[0055] In this document, unless otherwise stated, the term "multiple" means two or more.
[0056] In this article, the character " / " indicates that the objects before and after it are in an "or" relationship. For example, A / B means: A or B.
[0057] In this article, the term "and / or" describes an association between objects, indicating that three relationships can exist. For example, A and / or B means: A or B, or A and B.
[0058] It should be understood that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order constraint on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the diagram may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0059] The modules in the apparatus or system of this application can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0060] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0061] Figure 1 An embodiment of a foot parameter measurement and extraction method based on a parametric model according to the present invention is shown.
[0062] In this optional embodiment, the foot parameter measurement and extraction method based on a parametric model includes:
[0063] Step S101: Set up an industrial camera according to a preset layout and obtain the camera's internal and external parameters in a unified coordinate system. Based on the camera's internal and external parameters, simultaneously acquire the subject's foot depth image and foot color image.
[0064] Step S102: Based on the camera intrinsic and extrinsic parameters and foot depth image, generate multi-view initial point cloud, and based on the multi-view initial point cloud, generate complete foot point cloud; convert the multi-view initial point cloud into a foot triangular mesh model.
[0065] Step S103: Based on the foot triangular mesh model, a standard mesh template is selected, and the foot triangular mesh model is aligned with the standard mesh template at the vertex level to obtain a standardized foot mesh; and based on the complete foot point cloud, anthropometry indicators are calculated, and a linear relationship between the vertex coordinates of the standardized foot mesh and the anthropometry indicators is constructed to generate a measurement mapping matrix.
[0066] Step S104: Calculate the vertex displacement field of the standardized foot mesh relative to the standard mesh template, extract the shape principal component matrix through principal component analysis, establish a foot parametric model in combination with the standard mesh template, and construct graph structure features based on the point cloud topology of the foot parametric model.
[0067] Step S105: Construct a multi-head attention mechanism neural network that integrates graph structure features. Train the multi-head attention mechanism neural network with a foot color image as input and a foot triangular mesh model as labeled samples. After the multi-head attention mechanism neural network outputs the target foot mesh coordinates, extract foot anthropometric indicators by indexing the corresponding vertex coordinates through a measurement mapping matrix.
[0068] Figure 2 An embodiment of a foot parameter measurement and extraction system based on a parametric model according to the present invention is shown.
[0069] In this optional embodiment, the foot parameter measurement and extraction system based on a parametric model includes:
[0070] The data acquisition module 201 is used to set up an industrial camera according to a preset layout and acquire the camera's internal and external parameters in a unified coordinate system, and simultaneously acquire the subject's foot depth image and foot color image based on the camera's internal and external parameters.
[0071] Mesh generation module 202 is used to generate multi-view initial point clouds based on camera intrinsic and extrinsic parameters and foot depth images, and to generate complete foot point clouds based on multi-view initial point clouds; and to convert multi-view initial point clouds into foot triangular mesh models.
[0072] The mesh processing module 203 is used to select standard mesh templates based on the foot triangular mesh model, align the foot triangular mesh model with the standard mesh templates at the vertex level to obtain a standardized foot mesh; and calculate anthropometry indices based on the complete foot point cloud, construct a linear relationship between the vertex coordinates of the standardized foot mesh and the anthropometry indices, and generate a measurement mapping matrix.
[0073] The parameterization module 204 is used to calculate the vertex displacement field of the standardized foot mesh relative to the standard mesh template, extract the shape principal component matrix through principal component analysis, establish a foot parameterization model in combination with the standard mesh template, and construct graph structure features based on the point cloud topology of the foot parameterization model.
[0074] The network extraction module 205 is used to construct a multi-head attention mechanism neural network that integrates graph structure features. The multi-head attention mechanism neural network is trained with a foot color image as input and a foot triangular mesh model as labeled samples. After the multi-head attention mechanism neural network outputs the target foot mesh coordinates, the corresponding vertex coordinates are indexed by the measurement mapping matrix to extract foot anthropometric indicators.
[0075] In practical applications, four industrial cameras were set up in a rectangular layout at the four corners of the test site, with each camera approximately 40cm above the ground. The subject, wearing thin socks, stood at the center of the camera array. Simultaneous acquisition of depth and color images of the feet was triggered via a PC. To obtain the camera's intrinsic and extrinsic parameters in a unified coordinate system, a calibration plate was placed at the center of the camera array, and the Zhang Zhengyou calibration method was used to jointly calibrate the intrinsic and extrinsic parameters of the multiple cameras, thus obtaining the unified coordinate system's intrinsic and extrinsic parameters. Specifically, four cameras simultaneously acquire calibration board images in multiple poses (including multiple poses that can be observed by two or more cameras). The intrinsic parameters and initial extrinsic parameters relative to each calibration board are calculated separately for each camera. Based on this relative relationship, the pose relationship between the cameras is initially estimated. Then, the initial estimated parameters of all cameras, including their intrinsic parameters, distortion coefficients, and extrinsic parameters relative to a unified world coordinate system, are used to establish a global optimization model and simultaneously perform fine-grained iterations to minimize the total reprojection error between the corner projection positions and the actual observation positions in all images from all cameras. This results in a set of consistent and reliable multi-camera intrinsic and extrinsic parameters with high precision.
[0076] Based on camera intrinsic and extrinsic parameters and foot depth images, a multi-view initial point cloud is generated, and a complete foot point cloud is generated based on the multi-view initial point cloud. When converting the multi-view initial point cloud into a foot triangular mesh model, the three-dimensional coordinates of corresponding points of pixels can be calculated using a stereo matching algorithm based on the obtained camera intrinsic and extrinsic parameters and foot depth images to generate the multi-view initial point cloud. The spatial relationship of the camera intrinsic and extrinsic parameters is used to convert the multi-view initial point cloud to the same coordinate system. After deduplication and hole filling processing, it is fused into a complete foot point cloud. Then, the Poisson surface reconstruction algorithm is used on the complete foot point cloud. The isosurface is extracted by constructing a normal vector field and solving the Poisson equation, and it is converted into a closed foot triangular mesh model.
[0077] The stereo matching algorithm involves: based on the high-precision camera intrinsic and extrinsic parameters obtained through calibration, pixel-level corresponding point searches are performed between paired camera views. The coordinates in 3D space are calculated based on the principle of triangulation, thus generating initial foot point clouds from multiple perspectives. The multi-view point cloud fusion process involves: using the extrinsic parameters of all cameras relative to a unified world coordinate system obtained through calibration, all point cloud data from different perspectives, including overlapping and incomplete data, are transformed to the same global coordinate system for registration and fusion. By removing overlapping redundancy and filling holes caused by occlusion, a complete 3D foot point cloud model is finally generated. The Poisson surface reconstruction algorithm transforms the discrete point cloud and its normal vector field into a continuous function and extracts the isosurface of this function to generate a closed triangular mesh model. First, the normal vector of each point in the point cloud is estimated, and these normal vectors are constructed into a vector field. By solving the Poisson equation, an optimal implicit surface that best fits this vector field is reconstructed. Finally, by extracting the isosurface of this implicit surface, a watertight, smooth triangular mesh that can effectively fill small holes in the point cloud can be obtained.
[0078] When aligning the foot triangular mesh model with the standard mesh template at the vertex level to obtain a standardized foot mesh, a registration algorithm and a deformation algorithm are used to align the batch of foot triangular mesh models with the standard mesh template at the vertex level to obtain a standardized foot mesh; wherein, the registration algorithm is the nearest iteration point algorithm (NICP); and the deformation algorithm is the area-preserving deformation algorithm (ARAP).
[0079] Specifically, the NICP algorithm is used to perform vertex-level alignment between batch foot meshes and standard meshes, unifying the vertex topology. The NICP algorithm iteratively finds the nearest point on the target standard mesh for each source mesh vertex, then optimizes the vertex position using the ARAP algorithm, ultimately achieving high-precision geometric alignment between all source foot meshes and the standard mesh, and forcibly unifying their vertex topology. The ARAP algorithm achieves rigid deformation preservation from the source point cloud to the target point cloud through an iterative optimization process. Its process is as follows: first, the rotation matrix of each vertex is initialized; then, in each iteration, intermediate variables are generated through linear transformation and robust weights are calculated; next, the right-hand side of the system is constructed and the rotation matrix is updated, while simultaneously transforming the normal vector direction; subsequently, a global linear system is constructed and matrix decomposition is used to efficiently solve for vertex position updates; finally, after multiple iterations, optimized vertex coordinates that maintain local rigidity and satisfy normal constraints are output.
[0080] When calculating anthropometric indicators based on a complete foot point cloud, the local extremum detection algorithm is used to calculate the foot anthropometric indicators based on the complete foot point cloud. The local extremum detection algorithm calculates the foot length by indexing the maximum and minimum values of the coordinates of the complete foot point cloud in the foot length direction and calculates the arch height by detecting the height of the apex of the arch region. The foot anthropometric indicators include at least: foot length and arch height.
[0081] When calculating the vertex displacement field of the standardized foot mesh relative to the standard mesh template and extracting the shape principal component matrix through principal component analysis, the vertex displacement field of the standardized foot mesh relative to the standard mesh template is calculated, flattened into high-dimensional vectors, and combined into a displacement matrix. The displacement matrix is then centered, its covariance matrix is calculated, and eigenvalues are selected to extract the shape principal component matrix. The processing flow for the displacement matrix is as follows: first, the displacement matrix is centered to eliminate the influence of the overall offset; then, the covariance matrix of the centered displacement matrix is calculated; finally, the eigenvalues and eigenvectors of the covariance matrix are solved, and the top k eigenvectors are selected in descending order of eigenvalues to form the shape principal component matrix.
[0082] In summary, the Principal Component Analysis (PCA) process is as follows: The vertex displacement field of each matched grid relative to the standard grid is calculated, where the displacement vector of each sample is flattened into a high-dimensional vector and combined into a displacement matrix. Next, the displacement matrix is centered (mean subtraction), then its covariance matrix is calculated, and the eigenvalues and eigenvectors of the covariance matrix are solved. The top k principal eigenvectors, sorted in descending order of eigenvalues, are selected as shape principal components. These principal components capture the key variance patterns of foot grid deformation. Finally, these principal components are combined into a shape principal component matrix. The local extremum detection method, on the other hand, searches for corresponding points on the original foot point cloud for calculations based on different foot measurement points. Taking foot length as an example, the difference between the maximum and minimum coordinates of the foot point cloud along the foot length direction is calculated to obtain foot length information. The mapping matrix is a matrix composed of the index values of the foot measurement points in the original foot point cloud. The graph structure feature is a tree diagram constructed based on the topological relationships of the point cloud from the foot's parametric model.
[0083] In practical applications, the subjects are selected as adolescents, and the foot parametric model is constructed based on standardized foot mesh data from a number of adolescents (e.g., 10,000) to characterize the growth characteristics of adolescent feet and the shape changes of unhealthy foot types. The multi-head attention mechanism neural network is designed as a foot parametric mesh regression network for single-view images. This network takes a single-view image as input and, combined with an established measurement mapping matrix, can accurately acquire foot measurement indicators. Through training with a large amount of sample data, the network parameters are optimized to improve the model's generalization ability and measurement accuracy. The multi-head attention mechanism neural network enhances the ability to model foot surface deformation by projecting foot color images and graph structure features in parallel to multiple subspaces and fusing multi-dimensional attention information. The training method of the multi-head attention mechanism neural network is supervised learning, and the network parameters are optimized to minimize the coordinate error between the output target foot mesh coordinates and the foot triangular mesh model of the labeled sample.
[0084] Figure 3 An embodiment of a computer device according to the present invention is shown. The computer device may be a server, and includes a processor, memory, and a network interface connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database of the computer device stores static and dynamic information data. The network interface of the computer device is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the steps in the above-described method embodiment.
[0085] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present invention and does not constitute a limitation on the computer device to which the present invention is applied. A specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0086] In addition, the present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0087] In addition, the present invention also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the above method embodiments.
[0088] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the methods described above. Any references to memory, storage, databases, or other media used in the embodiments provided by this invention can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, or optical storage, etc. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0089] This invention is not limited to the structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this invention is limited only by the appended claims.
Claims
1. A method for measuring and extracting foot parameters based on a parametric model, characterized in that, include: An industrial camera is set up according to a preset layout, and the camera's internal and external parameters in a unified coordinate system are obtained. Based on the camera's internal and external parameters, the subject's foot depth image and foot color image are acquired synchronously. Based on camera intrinsic and extrinsic parameters and foot depth images, a multi-view initial point cloud is generated, and based on the multi-view initial point cloud, a complete foot point cloud is generated. Convert the initial point cloud from multiple perspectives into a foot triangular mesh model; Based on the foot triangular mesh model, standard mesh templates are selected, and the foot triangular mesh model is aligned with the standard mesh templates at the vertex level to obtain a standardized foot mesh. Based on the complete foot point cloud, anthropometry indicators are calculated, a linear relationship between the vertex coordinates of the standardized foot mesh and the anthropometry indicators is constructed, and a measurement mapping matrix is generated. The vertex displacement field of the standardized foot mesh relative to the standard mesh template is calculated, the shape principal component matrix is extracted by principal component analysis, a parametric model of the foot is established by combining the standard mesh template, and graph structure features are constructed based on the point cloud topology of the foot parametric model. A multi-head attention mechanism neural network with fused graph structure features is constructed. The multi-head attention mechanism neural network is trained with a foot color image as input and a foot triangular mesh model as labeled samples. After the multi-head attention mechanism neural network outputs the target foot mesh coordinates, the corresponding vertex coordinates are indexed by the measurement mapping matrix to extract foot anthropometric indicators. Wherein, the measurement mapping matrix is a complete foot point cloud sequence index matrix corresponding to foot anthropometric indicators; the graph structure feature is a tree diagram constructed based on the point cloud topology relationship of the foot parametric model.
2. The foot parameter measurement and extraction method based on a parametric model according to claim 1, characterized in that, The industrial cameras consist of four units, arranged in a rectangular layout with all units at the same height from the ground.
3. The method for measuring and extracting foot parameters based on a parametric model according to claim 1, characterized in that, Obtaining camera intrinsic and extrinsic parameters in a unified coordinate system includes: A calibration board is placed at the center of the camera array, and the Zhang Zhengyou calibration method is used to jointly calibrate the intrinsic and extrinsic parameters of multiple cameras to obtain the intrinsic and extrinsic parameters of the cameras in a unified coordinate system.
4. The method for measuring and extracting foot parameters based on a parametric model according to claim 1, characterized in that, Based on camera intrinsic and extrinsic parameters and foot depth images, a multi-view initial point cloud is generated, and based on the multi-view initial point cloud, a complete foot point cloud is generated. The process of converting multi-view initial point clouds into a foot triangular mesh model includes: Based on the obtained camera intrinsic and extrinsic parameters and foot depth images, the three-dimensional coordinates of corresponding points of pixels are calculated using a stereo matching algorithm to generate multi-view initial point clouds. The spatial relationship of the camera intrinsic and extrinsic parameters is used to transform the multi-view initial point clouds to the same coordinate system. After deduplication and hole filling, the points are fused into a complete foot point cloud. Then, the Poisson surface reconstruction algorithm is applied to the complete foot point cloud. The isosurface is extracted by constructing a normal vector field and solving the Poisson equation, and the point cloud is transformed into a closed foot triangular mesh model.
5. The method for measuring and extracting foot parameters based on a parametric model according to claim 1, characterized in that, To obtain a standardized foot mesh, the process involves aligning the foot triangular mesh model with a standard mesh template at the vertex level using a registration algorithm and a deformation algorithm. The registration algorithm is the nearest iteration point algorithm, and the deformation algorithm is an area-preserving deformation algorithm.
6. The method for measuring and extracting foot parameters based on a parametric model according to claim 1, characterized in that, Based on the complete foot point cloud, anthropometric parameters were calculated, including: Based on the complete foot point cloud, foot anthropometry indices are calculated using a local extremum detection algorithm. The local extremum detection algorithm is a foot anthropometric index obtained by calculating the foot length by indexing the maximum and minimum differences of the coordinates of the complete foot point cloud in the foot length direction and calculating the arch height by detecting the height of the apex of the arch region. The foot anthropometric parameters include at least: foot length and arch height.
7. The method for measuring and extracting foot parameters based on a parametric model according to claim 1, characterized in that, The vertex displacement field of the normalized foot mesh relative to the standard mesh template was calculated, and the shape principal component matrix was extracted by principal component analysis, including: The vertex displacement field of the standardized foot mesh relative to the standard mesh template is calculated, the vertex displacement field is flattened into a high-dimensional vector and combined into a displacement matrix, and the shape principal component matrix is extracted after centering, covariance matrix calculation and eigenvalue filtering of the displacement matrix.
8. The method for measuring and extracting foot parameters based on a parametric model according to claim 1, characterized in that, The subjects were adolescents, and the foot parametric model was constructed based on standardized foot mesh data of several adolescents to characterize the growth characteristics of adolescent feet and the shape changes of unhealthy foot types.
9. A foot parameter measurement and extraction system based on a parametric model, characterized in that, include: The data acquisition module is used to set up an industrial camera according to a preset layout and acquire the camera's internal and external parameters in a unified coordinate system. Based on the camera's internal and external parameters, the module synchronously acquires the subject's foot depth image and foot color image. The mesh generation module is used to generate multi-view initial point clouds based on camera intrinsic and extrinsic parameters and foot depth images, and to generate complete foot point clouds based on multi-view initial point clouds; and to convert multi-view initial point clouds into foot triangular mesh models. The mesh processing module is used to filter standard mesh templates based on the foot triangular mesh model, and align the foot triangular mesh model with the standard mesh templates at the vertex level to obtain a standardized foot mesh. Based on the complete foot point cloud, anthropometry indicators are calculated, a linear relationship between the vertex coordinates of the standardized foot mesh and the anthropometry indicators is constructed, and a measurement mapping matrix is generated. The parameterization module is used to calculate the vertex displacement field of the standardized foot mesh relative to the standard mesh template, extract the shape principal component matrix through principal component analysis, establish a foot parameterization model in combination with the standard mesh template, and construct graph structure features based on the point cloud topology of the foot parameterization model. The network extraction module is used to construct a multi-head attention mechanism neural network that integrates graph structure features. The multi-head attention mechanism neural network is trained with a foot color image as input and a foot triangular mesh model as labeled samples. After the multi-head attention mechanism neural network outputs the target foot mesh coordinates, the corresponding vertex coordinates are indexed by the measurement mapping matrix to extract foot anthropometric indicators. Wherein, the measurement mapping matrix is a complete foot point cloud sequence index matrix corresponding to foot anthropometric indicators; the graph structure feature is a tree diagram constructed based on the point cloud topology relationship of the foot parametric model.
Citation Information
Patent Citations
Measuring method for foot three-dimensional foot-type information and three-dimensional reconstruction model by means of RGB-D camera
CN103971409A
Human body measurement method based on 3D camera
CN117652740A