Multi-view auricle three-dimensional reconstruction system for multiple tasks
By combining deformation-aware multi-view registration and anatomical prior-guided point cloud enhancement with a task-aware attention mechanism, the problems of non-rigid deformation and multi-task application in auricular 3D reconstruction were solved, achieving high-precision, high-quality auricular 3D reconstruction and intelligent multi-task application.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- UNIV OF SCI & TECH BEIJING
- Filing Date
- 2025-11-07
- Publication Date
- 2026-04-24
AI Technical Summary
Existing methods for 3D reconstruction of the auricle do not consider non-rigid deformation during image registration, resulting in insufficient registration and fusion accuracy; point cloud reconstruction lacks anatomical prior guidance, leading to low accuracy in reconstruction details; and multi-task applications lack a unified framework and feature sharing, resulting in computational redundancy and inefficiency.
We employ a deformation-aware multi-view registration mechanism, combined with a point cloud enhancement method based on anatomical prior knowledge base, and achieve adaptive feature representation through a task-aware attention mechanism to construct a unified multi-task processing framework.
It achieves high-precision, high-quality, and high-efficiency three-dimensional reconstruction of the auricle, solves the problems of non-rigid deformation modeling, accurate expression of anatomical structure, and unified feature representation for multiple tasks, and improves the overall efficiency and effectiveness of multi-task applications.
Smart Images

Figure CN121921433A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of auricular three-dimensional reconstruction technology, and in particular to a multi-view auricular three-dimensional reconstruction system for multi-task applications. Background Technology
[0002] Three-dimensional auricular information has significant application value in fields such as biometrics, 3D personalized printing, wearable device customization, auricular visualization, and film and television special effects. 3D auricular models provide precise anatomical information, unaffected by lighting, posture, or occlusion, and offer stronger feature representation capabilities compared to 2D images. In biometrics, 3D auricular features comprehensively reflect the spatial structure of the ear, exhibiting higher stability and anti-counterfeiting capabilities compared to 2D images, effectively improving the accuracy and security of identity authentication. In ear plastic surgery, 3D auricular models accurately recreate the three-dimensional structural features of the ear, offering higher geometric accuracy and operational feasibility compared to 2D image data, providing reliable data support for clinical applications such as auricular surgery planning, personalized prosthetic ear customization, and 3D printing. In auricular visualization, 3D morphological analysis of the auricle not only aids in the precise location of acupoints but also more clearly identifies complex spatial abnormalities such as bulges and cord-like structures, demonstrating greater accuracy and clinical value compared to 2D planar analysis. In the consumer electronics field, 3D auricular data can capture geometric information of key parts such as the concha and helix, which is more accurate than 2D data and can be used for personalized design and comfort optimization of wearable devices such as customized headphones and hearing aids.
[0003] Compared to methods that directly acquire 3D auricular data using 3D scanning instruments, 3D auricular reconstruction based on 2D multi-view images offers greater flexibility and practicality. While 3D scanners provide high-precision geometric information, their high equipment cost, complex operation, long scanning time, and stringent environmental requirements make them unsuitable for large-scale applications. This is particularly true in medical applications such as post-operative recovery assessment after auricular reconstruction, long-term follow-up monitoring in auricular visual diagnosis, and daily health management, where the use of specialized 3D scanning equipment is inconvenient for patients. In contrast, 2D image-based reconstruction methods require only a regular camera or smartphone for data acquisition, offering significant advantages such as low cost, convenience, and rapid deployment. Furthermore, it allows patients to collect data at home, enabling remote medical monitoring and providing a convenient technological pathway for long-term follow-up and daily health management.
[0004] Compared to the relatively mature 3D facial reconstruction, research on 3D auricular reconstruction is still in its developmental stage. Early 3D deformation modeling (3DMM) methods were based on PCA for parametric modeling, but due to the linear assumption, they struggled to handle nonlinear geometric structures. Subsequent studies, while introducing techniques such as keypoint detection and data augmentation, still lack robust 2D-to-3D constraints and curvature guidance, failing to capture complex nonlinear geometric changes. In recent years, deep learning-based methods such as HERA, Cycle3D, and Image2Life have made progress, but they are still limited by simplified geometric constraints and lack effective mechanisms for complex anatomical variations.
[0005] The task of 3D reconstruction of the auricle based on multi-view images involves key issues such as image registration, point cloud reconstruction, and multi-task application. However, existing methods still have the following problems:
[0006] (1) In terms of image registration, existing methods do not consider the non-rigid deformation characteristics of the auricle, resulting in insufficient registration and fusion accuracy. The cartilage and skin structure of the human ear itself gives it a certain degree of flexibility and elasticity. Existing multi-view 3D reconstruction methods are usually based on the assumption of rigid objects for registration and fusion, without fully considering the slight deformation of the auricle caused by external factors (such as changes in shooting angle, equipment contact, etc.) during the acquisition process. In addition, for complex 3D curved surface structures such as the helix, antihelix, concha, and triangular fossa, traditional feature point matching and registration algorithms are prone to feature registration deviations when dealing with the complex curved surfaces of the auricle. Furthermore, existing methods lack effective modeling and constraint mechanisms for the consistency of auricle deformation during multi-view fusion, and cannot effectively distinguish between real anatomical structural changes and non-rigid deformations introduced during the acquisition process, resulting in decreased registration accuracy, geometric distortion of the reconstruction model, surface discontinuities, and other problems, which affect the accuracy and quality of 3D reconstruction.
[0007] (2) In terms of point cloud reconstruction, existing methods lack anatomical prior guidance and are insufficient in noise processing, resulting in low accuracy of reconstruction details and unreasonable structures. Existing point cloud reconstruction methods mainly adopt a pure data-driven approach, lacking prior knowledge guidance on human ear anatomy (such as the concha, helix, tragus, etc.), resulting in insufficient accuracy of the reconstruction model in key anatomical details. In addition, the initial point cloud generated from multi-view images often has problems such as high noise, uneven density distribution, and sparse or missing point clouds in weak texture areas. Existing methods lack effective point cloud enhancement, denoising, and repair methods for human ear geometric features, and cannot fully consider its surface characteristics. Furthermore, existing technologies do not utilize the curvature changes between different anatomical parts of the auricle to resolve structural constraints, resulting in unreasonable structural phenomena such as abrupt changes in the curvature of anatomical parts in the reconstruction results, which reduces the quality and physiological rationality of the model.
[0008] (3) In terms of multi-task applications, the three-dimensional auricle model needs to serve multiple tasks such as identity recognition, 3D printing, auricular segmentation and key point localization, and film and television special effects. The requirements of different tasks are significantly different. However, existing methods are designed independently for a single task, lacking a unified framework and adaptive mechanism. When applying multiple tasks, there is computational redundancy and insufficient feature sharing, which limits the application efficiency.
[0009] For example, in traditional medical auxiliary diagnosis scenarios, point cloud structural feature extraction can be used to detect abnormal features of the auricle, such as bulges, depressions, nodules, and cords, to assist doctors in diagnosis. Simultaneously, the auricle's identification function can be used for medical record management, and precise segmentation of key areas such as the concha and antihelix can enable personalized customization of wearable devices related to auricular acupoints. In practical applications, these tasks often need to work collaboratively, establishing a multi-task mechanism. Multi-task collaboration not only improves the independent performance of each task but also optimizes the overall system performance through shared features and information flow. However, existing methods design each task independently, lacking a unified feature representation framework, leading to repeated feature extraction in multi-task applications and wasting computational resources. Furthermore, traditional 3D features have poor adaptability, unable to adjust the expression focus according to different scenario requirements, and lack an effective feature sharing mechanism between tasks, making it difficult to simultaneously guarantee global discriminative power and local precision within a unified framework, thus limiting the overall efficiency of multi-task applications. Summary of the Invention
[0010] To address the aforementioned problems, the present invention aims to provide a multi-view auricle 3D reconstruction system for multiple tasks. It solves the deformation modeling problem through a deformation-aware multi-view registration mechanism, improves point cloud quality and anatomical accuracy by integrating a point cloud enhancement method with an anatomical prior knowledge base, and achieves unified multi-task processing through an adaptive feature representation framework with task-aware attention. This enables high-precision, high-quality, and high-efficiency auricle 3D reconstruction and intelligent multi-task applications.
[0011] To solve the above-mentioned technical problems, the present invention provides the following technical solution: A multi-view three-dimensional reconstruction system for auricles oriented to multi-task, the system comprising: a multi-view registration and reconstruction network, an anatomical prior-guided point cloud enhancement module, and a multi-task-oriented feature extraction module; The multi-view registration and reconstruction network includes a deformation perception module and a progressive feature perturbation module. For the input multi-view auricle image, the deformation perception module extracts multi-view features of the auricle through spatial feature branches and frequency feature branches, and fuses them with the noise-perturbed features generated by the progressive feature perturbation module through a multi-view cross-attention mechanism to reconstruct the three-dimensional mesh model of the auricle image. The anatomical prior-guided point cloud enhancement module introduces the auricular anatomy prior knowledge base, which includes mean auricular point cloud, PCA principal components, key point locations and curvature features. Anatomical prior features are embedded through four independent feature codes to optimize the initially reconstructed 3D mesh model and generate enhanced 3D auricular point cloud. The multi-task-oriented feature extraction module generates task semantic vectors using a task encoder based on different downstream task requirements. It also dynamically adjusts the multi-scale feature weights of the 3D auricular point cloud through a cross-task perceptual attention mechanism, adaptively extracts and organizes feature representations, and applies the enhanced 3D auricular point cloud to multi-task scenarios.
[0012] Optionally, in the deformation perception module, the spatial feature branch extracts spatial features from multi-view auricle images, capturing local texture information, edge contours, and geometric details of the auricle through a multi-layer convolutional network to generate spatial features. ; The frequency domain feature branch performs a two-dimensional Fourier transform on multi-view auricle images, converting the images from the spatial domain to the frequency domain to obtain the spectral features of the images; then, a graph convolutional network is used to process the spectral features to extract global frequency domain features. .
[0013] Optionally, in the multi-view registration and reconstruction network, the spatial features obtained by the spatial feature branch are... Frequency domain features obtained from frequency domain feature branches After adaptive dynamic weighted fusion, the input to the multi-view feature extraction module generates unified multi-view features; Camera intrinsic parameters are introduced, including rotation matrix R, translation vector t, and scaling factor μ. The camera intrinsic parameters are embedded with feature information through feature encoding and alignment modules. At the same time, multi-view features are combined with position encoding embedded information, and enhanced features containing spatial position information are formed by fusion. Then, a query vector Q is generated through linear transformation. In the progressive feature perturbation module, firstly, multi-view image features are extracted from the input multi-view auricle image by an encoder, then layer normalization is performed to standardize the feature distribution, then feature enhancement is performed through an attention mechanism, and the dynamic range of the features is adjusted through scaling and translation operations; during the training phase, Gaussian perturbation noise is added to the processed features, and the perturbation features with added noise generate key vector K and value vector V. The query vector Q, key vector K, and value vector V are input into the multi-view cross-attention mechanism module. Attention weights are obtained by calculating the dot product of Q and K. After being normalized by softmax, the attention weights are weighted and aggregated with V to generate fused features. The fused features are input into the MLP network, which maps the high-dimensional features to a three-dimensional geometric space through multiple fully connected layers to generate shape feature vectors. Key vector K' and value vector V' are extracted from these vectors for subsequent decoding and reconstruction processes.
[0014] Optionally, the three-dimensional mesh model for reconstructing the auricle image specifically includes: The mesh generation rule encoding generates a regularized query vector Q' in three-dimensional space according to predefined rules. The predefined rules refer to discretizing the space within the boundary of a cube with the auricle center as the origin, according to a uniform mesh division strategy, to obtain a regularly distributed set of sampling points in three-dimensional space. The sampling points are subsequently used to guide the query of the three-dimensional feature space and the generation of mesh vertices, thereby realizing a structured three-dimensional reconstruction process. The query vector Q', along with the key vector K' and value vector V' extracted from the shape feature vector, are input into the decoder. The decoder uses cross-attention calculation to map the geometric information in the shape feature vector to the query point location. The decoded feature vector is then processed by a surface extraction algorithm to extract isosurfaces from the three-dimensional voxel mesh, converting the continuous implicit field representation into a discrete three-dimensional mesh model.
[0015] Optionally, the anatomical prior-guided point cloud enhancement module includes two branches, which respectively perform anatomical prior feature extraction and multi-scale point cloud feature extraction, and then generate an enhanced three-dimensional auricle point cloud through feature fusion. In the branch of anatomical prior feature extraction, an auricular anatomical prior knowledge base containing mean auricular point cloud, PCA principal component, key point location and curvature features is introduced. It is processed by different encoders to extract mean auricular point cloud features, PCA principal component features, key point features and curvature features respectively.
[0016] Optionally, the mean auricle point cloud feature is encoded by a multi-scale point cloud feature extraction module, specifically including: the multi-scale point cloud feature extraction module uses a multi-level point cloud convolutional network to extract local and global geometric features of the point cloud at different spatial scales, generating a mean auricle point cloud feature containing standard shape information. The PCA principal component feature extraction specifically includes: performing principal component decomposition on a preset number of auricle samples, then encoding them through an MLP network, mapping the principal component coefficients to high-dimensional feature vectors, and obtaining PCA principal component features; The keypoint features are encoded by a keypoint encoder, specifically including: labeling 3D keypoints; the keypoint coordinates (x, y, z) are first encoded by a coordinate encoding module to map the 3D spatial coordinates into a high-dimensional feature representation; simultaneously, the semantic information of the keypoints is input into a semantic encoding module to encode discrete semantic labels into continuous semantic feature representations; the key vector K'' and value vector V'' output from the coordinate encoding module and the query vector Q'' output from the semantic encoding module are input into an attention mechanism module; the attention mechanism module calculates the similarity between Q'' and K'' to generate attention weights, and then performs weighted aggregation with V'' to achieve the association mapping between semantic information and spatial location information; the output of the attention mechanism module is aggregated through global average pooling, and then processed through a fully connected layer and normalization to generate keypoint features; The curvature features are encoded by a curvature encoder, specifically including: generating a curvature heatmap by calculating the Gaussian curvature and average curvature of the point cloud surface; inputting the curvature heatmap into the curvature encoder; firstly, feature extraction is performed by the Gaussian curvature calculation module to capture the local curvature of the surface; then, it is processed by the point-by-point curvature encoding module, which uses 1D convolution operations to encode the curvature value sequence; the features after multiple 1D convolutions are input into a max pooling layer for feature aggregation; the pooled features are then transformed and mapped through a fully connected layer, finally outputting the curvature features.
[0017] Optionally, the extracted mean auricular point cloud features, PCA principal component features, key point features, and curvature features together constitute anatomical prior features, which are fused with multi-scale point cloud features. The fused features are modeled globally through a self-attention mechanism module. By calculating the correlation between different locations in the point cloud, global inconsistencies and structural defects in the point cloud are identified and corrected. The optimized point cloud features are then upsampled to map from low-resolution features back to high-resolution features of the original point cloud, restoring the complete sampling density of the point cloud and generating an enhanced three-dimensional auricular point cloud.
[0018] Optionally, in the multi-task-oriented feature extraction module, an enhanced 3D auricular point cloud is used as input. Through a multi-scale feature extraction module, a cross-task perception attention module guided by a task encoder, and a shared feature layer, the final output is a feature representation and application result optimized for different downstream tasks.
[0019] Optionally, the multi-scale feature extraction module employs a hierarchical point cloud processing network to extract geometric features of the point cloud at different spatial resolutions and receptive field scales, generating hierarchical multi-scale features that include local details and global structure. The extracted hierarchical multi-scale features are then input into the cross-task perception attention module for task-oriented feature adaptation processing. In the cross-task perception attention module, a task encoder is introduced to encode different downstream tasks and generate task-specific query vectors Q'''. At the same time, the hierarchical multi-scale features are used as input to the cross-task perception attention module and are transformed into key vectors K''' and value vectors V''' through linear transformation. The task encoder generates Q''' and performs a dot product operation with K''' generated by hierarchical multi-scale features to calculate the correlation score between the task and the features of each layer; the correlation score is normalized by softmax to generate attention weights, and the attention weights are then weighted with V''' to achieve task-oriented feature selection and fusion; The output features of the cross-task perception attention module are input into the MLP network for feature decoding and dimensionality transformation to generate task-adaptive features. These task-adaptive features are then input into the shared feature layer for further feature abstraction and representation learning. The features processed by the shared feature layer are distributed to specific output branches of the multi-task application module to achieve multi-task application.
[0020] Optionally, the multitasking application module includes: The system includes: an identity recognition branch for individual matching and authentication; a 3D printing branch for generating STL format files that meet printing requirements for subsequent ear reshaping customization; an ear acupoint segmentation unit for outputting segmentation mask maps to provide anatomical references for acupoint location; a key point detection unit for outputting key point coordinate sets to provide reference points for anatomical measurement and morphological analysis; and a visual effects unit for generating realistic renderable models for import into film and television production software and game engines.
[0021] The beneficial effects of the technical solution provided by this invention include at least the following: (1) Multi-view registration and reconstruction of auricular deformation perception; This invention addresses the low registration accuracy of existing multi-view 3D reconstruction methods for non-rigid deformation of the auricle and complex local surfaces. It proposes a multi-view registration and reconstruction method based on deformation perception and progressive feature perturbation fusion. This method extracts spatial geometric features from multi-view images through spatial domain feature branches and combines them with two-dimensional Fourier transform spectral features extracted through frequency domain feature branches to construct a multi-view feature fusion mechanism. Simultaneously, camera intrinsic parameters (rotation matrix R, translation vector t, scaling factor μ) are introduced for feature encoding and alignment, and positional encoding is embedded. Then, a multi-view cross-attention mechanism is used with a progressive feature perturbation module to calculate the correlation between features, establishing deformation consistency constraints between viewpoints. Finally, a high-quality 3D mesh model is generated through mesh generation rules and surface extraction algorithms. This invention solves the problems of registration deviation, geometric distortion, and surface discontinuity in traditional methods when dealing with complex surfaces such as the helix, antihelix, and concha, achieving deformation-perceptive multi-view registration and 3D reconstruction.
[0022] (2) Enhancement and optimization of multi-scale point clouds by integrating anatomical prior knowledge base; This invention addresses the problems of existing point cloud reconstruction methods, such as lack of anatomical prior guidance, poor initial point cloud quality, and unreasonable topological structure. It proposes a multi-scale point cloud enhancement and optimization method that integrates an anatomical prior knowledge base. This method constructs an anatomical prior knowledge base containing multi-dimensional information such as mean auricular point cloud, PCA principal components, key point locations, and curvature features, embedding the anatomical prior into the point cloud processing flow. It introduces anatomical prior feature embedding and a self-attention mechanism for multi-scale feature fusion, solving problems such as insufficient accuracy of key anatomical details, missing point clouds in weakly textured regions, and incorrect anatomical connectivity in the first-stage reconstructed point cloud. This ensures the anatomical accuracy, surface continuity, and physiological rationality of the reconstructed model.
[0023] (3) Adaptive multi-task feature extraction and representation based on task-aware attention mechanism; This invention addresses the problems of existing methods, such as independent design for single tasks, lack of a unified framework, and inability to adaptively adjust feature representations. It proposes an adaptive multi-task feature extraction and representation method based on a cross-task perceptual attention mechanism. Starting with high-quality 3D auricular point clouds, this method acquires hierarchical multi-scale features through a multi-scale feature extraction module and constructs a cross-task perceptual attention module containing hierarchical multi-scale features and a task encoder. A task description text encoder is designed to encode different task requirements (identity recognition, 3D printing, auricular segmentation, keypoint detection, and film and television special effects) into task-specific vectors. The feature weights are dynamically adjusted through the cross-task attention mechanism, enabling the MLP network to output shared feature layers optimized for different tasks. For different tasks, various differentiated results are output: identity recognition tasks output identity IDs, 3D printing tasks generate STL mesh files, auricular segmentation tasks provide segmentation masks, keypoint detection tasks locate keypoint coordinates, and film and television special effects tasks generate renderable models, etc. This invention solves the problems of repetitive computation, fixed feature representation, and low efficiency in multi-task applications in traditional methods. It achieves feature sharing and adaptive representation through a unified task-aware framework, ensuring global discriminability while taking into account local fineness, thus significantly improving the overall efficiency and effectiveness of multi-task applications. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1 This is a schematic diagram of the framework of the multi-view three-dimensional reconstruction system for auricles, which is oriented towards multiple tasks, provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of the overall structure of the multi-view three-dimensional reconstruction system for auricles, which is oriented towards multiple tasks, provided in an embodiment of the present invention. Figure 3 This is a schematic diagram of the structure of the multi-view registration and reconstruction network provided in the embodiments of the present invention; Figure 4 This is a schematic diagram of the structure of the anatomical prior-guided point cloud enhancement module provided in an embodiment of the present invention; Figure 5 This is a schematic diagram of the multi-scale point cloud feature extraction process provided in an embodiment of the present invention; Figure 6 This is a schematic diagram of the key point encoder principle provided in an embodiment of the present invention; Figure 7 This is a schematic diagram of the curvature encoder principle provided in an embodiment of the present invention; Figure 8This is a schematic diagram of the structure of the multi-task-oriented feature extraction module provided in an embodiment of the present invention. Detailed Implementation
[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention without creative effort are within the scope of protection of the present invention.
[0027] In embodiments of the present invention, words such as "exemplarily" and "for example" are used to indicate that something is an example, illustration, or description. Any embodiment or design described as "exemplary" in the present invention should not be construed as being more preferred or advantageous than other embodiments or designs. Rather, the use of the term "exemplary" is intended to present concepts in a specific manner. In embodiments of the present invention, "image" and "picture" may sometimes be used interchangeably, and it should be noted that their intended meanings are consistent unless the distinction is emphasized.
[0028] This invention provides a multi-view three-dimensional reconstruction system for auricles oriented towards multiple tasks, with reference to... Figure 1 and Figure 2 As shown, the system includes: a multi-view registration and reconstruction network, an anatomy-guided point cloud enhancement module, and a multi-task-oriented feature extraction module. For the acquired raw ear images, the ear region is first located and segmented to obtain multi-view auricle images, which are then used as inputs to the multi-view registration and reconstruction network.
[0029] The multi-view registration and reconstruction network includes a deformation perception module and a progressive feature perturbation module. For the input multi-view auricle image, the deformation perception module extracts multi-view features of the auricle through spatial feature branches and frequency feature branches, and fuses them with the noise-perturbed features generated by the progressive feature perturbation module through a multi-view cross-attention mechanism to reconstruct the three-dimensional mesh model of the auricle image. The anatomical prior-guided point cloud enhancement module introduces the auricular anatomy prior knowledge base, which includes mean auricular point cloud, PCA principal components, key point locations and curvature features. Anatomical prior features are embedded through four independent feature codes to optimize the initially reconstructed 3D mesh model and generate enhanced 3D auricular point cloud. The multi-task-oriented feature extraction module generates task semantic vectors using a task encoder based on different downstream task requirements. It also dynamically adjusts the multi-scale feature weights of the 3D auricular point cloud through a cross-task perceptual attention mechanism, adaptively extracts and organizes feature representations, and applies the enhanced 3D auricular point cloud to multi-task scenarios.
[0030] In this embodiment of the invention, a technical system is constructed that combines multi-view registration and reconstruction based on deformation perception and progressive feature perturbation, multi-scale point cloud enhancement and optimization fused with anatomical prior knowledge base, and adaptive multi-task feature extraction and representation based on task-aware attention mechanism. This system extends from multi-view image acquisition, camera pose calibration, deformation-aware registration, and anatomical prior-guided point cloud optimization to task-aware adaptive multi-task application. It solves key problems in existing technologies such as non-rigid deformation modeling of the auricle, accurate expression of anatomical structures, and unified feature representation for multiple tasks, and achieves high-precision, high-quality, and high-efficiency three-dimensional reconstruction of the auricle and intelligent multi-task application.
[0031] Furthermore, such as Figure 3 The diagram shows a multi-view registration and reconstruction network. The deformation perception module extracts multi-view features of the auricle through a dual-branch parallel feature extraction method, and fuses them with the progressive feature perturbation module through a multi-view cross-attention mechanism, thereby achieving end-to-end reconstruction from multi-view auricle images to a 3D mesh model.
[0032] The network input consists of multiple auricle images from different perspectives (view 1, view 2, view 3, view 4), and multimodal features are extracted from each perspective image through two parallel feature extraction branches (spatial domain feature branch and frequency domain feature branch).
[0033] The spatial feature branch extracts spatial features from multi-view auricle images, capturing local texture information, edge contours, and geometric details of the auricle through a multi-layer convolutional network to generate spatial features. Spatial domain features mainly focus on the spatial geometric information and local morphological features of an image.
[0034] The frequency domain feature branch performs a two-dimensional Fourier transform on multi-view auricle images, converting the images from the spatial domain to the frequency domain to obtain the spectral features of the images; then, a graph convolutional network is used to process the spectral features to extract global frequency domain features. Frequency domain features can characterize the global periodic patterns and frequency distribution characteristics of auricular images, complementing spatial domain features.
[0035] Furthermore, the spatial features obtained from the spatial feature branch Frequency domain features obtained from frequency domain feature branches The features are integrated through an adaptive dynamic weighted fusion operation, where α and β represent the learnable weight coefficients of the spatial and frequency domain features, respectively. The fused features simultaneously contain local details and global structural information. The fused features are then input into a multi-view feature extraction module, which integrates and registers features from four different perspectives to generate unified multi-view features that include complementary information from these different perspectives. .
[0036] To obtain the true size of the auricle, camera intrinsic parameters are introduced, including the rotation matrix R, translation vector t, and scaling factor μ. These parameters are used to unify image features from different viewpoints into the same three-dimensional spatial coordinate system, achieving scale recovery. The camera intrinsic parameters embed feature information through feature encoding and alignment modules to ensure spatial consistency of features from different viewpoints. Simultaneously, multi-view features are combined with position-encoded embedding information to form enhanced features containing spatial position information. This enhanced feature is then transformed linearly to generate a query vector Q, which is used for multi-view cross-attention mechanism computation.
[0037] To improve the network's ability to decode point clouds of auricle shape, a progressive feature perturbation module is introduced. By introducing noise perturbation into the feature space, the decoder is forced to learn robust feature representations that are insensitive to noise, thus enhancing the model's generalization ability to input variations. First, multi-view image features are extracted from the input multi-view auricle images using an encoder. Then, layer normalization is performed to standardize the feature distribution. Next, an attention mechanism is used for feature enhancement to capture the global correlation between features, and scaling and translation operations are used to adjust the dynamic range of the features. During the training phase, Gaussian perturbation noise is added to the processed features for noise enhancement; this is the forward process of the progressive feature perturbation module. The noise-enhanced perturbed features generate a key vector K and a value vector V, which contain feature information with noise perturbation.
[0038] Furthermore, the query vector Q, key vector K, and value vector V are input into the multi-view cross-attention mechanism module. Attention weights are obtained by calculating the dot product of Q and K, representing the correlation between the query location and the noisy features. After softmax normalization, the attention weights are weighted and aggregated with V to generate fused features, achieving multi-view feature fusion with feature perturbation characteristics. The network utilizes the cross-attention mechanism to denoise and reconstruct the noisy features, enabling the model to gradually recover clear geometric information from noisy features. The fused features are input into the MLP network, where multiple fully connected layers map the high-dimensional features to a three-dimensional geometric space, generating shape feature vectors. Key vector K' and value vector V' are extracted from these vectors for subsequent decoding and reconstruction.
[0039] The final process of decoding and reconstructing the 3D mesh model of the auricle image is as follows: The mesh generation rule encoding generates a regularized query vector Q' in three-dimensional space according to predefined rules. The predefined rules refer to discretizing the space within the boundary of a cube with the auricle center as the origin, according to a uniform mesh division strategy, to obtain a regularly distributed set of sampling points in three-dimensional space. The sampling points are subsequently used to guide the query of the three-dimensional feature space and the generation of mesh vertices, thereby realizing a structured three-dimensional reconstruction process.
[0040] The query vector Q', along with the key vector K' and value vector V' extracted from the shape feature vector, are input into the decoder. The decoder uses cross-attention calculation to map the geometric information in the shape feature vector to the query point location. The decoded feature vector is then processed by a surface extraction algorithm to extract isosurfaces in the three-dimensional voxel mesh, converting the continuous implicit field representation into a discrete three-dimensional mesh model.
[0041] Furthermore, such as Figure 4 The diagram shows a schematic of the anatomical prior-guided point cloud enhancement module. This module takes the initial reconstructed 3D mesh model generated by multi-view registration and reconstruction network as input, and uses prior knowledge of auricular anatomy to refine and optimize the initial reconstructed point cloud, correcting problems such as curvature error and local detail distortion in the initial 3D mesh model.
[0042] The anatomical prior-guided point cloud enhancement module comprises two branches: anatomical prior feature extraction and multi-scale point cloud feature extraction, respectively. Then, feature fusion is used to generate a refined, enhanced 3D auricular point cloud. The anatomical prior feature extraction branch incorporates an auricular anatomical prior knowledge base containing mean auricular point cloud, PCA principal component features, keypoint locations, and curvature features. This knowledge base is processed using different encoders to extract mean auricular point cloud features, PCA principal component features, keypoint features, and curvature features, respectively.
[0043] The first part is the extraction of mean auricular point cloud features. The mean auricular point cloud represents the statistical average shape of a large amount of real auricular data, providing a standard geometric template for the auricle. The mean auricular point cloud features are encoded using a multi-scale point cloud feature extraction module, such as... Figure 5 As shown, this module uses a multi-layered point cloud convolutional network to extract local and global geometric features of the point cloud at different spatial scales, generating mean auricular point cloud features containing standard shape information.
[0044] The second part is PCA principal component feature extraction. By performing principal component decomposition on a preset number (large number) of auricle samples, the main patterns of auricle shape changes are extracted, which include the statistical distribution characteristics of individual differences in auricle. Then, the principal component coefficients are encoded through a multilayer perceptron (MLP) network and mapped into high-dimensional feature vectors to obtain PCA principal component features, which are used to characterize the main changing trends and statistical laws of auricle shape.
[0045] The third part is keypoint feature extraction. Keypoint features are encoded using a keypoint encoder. First, 3D keypoints are labeled; these keypoints identify important landmarks on the auricle (such as points on important marginal sulci, anatomically important acupoints, etc.). For example... Figure 6 As shown, the key point coordinates (x, y, z) are first encoded by the coordinate encoding module to map the three-dimensional spatial coordinates into a high-dimensional feature representation; at the same time, the semantic information of the key points (such as the key point names such as "heart" and "stomach") are input into the semantic encoding module to encode the discrete semantic labels into a continuous semantic feature representation.
[0046] The key vector K'' and value vector V'' output from the coordinate encoding module, along with the query vector Q'' output from the semantic encoding module, are input into the attention mechanism module. The attention mechanism module calculates the similarity between Q'' and K'', generates attention weights, and then performs weighted aggregation with V'' to achieve a mapping between semantic information and spatial location information. The output of the attention mechanism module undergoes global average pooling for feature aggregation, and then passes through a fully connected layer and standardization to generate keypoint features. This feature vector simultaneously encodes the spatial location and anatomical semantic information of the keypoints.
[0047] The fourth part is curvature feature extraction, which characterizes the local geometric properties of the auricle surface. For example... Figure 7 As shown, curvature features are encoded using a curvature encoder. The input to the curvature encoder is a curvature heatmap, which is generated by calculating the Gaussian curvature and average curvature of the point cloud surface. This heatmap visualizes the concavity and convexity variations and local morphological features of the auricle surface using color coding. The curvature heatmap is input into the curvature encoder, first undergoing feature extraction via a Gaussian curvature calculation module to capture the local curvature of the surface. Then, it is processed by a point-by-point curvature encoding module, which uses 1D convolution operations to encode the curvature value sequence. The 1D convolution operation is repeated three times, with each convolutional layer capturing curvature change patterns at different scales. The features after multiple 1D convolutions are input into a max-pooling layer for feature aggregation. The pooled features undergo dimensionality transformation and feature mapping through a fully connected layer (FC), ultimately outputting curvature features. This feature vector encodes the complete curvature distribution information and local geometric characteristics of the auricle surface.
[0048] The extracted mean auricle point cloud features, PCA principal component features, keypoint features, and curvature features together constitute anatomical prior features, which are then fused with multi-scale point cloud features. Multi-scale point cloud feature extraction is as follows: Figure 5 As shown, when extracting features from multi-scale point clouds, a multi-level analysis is performed on the defects in the initial point cloud.
[0049] Finally, multi-scale point cloud features are fused with anatomical prior features. The fused features include data-driven geometric features and anatomically constrained structural priors, enabling the enhancement process to both repair local defects in the initial point cloud and ensure that the repaired point cloud conforms to the true anatomical patterns of the auricle.
[0050] The fusion feature is modeled globally through a self-attention mechanism module. By calculating the correlation between different locations in the point cloud, global inconsistencies and structural defects in the point cloud are identified and corrected. The optimized point cloud features are then upsampled (inverse distance weight interpolation) to map from low-resolution features back to high-resolution features of the original point cloud, restoring the complete sampling density of the point cloud and generating a refined and enhanced 3D auricular point cloud.
[0051] Furthermore, such as Figure 8 The diagram shows a multi-task-oriented feature extraction module. After obtaining the enhanced 3D auricle point cloud, to achieve adaptive representation and multi-task processing capabilities of the 3D auricle model in different application scenarios, this invention proposes a multi-task-oriented feature extraction module. This module, through a cross-task perceptual attention mechanism, adaptively extracts and organizes feature representations according to different downstream task requirements, enabling applications such as identity recognition, 3D printing, auricular segmentation, keypoint detection, and film and television special effects.
[0052] In the multi-task-oriented feature extraction module, enhanced 3D auricular point cloud is used as input. Through multi-scale feature extraction module, cross-task perception attention module guided by task encoder and shared feature layer processing, the final output is feature representation and application results optimized for different downstream tasks.
[0053] First, multi-scale feature extraction is performed on the input high-quality enhanced 3D auricular point cloud. The multi-scale feature extraction module uses a hierarchical point cloud processing network to extract the geometric features of the point cloud at different spatial resolutions and receptive field scales, generating hierarchical multi-scale features that include local details and global structure.
[0054] The extracted hierarchical multi-scale features are then input into a cross-task-aware attention module for task-oriented feature adaptation processing. In this module, a task encoder is introduced to encode representations for different downstream tasks. The task encoder receives a task description as input, which defines specific application requirements in text form, such as "identity recognition," "3D printing," and "ear acupoint segmentation." The task description is encoded using a text encoder, which employs a pre-trained language model to convert the natural language description into a high-dimensional task semantic vector. The encoded task semantic vector is then input into the task encoder module to generate a task-specific query vector Q'''. This query vector represents the current task's feature requirements and focus. The algorithm employs multi-task joint optimization, designing metric loss, geometric reconstruction loss, and segmentation loss for tasks such as identity recognition, 3D reconstruction, and ear acupoint segmentation respectively. A dynamic weighting mechanism balances the contributions of each task, enabling the task-aware attention module to learn adaptive feature representations that balance commonalities and specificities.
[0055] Simultaneously, hierarchical multi-scale features serve as input to the cross-task perceptual attention module, generating key vector K''' and value vector V''' through linear transformation. The task encoder's Q''' is multiplied by the hierarchical multi-scale features' generated K''' to calculate the relevance score between the task and each layer's features. This relevance score is then normalized using softmax to generate attention weights, which characterize the importance distribution of the current task to features at different levels. These attention weights are then weighted with V''' to achieve task-oriented feature selection and fusion. Through this mechanism, the network can adaptively emphasize important feature levels and geometric attributes while suppressing irrelevant feature components, based on the needs of different downstream tasks.
[0056] The output features of the cross-task-aware attention module are input into the MLP network for feature decoding and dimensionality transformation. The MLP maps the cross-task-aware features to a unified high-dimensional space through a multi-layer fully connected network, generating task-adaptive features. These feature vectors contain geometric information and semantic representations optimized for the current task, providing a foundation for subsequent multi-task processing. The task-adaptive features are then input into a shared feature layer for further feature abstraction and representation learning. The shared feature layer employs a multi-layer neural network structure, gradually enriching the feature representation through connections and feature transfer between layers. The shared feature layer enables different tasks to share the underlying geometric feature representations, improving the model's parameter efficiency and generalization ability. The features processed by the shared feature layer are then distributed to specific output branches of the multi-task application module, enabling multi-task applications.
[0057] Specifically, the multitasking application module includes: The system comprises several functionalities: an identity recognition branch for individual matching and authentication; a 3D printing branch for generating STL format files that meet printing requirements for subsequent customized auricular reshaping; an auricular acupoint segmentation unit for outputting segmentation mask maps to provide anatomical references for acupoint location; a keypoint detection unit for outputting keypoint coordinate sets to provide reference points for anatomical measurement and morphological analysis; and a visual effects unit for generating realistic, renderable models for import into film and television production software and game engines. Through multi-task-oriented feature extraction and adaptive representation, the system achieves multi-task feature sharing of the 3D reconstruction model across various application scenarios, moving beyond traditional single-application models and providing a unified application framework for fields such as biometric recognition, auricular reshaping, auricular diagnosis, and film and television production.
[0058] In summary, this invention constructs a unified framework from multi-view image acquisition, deformation-aware registration, anatomical prior-guided optimization to multi-task adaptive application through synergistic improvements in three core aspects: multi-view registration and reconstruction by fusing auricular deformation perception with progressive feature perturbation, multi-scale point cloud enhancement and optimization by introducing an anatomical prior knowledge base, and adaptive multi-task feature extraction and representation based on a cross-task perception attention mechanism. This solves problems such as auricular deformation modeling, accurate representation of anatomical structures, and unified feature representation for multiple tasks, and achieves high-fidelity 3D reconstruction of the auricle and multi-task applications.
[0059] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Unless otherwise specified, an element defined by the phrase "comprising..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.
[0060] The use of terms such as "an embodiment," "an embodiment," "an exemplary embodiment," and "some embodiments" in the specification indicates that the described embodiment may include a specific feature, structure, or characteristic, but not every embodiment necessarily includes that specific feature, structure, or characteristic. Furthermore, when a specific feature, structure, or characteristic is described in connection with an embodiment, implementing such a feature, structure, or characteristic in conjunction with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the art.
[0061] It should be understood that, in the various embodiments of the present invention, the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0062] In the embodiments provided by this invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0063] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0064] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0065] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0066] This invention encompasses any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of this invention. To provide the public with a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments; however, those skilled in the art will fully understand the invention even without these details. Furthermore, to avoid unnecessary misunderstanding of the essence of this invention, well-known methods, processes, procedures, components, and circuits are not described in detail.
[0067] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A multi-view three-dimensional reconstruction system for auricles oriented towards multiple tasks, characterized in that, The system includes: a multi-view registration and reconstruction network, an anatomy-guided point cloud enhancement module, and a multi-task-oriented feature extraction module; The multi-view registration and reconstruction network includes a deformation perception module and a progressive feature perturbation module. For the input multi-view auricle image, the deformation perception module extracts multi-view features of the auricle through spatial feature branches and frequency feature branches, and fuses them with the noise-perturbed features generated by the progressive feature perturbation module through a multi-view cross-attention mechanism to reconstruct the three-dimensional mesh model of the auricle image. The anatomical prior-guided point cloud enhancement module introduces the auricular anatomy prior knowledge base, which includes mean auricular point cloud, PCA principal components, key point locations and curvature features. Anatomical prior features are embedded through four independent feature codes to optimize the initially reconstructed 3D mesh model and generate enhanced 3D auricular point cloud. The multi-task-oriented feature extraction module generates task semantic vectors using a task encoder based on different downstream task requirements. It also dynamically adjusts the multi-scale feature weights of the 3D auricular point cloud through a cross-task perceptual attention mechanism, adaptively extracts and organizes feature representations, and applies the enhanced 3D auricular point cloud to multi-task scenarios.
2. The multi-view three-dimensional reconstruction system for auricles oriented towards multiple tasks according to claim 1, characterized in that, In the deformation perception module, the spatial feature branch extracts spatial features from multi-view auricle images. It captures local texture information, edge contours, and geometric details of the auricle through a multi-layer convolutional network to generate spatial features. ; The frequency domain feature branch performs a two-dimensional Fourier transform on multi-view auricle images, transforming the images from the spatial domain to the frequency domain to obtain the spectral features of the images. Furthermore, the spectral features are processed using a graph convolutional network to extract global frequency domain features. .
3. The multi-view three-dimensional reconstruction system for auricles oriented towards multiple tasks according to claim 1, characterized in that, In the multi-view registration and reconstruction network, the spatial features obtained by the spatial feature branch are... Frequency domain features obtained from frequency domain feature branches After adaptive dynamic weighted fusion, the input to the multi-view feature extraction module generates unified multi-view features; Camera intrinsic parameters are introduced, including rotation matrix R, translation vector t, and scaling factor μ. The camera intrinsic parameters are embedded with feature information through feature encoding and alignment modules. At the same time, multi-view features are combined with position encoding embedded information, and enhanced features containing spatial position information are formed by fusion. Then, a query vector Q is generated through linear transformation. In the progressive feature perturbation module, firstly, multi-view image features are extracted from the input multi-view auricle image by an encoder, then layer normalization is performed to standardize the feature distribution, then feature enhancement is performed through an attention mechanism, and the dynamic range of the features is adjusted through scaling and translation operations; during the training phase, Gaussian perturbation noise is added to the processed features, and the perturbation features with noise are used to generate key vector K and value vector V. The query vector Q, key vector K, and value vector V are input into the multi-view cross-attention mechanism module. The attention weights are obtained by calculating the dot product of Q and K. After the attention weights are normalized by softmax, they are weighted and aggregated with V to generate fused features. The fusion feature input MLP network maps high-dimensional features to three-dimensional geometric space through multiple fully connected layers, generating shape feature vectors, and extracting key vector K' and value vector V' from them for subsequent decoding and reconstruction processes.
4. The multi-view three-dimensional reconstruction system for auricles oriented towards multiple tasks according to claim 3, characterized in that, The 3D mesh model for reconstructing the auricle image specifically includes: The mesh generation rule encoding generates a regularized query vector Q' in three-dimensional space according to predefined rules. The predefined rules refer to discretizing the space within the boundary of a cube with the auricle center as the origin, according to a uniform mesh division strategy, to obtain a regularly distributed set of sampling points in three-dimensional space. The sampling points are subsequently used to guide the query of the three-dimensional feature space and the generation of mesh vertices, thereby realizing a structured three-dimensional reconstruction process. The query vector Q', along with the key vector K' and value vector V' extracted from the shape feature vector, are input into the decoder. The decoder uses cross-attention calculation to map the geometric information in the shape feature vector to the query point location. The decoded feature vector is then processed by a surface extraction algorithm to extract isosurfaces from the three-dimensional voxel mesh, converting the continuous implicit field representation into a discrete three-dimensional mesh model.
5. The multi-view three-dimensional reconstruction system for auricles oriented towards multiple tasks according to claim 1, characterized in that, The anatomical prior-guided point cloud enhancement module includes two branches, which perform anatomical prior feature extraction and multi-scale point cloud feature extraction respectively, and then generate an enhanced three-dimensional auricle point cloud through feature fusion. In the branch of anatomical prior feature extraction, an auricular anatomical prior knowledge base containing mean auricular point cloud, PCA principal component, key point location and curvature features is introduced. It is processed by different encoders to extract mean auricular point cloud features, PCA principal component features, key point features and curvature features respectively.
6. The multi-view three-dimensional reconstruction system for auricles oriented towards multiple tasks according to claim 5, characterized in that, The mean auricle point cloud feature is encoded by a multi-scale point cloud feature extraction module, which specifically includes: the multi-scale point cloud feature extraction module uses a multi-level point cloud convolutional network to extract local and global geometric features of the point cloud at different spatial scales, generating a mean auricle point cloud feature containing standard shape information. The PCA principal component feature extraction specifically includes: performing principal component decomposition on a preset number of auricle samples, then encoding them through an MLP network, mapping the principal component coefficients to high-dimensional feature vectors, and obtaining PCA principal component features; The keypoint features are encoded by a keypoint encoder, specifically including: labeling 3D keypoints; the keypoint coordinates (x, y, z) are first encoded by a coordinate encoding module to map the 3D spatial coordinates into a high-dimensional feature representation; simultaneously, the semantic information of the keypoints is input into a semantic encoding module to encode discrete semantic labels into continuous semantic feature representations; the key vector K'' and value vector V'' output from the coordinate encoding module and the query vector Q'' output from the semantic encoding module are input into an attention mechanism module; the attention mechanism module calculates the similarity between Q'' and K'' to generate attention weights, and then performs weighted aggregation with V'' to achieve the association mapping between semantic information and spatial location information; the output of the attention mechanism module is aggregated through global average pooling, and then processed through a fully connected layer and normalization to generate keypoint features; The curvature features are encoded by a curvature encoder, specifically including: generating a curvature heatmap by calculating the Gaussian curvature and average curvature of the point cloud surface; inputting the curvature heatmap into the curvature encoder; firstly, feature extraction is performed by the Gaussian curvature calculation module to capture the local curvature of the surface; then, it is processed by the point-by-point curvature encoding module, which uses 1D convolution operations to encode the curvature value sequence; the features after multiple 1D convolutions are input into a max pooling layer for feature aggregation; the pooled features are then transformed and mapped through a fully connected layer, finally outputting the curvature features.
7. The multi-view three-dimensional reconstruction system for auricles oriented towards multiple tasks according to claim 5, characterized in that, The extracted mean auricular point cloud features, PCA principal component features, key point features, and curvature features together constitute the anatomical prior features, which are fused with multi-scale point cloud features. The fused features are modeled globally through a self-attention mechanism module. By calculating the correlation between different locations in the point cloud, global inconsistencies and structural defects in the point cloud are identified and corrected. The optimized point cloud features are then upsampled to map from low-resolution features back to high-resolution features of the original point cloud, restoring the complete sampling density of the point cloud and generating an enhanced three-dimensional auricular point cloud.
8. The multi-view three-dimensional reconstruction system for auricles oriented towards multiple tasks according to claim 1, characterized in that, In the multi-task-oriented feature extraction module, enhanced 3D auricular point cloud is used as input. Through multi-scale feature extraction module, cross-task perception attention module guided by task encoder and shared feature layer processing, the final output is feature representation and application results optimized for different downstream tasks.
9. The multi-view three-dimensional reconstruction system for auricles oriented towards multiple tasks according to claim 8, characterized in that, The multi-scale feature extraction module employs a hierarchical point cloud processing network to extract geometric features of point clouds at different spatial resolutions and receptive field scales, generating hierarchical multi-scale features that include local details and global structure. The extracted hierarchical multi-scale features are then input into the cross-task perception attention module for task-oriented feature adaptation processing. In the cross-task perception attention module, a task encoder is introduced to encode different downstream tasks and generate task-specific query vectors Q'''. At the same time, the hierarchical multi-scale features are used as input to the cross-task perception attention module and are transformed into key vectors K''' and value vectors V''' through linear transformation. The task encoder generates Q''' and performs a dot product operation with K''' generated by hierarchical multi-scale features to calculate the correlation score between the task and the features of each layer; the correlation score is normalized by softmax to generate attention weights, and the attention weights are then weighted with V''' to achieve task-oriented feature selection and fusion; The output features of the cross-task perception attention module are input into the MLP network for feature decoding and dimensionality transformation to generate task-adaptive features. These task-adaptive features are then input into the shared feature layer for further feature abstraction and representation learning. The features processed by the shared feature layer are distributed to specific output branches of the multi-task application module to achieve multi-task application.
10. The multi-view three-dimensional reconstruction system for auricles oriented towards multiple tasks according to claim 9, characterized in that, The multitasking application module includes: The system includes: an identity recognition branch for individual matching and authentication; a 3D printing branch for generating STL format files that meet printing requirements for subsequent ear reshaping customization; an ear acupoint segmentation unit for outputting segmentation mask maps to provide anatomical references for acupoint location; a key point detection unit for outputting key point coordinate sets to provide reference points for anatomical measurement and morphological analysis; and a visual effects unit for generating realistic renderable models for import into film and television production software and game engines.