A method and system for skin extraction

By using nonlinear regression fitting and interlayer geometric constraints, a skin cutting threshold and masking region are generated, which solves the problems of accuracy and consistency in skin extraction in CBCT images and achieves high-precision skin region extraction, which is suitable for data registration and fusion in digital oral diagnosis and treatment.

CN122134732AActive Publication Date: 2026-06-02BEIJING STOMATOLOGY HOSPITAL CAPITAL MEDICAL UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING STOMATOLOGY HOSPITAL CAPITAL MEDICAL UNIV
Filing Date
2026-05-08
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies for extracting facial skin tissue from CBCT images are easily affected by factors such as differences in scanning equipment parameters, patient tilt, and intraoral metal artifacts, resulting in poor consistency in the accuracy of skin extraction results and weak cross-sample adaptability, making it difficult to meet the needs of high-precision data registration and fusion.

Method used

By performing nonlinear regression fitting on the prior anatomical features of skin tissue in historical oral images, a skin cutting threshold is generated. Combined with the cross-sectional feature extraction and inter-layer geometric consistency constraints of the CBCT image to be processed, a target skin masking region is generated. Background value suppression and isosurface mesh processing are then performed to generate the skin extraction result.

Benefits of technology

It improves the accuracy, consistency, and stability of skin extraction results, reduces interference from irrelevant data, enhances the interlayer continuity and boundary integrity of the skin area, and meets the data registration and fusion requirements in digital oral diagnosis and treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122134732A_ABST
    Figure CN122134732A_ABST
Patent Text Reader

Abstract

This application belongs to the field of image processing and relates to a skin extraction method and system. The method includes: performing nonlinear regression fitting on the prior anatomical features of skin tissue in historical oral images to obtain a skin cutting threshold; acquiring a CBCT image of the oral cavity to be processed; extracting boundary features from the cross-sectional image layers of the CBCT image to be processed based on the skin cutting threshold to obtain a dynamic candidate masking region feature map; applying inter-layer geometric consistency constraints to the dynamic candidate masking region feature map to generate a target skin masking region; extracting the original pixel values ​​within the target skin masking region and suppressing their background values; performing isosurface evolution processing on the suppressed original pixel values ​​to generate an isosurface grid; and performing patch validity screening on the isosurface grid to generate the skin extraction result. This application can reduce the interference of irrelevant data on the extraction result during the skin extraction process and ensure the quality of skin tissue extraction as much as possible.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing, and more particularly to a skin extraction method and system. Background Technology

[0002] In the fields of digital orthodontics, orthognathic surgery, and facial aesthetic restoration, the precise registration and fusion of cone-beam computed tomography (CBCT) data and facial 3D scan data is a core prerequisite for achieving accurate design, simulation, and efficacy prediction of digital treatment plans. Among these, the precise extraction of facial skin tissue from CBCT 3D volumetric data is a crucial foundation for ensuring efficient and high-precision registration and fusion of the two types of data. With the widespread application of digital oral diagnostic technologies, the industry has placed higher demands on the automation level, result stability, and boundary accuracy of CBCT skin extraction.

[0003] Current mainstream techniques for facial skin extraction in CBCT mostly employ segmentation based on global grayscale clustering of 3D volumetric data. This involves performing indiscriminate clustering analysis on the grayscale distribution of the full-view CBCT volumetric data to divide the grayscale distribution intervals corresponding to different human tissues. Based on the tissue grayscale intervals obtained from clustering, this approach performs global binarization on the full-view CBCT 3D volumetric data and extracts the outermost closed connected components as the facial skin tissue region. To optimize the integrity of the segmentation results, this approach typically performs global morphological opening and closing operations and smoothing on the extracted connected components to eliminate local noise, holes, and burrs in the segmentation results.

[0004] This type of segmentation scheme based on global gray-level clustering is highly susceptible to differences in CBCT scanning equipment parameters, patient tilt, intraoral metal artifacts, and gray-level interference from non-target tissues behind the skull. It cannot stably lock the precise gray-level boundaries of facial skin tissue, resulting in poor consistency of skin extraction results, weak cross-sample adaptability, and problems such as inter-layer contour misalignment and blurred boundaries. It is difficult to meet the high-precision registration and fusion requirements of two types of data in digital oral diagnosis and treatment. Summary of the Invention

[0005] This application provides a skin extraction method and system, which aims to improve the accuracy and consistency of skin extraction results, reduce the interference of irrelevant data on the extraction results, and ensure the quality of skin tissue extraction as much as possible during the skin extraction process.

[0006] To address the aforementioned technical problems, this application provides the following technical solutions:

[0007] A skin extraction method, comprising: Step S1: Perform nonlinear regression fitting on the prior anatomical features of skin tissue in historical oral images to obtain the skin cutting threshold; Step S2: Obtain the oral CBCT image to be processed. Based on the skin cutting threshold, extract the boundary features of the cross-sectional image layer of the oral CBCT image to be processed to obtain a dynamic candidate masking region feature map. Apply inter-layer geometric consistency constraints to the dynamic candidate masking region feature map to generate the target skin masking region. Step S3: Extract the original pixel values ​​within the target skin masking area and suppress their background values. Perform isosurface evolution processing on the suppressed original pixel values ​​to generate an isosurface mesh. Perform patch validity screening processing on the isosurface mesh to generate the skin extraction result.

[0008] On the other hand, a skin extraction system is also provided, including: The threshold generation module is used to obtain the skin cutting threshold by performing nonlinear regression fitting on the prior anatomical features of skin tissue in historical oral images. The masking region generation module is used to acquire the oral CBCT image to be processed, extract the boundary features of the cross-sectional image layer of the oral CBCT image to be processed based on the skin cutting threshold to obtain a dynamic candidate masking region feature map, and perform inter-layer geometric consistency constraints on the dynamic candidate masking region feature map to generate the target skin masking region. The skin result generation module is used to extract the original pixel values ​​within the target skin masking area and suppress their background values. The suppressed original pixel values ​​are then subjected to isosurface evolution processing to generate an isosurface mesh. Finally, the isosurface mesh is subjected to patch validity screening to generate the skin extraction result.

[0009] The beneficial effects of this application are as follows: 1. This application obtains the skin cutting threshold by using nonlinear regression fitting based on the prior anatomical features of skin tissue from historical oral images. Unlike the traditional approach of indiscriminate global gray-level clustering, this approach relies on the prior anatomical information of oral skin tissue to reduce the adverse effects of interference factors such as differences in scanning equipment parameters and intraoral artifacts on threshold determination. This improves the adaptability of the skin cutting threshold across different samples and makes the locking of skin gray-level boundaries highly stable, providing a more stable processing benchmark for subsequent skin region extraction. This approach alleviates the problems of low consistency and weak cross-sample adaptability of skin extraction results in traditional approaches from the threshold determination stage.

[0010] 2. Based on the obtained skin cutting threshold, this application performs boundary feature extraction and inter-layer geometric consistency constraint on the cross-sectional image layer of the oral CBCT image to be processed to obtain the target skin masking region. Compared with the traditional method of global binarization of the entire field of view volume data, the boundary feature extraction of the cross-sectional image layer can focus on the target area of ​​facial skin and reduce the interference of non-target tissues behind the skull on the segmentation results. The inter-layer geometric consistency constraint can alleviate the problem of inter-layer contour misalignment caused by patient tilt, improve the inter-layer continuity of the skin area, and obtain a target skin masking region with a relatively complete boundary. At the same time, it reduces the effective range of subsequent processing and further reduces the interference of irrelevant data on the extraction results.

[0011] 3. This application first performs background value suppression processing on the original pixels within the target skin masking area, and then performs isosurface mesh generation and patch validity screening to generate skin extraction results. Compared with the traditional approach of performing overall morphological smoothing on the global segmentation results, background value suppression on the original pixels within the target area can reduce the adverse effects of background noise on subsequent mesh generation. Then, through isosurface mesh generation and patch validity screening, invalid mesh patches can be eliminated, improving the topological continuity and boundary integrity of the skin extraction results. This further alleviates the problem of blurred boundaries in the skin extraction results of traditional approaches, making it more suitable for the application needs of CBCT and facial scan data registration and fusion in digital oral diagnosis and treatment. Attached Figure Description

[0012] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0013] Figure 1 This is a schematic diagram of the overall process of a skin extraction method provided in an embodiment of this application; Figure 2 This is a schematic diagram illustrating the generation of sampling contour lines for each layer of cross-sectional image from a top-view perspective in a skin extraction method provided in this application embodiment; Figure 3 This is a schematic diagram of the constrained closed contour of the skin and the target skin mask sub-region in each layer of the cross-sectional image of a skin extraction method provided in this application embodiment; Figure 4 This is a schematic diagram illustrating triangular mesh transformation and patch selection processing in a skin extraction method provided in an embodiment of this application. Detailed Implementation

[0014] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0015] Reference Figure 1 This application provides a skin extraction method, including: Step S1: Perform nonlinear regression fitting on the prior anatomical features of skin tissue in historical oral images to obtain the skin cutting threshold; Step S2: Obtain the oral CBCT image to be processed. Based on the skin cutting threshold, extract the boundary features of the cross-sectional image layer of the oral CBCT image to be processed to obtain a dynamic candidate masking region feature map. Apply inter-layer geometric consistency constraints to the dynamic candidate masking region feature map to generate the target skin masking region. Step S3: Extract the original pixel values ​​within the target skin masking area and suppress their background values. Perform isosurface evolution processing on the suppressed original pixel values ​​to generate an isosurface mesh. Perform patch validity screening processing on the isosurface mesh to generate the skin extraction result.

[0016] Optionally, step S1 includes: Step S11: Obtain historical oral cavity images and perform multi-scale gradient field mapping processing on the historical oral cavity images to construct a gray-scale gradient evolution map of skin tissue. Step S12: Determine the multidimensional morphological feature points of the skin tissue based on the gray-scale gradient evolution map of the skin tissue, and generate prior anatomical features of the skin tissue accordingly. Step S13: Input the generated prior anatomical features of the skin tissue into a preset tissue-air interface grayscale anchoring mapping model for nonlinear regression fitting to obtain the skin cutting threshold.

[0017] Optionally, the step of determining multidimensional morphological feature points of skin tissue based on the gray-level gradient evolution map of skin tissue, and generating prior anatomical features of skin tissue accordingly, includes: Step S121: Determine the multidimensional morphological feature points of the skin tissue based on the gray-scale gradient evolution map of the skin tissue; Step S122: Perform cross-scale anatomical topology nesting processing on the multidimensional morphological feature points of the skin tissue to obtain a cross-scale nested feature set; Step S123: Generate prior features of skin tissue anatomy based on the feature set nested across scales.

[0018] Preferably, in step S11, the object of processing is the historical oral head cone-beam CT (CBCT for short, an imaging device that acquires three-dimensional structural data of the oral head through cone-beam X-ray scanning) data that has been registered with the facial three-dimensional surface scan data in the context of digital oral diagnosis and treatment. The historical oral images are the set of such three-dimensional volume data.

[0019] First, historical oral CBCT 3D volumetric data were acquired from multiple centers and multiple devices. The acquired volumetric data were then screened for validity, and invalid samples with severe metal artifacts or scan ranges that did not cover the entire facial area were removed.

[0020] Subsequently, the selected valid samples were normalized using Hounsfield Units (HU, a standardized unit characterizing the attenuation coefficient of tissue to X-rays in CT images; the HU value of air is approximately -1000, the HU value of water is 0, and the HU value of bone is usually greater than 400). This normalization mapped the HU values ​​of all samples to a preset grayscale range of [-1000, 1000], resulting in standardized historical oral cavity images. The normalization method used here is min-max normalization, calculated as follows: Normalized HU value = (Original HU value - Minimum value in the range) / (Maximum value in the range - Minimum value in the range) × 2000 - 1000. For example, if an original pixel has a HU value of 0, a minimum value of -1000, and a maximum value of 1000, it will still be 0 after normalization; if an original pixel has a HU value of -2000, it will be -1000 after normalization.

[0021] Subsequently, the standardized historical oral images are subjected to spatial resolution normalization. Three-dimensional linear interpolation is used to resample the voxel sizes of all samples to a preset baseline voxel size (e.g., isotropic 0.3mm × 0.3mm × 0.3mm) to eliminate the impact of differences in scanning resolution between different devices on subsequent multi-scale analysis. Then, a Gaussian pyramid multi-scale space is constructed from the spatially normalized historical oral images to obtain multi-scale volume data at different resolutions. The Gaussian pyramid multi-scale space is a series of volume data sets with progressively decreasing resolution obtained by performing multiple Gaussian smoothing and downsampling operations on the original three-dimensional volume data. Each layer is called a scale; the larger the scale, the lower the volume data resolution and the more global facial contour features are preserved; the smaller the scale, the higher the resolution and the more skin boundary details are preserved. The specific construction process is as follows: Using standardized historical oral images as the 0th layer of the pyramid (scale 0, original resolution); smoothing the volume data of the 0th layer with a 3×3×3 Gaussian kernel with a standard deviation σ=1.0, then performing half-sampling in the cross-sectional image layer plane to obtain the 1st layer of the pyramid (scale 1, cross-sectional image layer resolution is 1 / 2 of the original); the smallest imaging unit of the original cross-sectional image layer is the original high-resolution pixel, belonging to the native resolution cross-sectional image layer of the standardized historical oral image obtained in step S1. Each original high-resolution pixel corresponds one-to-one with a three-dimensional voxel in the three-dimensional volume data of the standardized historical oral image, and the gray value of the pixel is equal to the HU value of the corresponding voxel; the smallest imaging unit of the low-resolution cross-sectional image layer obtained after downsampling is the downsampled pixel, belonging to the multi-scale volume data of the corresponding scale; repeating the above smoothing + downsampling operation, a total of 3-5 scales of Gaussian pyramids are constructed, for example: the original cross-sectional image layer resolution is 512×512, the resolution of scale 1 is 256×256, and the resolution of scale 2 is 128×128.

[0022] Finally, for the volume data at each scale, the grayscale gradient magnitude and gradient direction of each pixel are calculated along the radial direction from the center of the head to the front of the face within the cross-sectional image layer. The center point of the head tissue is the centroid coordinate of the connected domain of the head tissue within that cross-sectional image layer. The gradient magnitude is calculated using the following formula: Gradient magnitude = √[(current pixel HU value - previous pixel HU value in the radial direction)² + (current pixel HU value - next pixel HU value in the radial direction)²]. The core logic of the gradient calculation here is: the gradient magnitude represents the degree of change in the grayscale value of a pixel. The grayscale value at the skin-air interface will suddenly change from the high HU value of the skin tissue to the low HU value of the air, so the gradient magnitude will have a significant peak. The gradient direction represents the direction of the grayscale value change. The gradient direction at the skin-air interface is fixed as the radial direction from the center of the head to the front of the face. The gradient calculation results at all scales are stacked along the slice axis from the top of the skull to the mandible to construct a grayscale gradient evolution map of skin tissue. The horizontal axis of the grayscale gradient evolution map of skin tissue corresponds to the HU value range of CBCT pixels (-1000 to 1000), and the vertical axis corresponds to the statistical frequency of gradient amplitude (0 to 100%). The elements at the intersection of rows and columns in the map represent the statistical frequency of the corresponding gradient amplitude at the skin-air interface under the corresponding HU value. For example, at the position where the horizontal axis HU value is 0 and the vertical axis gradient amplitude is 200, the element value is 85%, which means that in 85% of the samples, when the HU value of the skin-air interface is 0, the gradient amplitude is 200.

[0023] Preferably, in step S121, the object being processed is the skin tissue grayscale gradient evolution map obtained in step S11.

[0024] First, extract the HU value range corresponding to the peak values ​​whose gradient amplitude statistical frequency is greater than a preset threshold (usually 60%) from the gray-scale gradient evolution map of skin tissue to obtain the candidate feature range of skin tissue in the gray-scale dimension. For example, if the gradient peaks are concentrated in the range of HU values ​​from -50 to 50, this range is the candidate feature range of the gray-scale dimension.

[0025] Subsequently, for the standardized historical oral image obtained in step S11, within the gray-scale skin tissue candidate feature range, gradient magnitude abrupt change points (pixels with gradient magnitude greater than a preset threshold) along the radial direction in the cross-sectional image layer are extracted to obtain interface candidate feature points in the gradient dimension.

[0026] Subsequently, based on the anatomical prior information of the oral and head skin tissue, spatial filtering was performed on the interface candidate feature points in the gradient dimension. Only feature points located within the anatomical distribution range of the facial skin on the front half of the head were retained, while invalid feature points in the posterior cranial and cervical spine regions were removed, resulting in valid feature points in the anatomical spatial dimension. The anatomical prior information here is that the human facial skin is only distributed on the front half of the head, with the line connecting the two external auditory canals as the boundary, and there is no need for skin extraction on the back half of the head. Finally, the effective feature points of the grayscale dimension, gradient dimension, and anatomical space dimension are integrated and categorized. Using the same three-dimensional spatial coordinates as the sole anchoring reference, the three-dimensional feature attributes describing the physical location of the same skin interface are bound and fused to form a multi-dimensional feature vector F_i=[x_i,y_i,z_i,HU_i,Grad_i,Anat_i], where (x_i, y_i, z_i) are spatial coordinates, HU_i is the grayscale value, Grad_i is the gradient magnitude, and Anat_i is the anatomical space identifier (e.g., 0 / 1 indicates whether it is on the front of the face). Invalid features that do not meet the requirement of multi-dimensional synchronization are eliminated, and finally a four-dimensional feature set is formed in which each feature point carries spatial coordinate attributes + grayscale dimension attributes + gradient dimension attributes + anatomical space attributes, which is the multi-dimensional morphological feature point of skin tissue.

[0027] Preferably, in step S122, the processing object is the multidimensional morphological feature points of skin tissue obtained in step S121.

[0028] First, based on the fixed anatomical topological hierarchy of oral and facial skin tissue, three feature levels are defined from coarse to fine: the first level is the global facial contour topological level, corresponding to the overall morphological features of the outer contour of the entire anterior facial skin; the second level is the local organ neighborhood topological level, corresponding to the skin contour features around key facial organs such as the perinasal, perioral, and periorbital regions; and the third level is the pixel-level interface topological level, corresponding to the single-pixel-level gradient mutation features of the skin-air interface.

[0029] Subsequently, the multidimensional morphological feature points of the skin tissue were classified according to the three anatomical topological levels mentioned above, resulting in a global contour feature point set, a local neighborhood feature point set, and a pixel-level interface feature point set.

[0030] Subsequently, cross-scale anatomical topology nesting processing is performed on the feature point sets of the three levels. This cross-scale anatomical topology nesting processing maps fine-grained low-level features to a topological framework constructed from coarse-grained high-level features, forming a nested feature structure of "global framework - local details - pixel boundaries," which differs from the traditional simple feature splicing. The specific implementation process is as follows: First, the overall topological framework of facial skin tissue is constructed using the global contour feature point set to determine the overall range and shape of the facial contour; Second, by calculating the spatial nearest neighbor distance between local neighborhood feature points and the global contour feature framework, the local neighborhood feature point set is mapped to the corresponding anatomical positions of the perinasal, perioral, and periorbital regions within the topological framework, supplementing the contour details of key facial areas; Third, again through spatial nearest neighbor search, the pixel-level interface feature point set is mapped to the corresponding positions of the local neighborhood feature point set, supplementing the single-pixel boundary details of the skin-air interface, ultimately obtaining the cross-scale nested feature set. For example, the global contour feature point set determines that the face as a whole is an elliptical frame, the local neighborhood feature point set supplements the raised contour around the nose within the frame, and the pixel-level interface feature point set supplements the precise boundary points between the skin and the air on the raised contour.

[0031] Preferably, in step S123, the object being processed is the cross-scale nested feature set obtained in step S122.

[0032] First, statistical measures are calculated for the feature point sets corresponding to the three anatomical topology levels in the cross-scale nested feature set. The statistical measures include the mean HU value, the mean gradient magnitude, and the feature point coordinate range (minimum and maximum values ​​in each coordinate axis direction, a total of 6 values) for each feature point set, so as to obtain the feature statistics corresponding to each level.

[0033] Subsequently, the feature statistics of the three levels are subjected to min-max normalization to eliminate the dimensional differences between different feature dimensions. The calculation formula for min-max normalization is: normalized feature value = (original feature value - minimum feature value of this dimension) / (maximum feature value of this dimension - minimum feature value of this dimension). This can uniformly map all feature values ​​to the interval [0,1]. For example, if the feature value range of a certain dimension is 0-200, the original feature value is 100, and the normalized value is 0.5.

[0034] Next, the normalized feature statistics are concatenated in the order of global facial contour topology, local organ neighborhood topology, and pixel-level interface topology, converting them into a fixed-dimensional one-dimensional vector. For example, each level contains one HU mean, one gradient mean, and six coordinate range values, for a total of (1+1+6)×3=24-dimensional feature vectors across the three levels. Finally, this one-dimensional vector is used as the prior feature of skin tissue anatomy.

[0035] Preferably, in step S13, the object being processed is the prior anatomical features of the skin tissue obtained in step S123, and the processing tool is a preset tissue-air interface grayscale anchoring mapping model.

[0036] First, the prior anatomical features of the skin tissue are input into a pre-trained tissue-air interface grayscale anchoring mapping model. Then, through two parallel branches of the model, global feature encoding and local boundary anchoring processing are performed on the prior anatomical features of the skin tissue, respectively, yielding the output features of the two branches. Next, the output features of the two branches are fused through the regression branch of the model, and nonlinear regression fitting with prior regularization constraints on the anterior anatomical distribution of the oral and facial skin is performed. Finally, the skin cutting threshold output by the model is obtained; this skin cutting threshold is a grayscale threshold specifically used for skin-air interface segmentation in the oral CBCT image to be processed, and is used for boundary feature extraction of subsequent cross-sectional image layers.

[0037] Optionally, the step of extracting boundary features from the cross-sectional image layer of the oral CBCT image to be processed based on the skin cutting threshold to obtain a dynamic candidate masking region feature map includes: Step S21: Extract all cross-sectional image layers from the oral CBCT image to be processed using the constructed volume data space segmentation operator; Step S22: Based on the skin cutting threshold, perform boundary feature extraction on each cross-sectional image layer to determine the spatial topological correlation between adjacent cross-sectional image layers; Step S23: Generate a dynamic candidate masking region feature map based on spatial topological correlation.

[0038] Optionally, generating a dynamic candidate masking region feature map based on spatial topological correlation includes: Step S231: Perform asymmetric gradient transition kernel excitation response processing on each cross-sectional image layer to extract the corresponding discrete boundary candidate response points. Step S232: Perform dilatational morphological topology enhancement processing on the discrete boundary candidate response points; Step S233: Based on spatial topological correlation, perform maximum connected component extraction on the enhanced discrete boundary candidate response points to generate a dynamic candidate masking region feature map.

[0039] Optionally, the step of applying inter-layer geometric consistency constraints to the feature map of the dynamic candidate masking region to generate the target skin masking region includes: Step S24: Calculate the spatial dimension attention weight distribution matrix between the feature maps of the dynamic candidate masking regions corresponding to different cross-sectional image layers; Step S25: Based on the spatial dimension attention weight distribution matrix, perform inter-layer geometric consistency constraint processing on the feature maps of the dynamic candidate masking regions corresponding to all cross-sectional image layers to obtain the target skin masking region.

[0040] Preferably, in step S21, the object to be processed is the oral CBCT image to be processed in the context of digital oral diagnosis and treatment, which needs to be registered and fused with facial three-dimensional surface scan data, and the processing tool is a pre-constructed volume data spatial segmentation operator.

[0041] First, invalid background cropping is performed on the oral CBCT image to be processed to remove blank edge areas without tissue information. The specific method of invalid background cropping is as follows: based on a preset tissue threshold (e.g., HU value greater than -500), the largest head tissue connected component is extracted through three-dimensional connected component analysis, and the axial bounding box of this connected component is calculated. The original volume data is then cropped according to this bounding box to reduce the data range for subsequent processing. Next, the cropped oral CBCT image to be processed undergoes HU value range normalization processing, mapping the HU value of the volume data to the [-1000, 1000] tissue grayscale range consistent with the standardized historical oral images in step S11, eliminating grayscale value differences caused by different scanning devices, resulting in a preprocessed oral CBCT image to be processed. Then, the preprocessed oral CBCT image to be processed is input into a volume data space segmentation operator, which sequentially performs coordinate system alignment, downsampling, ROI limitation, inter-layer association labeling, and layer-by-layer segmentation processing. Finally, all cross-sectional image layers output by the volume data space segmentation operator are obtained; the cross-sectional image layer is a two-dimensional tomographic image perpendicular to the skull base direction, which is the basic processing unit for oral CBCT skin extraction.

[0042] Preferably, in step S22, the processing objects are the skin cutting threshold obtained in step S13 and all cross-sectional image layers obtained in step S21.

[0043] First, for each cross-sectional image layer, the pixels in the cross-sectional image layer are initially screened by grayscale threshold based on the skin cutting threshold. Pixels with HU values ​​greater than or equal to the skin cutting threshold are marked as tissue pixels, and pixels with HU values ​​less than the skin cutting threshold are marked as air pixels, thus obtaining the initial candidate pixel set for the skin-air interface.

[0044] Subsequently, spatial clustering is performed on the initial candidate pixel set, grouping spatially adjacent tissue pixels into one class to obtain the initial skin contour point set corresponding to each cross-sectional image layer (for example, using the DBSCAN density clustering algorithm, setting the neighborhood radius to 3 pixels and the minimum number of neighborhood samples to 5, to group spatially adjacent tissue pixels into one class, thus obtaining the initial skin contour point set corresponding to each cross-sectional image layer). Then, along the layer-cutting axis from the top of the skull to the mandible, the spatial coordinate correspondence between the initial skin contour point sets corresponding to adjacent cross-sectional image layers is calculated to obtain the contour mapping relationship between adjacent layers; specifically, for each contour point in the current layer, the nearest contour point in the adjacent layer is found, establishing a one-to-one correspondence between the two points.

[0045] Finally, based on the contour mapping relationship between adjacent slice layers, the spatial topological correlation between adjacent cross-sectional image layers is determined; the smaller the distance between corresponding points, the higher the spatial topological correlation between the two contour layers, and vice versa.

[0046] Preferably, in step S231, the object being processed is each cross-sectional image layer obtained in step S21.

[0047] First, refer to Figure 2 For each cross-sectional image layer, radial sampling contour lines are extracted from the center point of the head tissue toward the front of the face within that layer, resulting in a sequence of pixel HU values ​​corresponding to each radial sampling contour line. The radial sampling contour lines are generated as follows: taking the center point of the head tissue in the cross-sectional image layer as the origin and the front of the face as the 0-degree direction, a ray is generated every 5 degrees within a range of ±90 degrees, which is the radial sampling contour line. A total of 36 contour lines are generated, and each contour line corresponds to a set of pixel HU values ​​arranged in sequence, i.e., a sequence of pixel HU values.

[0048] Subsequently, using an asymmetric fifth-order gradient template [1,1,0,-1,-1] that matches the unidirectional gray-level step characteristics of the skin-air interface, pixel-by-pixel sliding convolution is performed along the pixel HU value sequence to obtain the gradient response value corresponding to each sliding position (i.e., to realize asymmetric gradient transition kernel excitation response processing), forming a gradient response sequence; here, the core technology will be explained in detail: The design principle of the asymmetric fifth-order gradient template: The grayscale change at the skin-air interface is unidirectional, exhibiting a step characteristic of "high-high-medium-low-low" from the high HU value of the skin tissue to the low HU value of the air. Therefore, the template weight is designed as [1,1,0,-1,-1], where: the 0 weight bit at the center of the template is the anchor position of the skin-air interface, the two 1 weight bits in front of the center correspond to the grayscale accumulation area on the skin tissue side, and the two -1 weight bits behind the center correspond to the grayscale decay area on the air side. This design is different from the symmetric weight design of the traditional Sobel operator and adapts to the unidirectional grayscale change at the skin-air interface.

[0049] Sliding convolution calculation rules: Align the 0 weight bits of the template with each pixel in the pixel HU value sequence in turn, multiply each weight bit of the template with the corresponding pixel HU value, and the sum of all products is the gradient response value at that position; when the template slides to the beginning and end of the sequence, some weight bits have no corresponding pixel values, and the pixel values ​​at the beginning and end of the sequence are used to fill in the missing pixel values.

[0050] For example, consider a radial sampling contour line with a pixel HU value sequence of [200, 150, 180, 200, 0, -100, -200, -180]. When the template weight 0 is aligned with the 5th pixel (HU value 0), the correspondence between the template weight and the pixel value is: 1→180, 1→200, 0→0, -1→-100, -1→-200. The gradient response value is 180×1+200×1+0×0+(-100)×(-1)+(-200)×(-1)=680. Finally, the pixel corresponding to the gradient response peak in the gradient response sequence is extracted as the discrete boundary candidate response point corresponding to the cross-sectional image layer. In the example above, the pixel with a HU value of 0 corresponds to the gradient response peak of 680, and this pixel is the discrete boundary candidate response point.

[0051] Preferably, in step S232, the processing object is the discrete boundary candidate response point corresponding to each cross-sectional image layer obtained in step S231.

[0052] First, for each cross-sectional image layer, the discrete boundary candidate response points of that layer are mapped to a two-dimensional blank matrix with the same size as the cross-sectional image layer. The rows and columns of the matrix correspond to the pixel rows and pixel columns of the cross-sectional image layer, respectively. The elements in the matrix corresponding to the positions of the discrete boundary candidate response points are assigned a value of 1, and the remaining elements are assigned a value of 0, thus obtaining a binary boundary candidate map. For example, if the resolution of the cross-sectional image layer is 512×512, then a blank matrix with 512 rows and 512 columns is constructed. If the discrete boundary candidate response point is located at the 100th row and 200th column, then the element at the 100th row and 200th column of the matrix is ​​assigned a value of 1.

[0053] Subsequently, a circular structuring element of a preset size is used to perform dilation morphological processing on the boundary candidate binary image, filling the gaps between discrete boundary candidate response points and enhancing the topological continuity of the skin contour. Dilation morphological processing is an image morphological operation that uses a preset structuring element to slide on the binary image, assigning the maximum value within the structuring element's coverage area to the pixel corresponding to the structuring element's center. This expands the white foreground area (pixels with a value of 1) outward, filling the gaps between adjacent foreground pixels. The radius of the circular structuring element can be set to 1-3 pixels. For example, a circular structuring element with a radius of 1, covering the center pixel and its 8 neighboring pixels, will connect two pixels with a value of 1 that were originally separated by 1 pixel after dilation of the boundary candidate binary image.

[0054] Next, the connected component screening is performed on the dilated binary image, retaining connected components with a pixel count greater than a preset threshold (usually 10 pixels). This preset threshold is dynamically determined based on the resolution of the cross-sectional image layer; for example, the threshold is set to 10 pixels on the image after downsampling in step S52, and proportionally enlarged (e.g., 40 pixels) on the original high-resolution image. This removes isolated noise points, resulting in a topologically enhanced boundary binary image. Finally, pixels assigned a value of 1 are extracted from the topologically enhanced boundary binary image to obtain enhanced discrete boundary candidate response points.

[0055] Preferably, in step S233, the processing objects are the spatial topological correlation between adjacent cross-sectional image layers obtained in step S22, and the enhanced discrete boundary candidate response points obtained in step S232.

[0056] First, for each cross-sectional image layer, the topologically enhanced binary boundary map corresponding to the enhanced discrete boundary candidate response points of that layer is used as the processing object. Then, the eight-neighborhood connectivity rule is used to detect and label all connected components in the topologically enhanced binary boundary map, obtaining the number of pixels and the location range of each connected component.

[0057] The definition of the eight-neighbor connectivity rule: For any pixel P in a two-dimensional image, its eight neighbors are the eight pixels directly adjacent to P, including the four pixels directly above, below, to the left, and to the right of P, and the four diagonally adjacent pixels to the upper left, upper right, lower left, and lower right. The eight-neighbor connectivity rule means that when two pixels are eight neighbors of each other and both pixels have a value of 1, the two pixels are considered to belong to the same connected component.

[0058] The specific method for connected component detection and labeling (two-pass scanning method): First pass: Traverse the topology-enhanced boundary binary graph row by row and column by column. When encountering a pixel with a value of 1, check its four scanned neighboring pixels to the left, upper left, directly above, and upper right. If all neighboring pixels have a value of 0, assign a new connected component number to the current pixel. If one or more neighboring pixels have a value of 1, assign the smallest connected component number among these neighboring pixels to the current pixel and record the equivalence relationships between different numbers. Second pass: Traverse the binary graph row by row and column by column again. Based on the equivalence relationships recorded in the first pass, unify the connected component numbers belonging to the same equivalence group to the smallest number, completing the labeling of all connected components.

[0059] For example, consider a 3×3 topology-enhanced binary boundary graph with the following numerical values: |1|1|0| |1|0|0| |0|0|1| First scan: The pixel value in row 1, column 1 is 1, with no neighboring pixels of 1, assigned number 1; the pixel value in row 1, column 2 is 1, with left neighbor number 1, assigned value 1; the pixel value in row 2, column 1 is 1, with upper neighbor number 1, assigned value 1; the pixel value in row 3, column 3 is 1, with no neighboring pixels of 1, assigned number 2. Second scan: Two connected components are obtained: Component 1 has 3 pixels, located in rows 1-2 and columns 1-2; Component 2 has 1 pixel, located in row 3 and column 3. Then, based on the spatial topological correlation between adjacent cross-sectional image layers, spatial consistency filtering is performed on all marked connected components. If the Euclidean distance between the center coordinates of a connected component in the current layer and the center coordinates of the corresponding connected component in the adjacent layer is greater than 15 pixels, it is determined to be an invalid connected component. Invalid connected components with no spatial correspondence to the contours of adjacent layers are removed. The connected component with the most pixels after filtering is extracted as the main connected component of the skin contour of that cross-sectional image layer.

[0060] Finally, the internal region of the main connected region of the skin contour is filled, and all pixels within the closed region enclosed by the main connected region are assigned a value of 1, while the rest are assigned a value of 0, thus obtaining the dynamic candidate masking region feature map corresponding to the cross-sectional image layer. The rows and columns of the dynamic candidate masking region feature map correspond one-to-one with the pixel rows and columns of the corresponding cross-sectional image layer. When the value of the element at the intersection of the row and column in the feature map is 1, it means that the position belongs to the candidate masking region of the facial skin represented by the dynamic candidate masking region feature map. When the value is 0, it means that the position belongs to the non-candidate region.

[0061] Preferably, in step S24, the processing object is the dynamic candidate masking region feature map corresponding to all cross-sectional image layers obtained in step S233.

[0062] First, based on the spatial topological correlation between adjacent cross-sectional image layers obtained in step S22, an adjacent layer sliding window is set for each cross-sectional image layer. The range of the sliding window includes the current cross-sectional image layer and the adjacent cross-sectional image layers 1-2 layers before and after the current layer; for example, if the current layer is the 10th cross-sectional image layer, the sliding window range is layers 8-12.

[0063] Subsequently, for each cross-sectional image layer within the sliding window, the coordinates of the skin boundary points corresponding to each radial sampling contour line in the feature map of the dynamic candidate masking region of that layer are extracted to obtain the boundary coordinate sequence of that layer.

[0064] Next, the Euclidean distance between the boundary coordinates of the radial sampling contour lines with the same sequence number in adjacent cross-sectional image layers within the sliding window is calculated. The formula for calculating the Euclidean distance is: for two-dimensional coordinate points (x1, y1) and (x2, y2), the Euclidean distance = √[(x1-x2)²+(y1-y2)²]. The smaller the Euclidean distance, the higher the inter-layer spatial topological correlation at that location. For example, the boundary point coordinates of the first contour line in the 10th layer are (100, 200), and the boundary point coordinates of the first contour line in the 11th layer are (101, 200). The Euclidean distance is 1, indicating a very high inter-layer correlation.

[0065] Next, based on the Euclidean distance of the boundary coordinates of adjacent layers, the attention weight value corresponding to the position of each cross-sectional image layer and each radial sampling contour line is calculated. The weight calculation formula is: attention weight value = 1 / (1 + Euclidean distance). The smaller the Euclidean distance, the higher the attention weight value, and vice versa. For example, when the Euclidean distance is 1, the attention weight value = 1 / (1+1) = 0.5; when the Euclidean distance is 0, the attention weight value = 1.

[0066] Finally, the attention weight values ​​corresponding to all cross-sectional image layers and all radial sampling contours are integrated to construct a spatial dimension attention weight distribution matrix. The rows of the spatial dimension attention weight distribution matrix correspond to the indices of all cross-sectional image layers arranged along the layer tangent axis, and the columns of the matrix correspond to the indices of all radial sampling contours within the same cross-sectional image layer. The element values ​​at the intersection of rows and columns in the matrix represent the weight ratio of the skin boundary at the corresponding cross-sectional image layer and the corresponding radial sampling contour position in the inter-layer consistency constraint. For example, the element value of the 10th row and 1st column of the matrix is ​​0.5, which means that the weight of the 1st contour position of the 10th layer is 0.5.

[0067] Preferably, in step S25, the objects being processed are the dynamic candidate masking region feature map obtained in step S233 and the spatial dimension attention weight distribution matrix obtained in step S11.

[0068] First, for each radial sampling contour line of each cross-sectional image layer, the boundary coordinates of the dynamic candidate masking region feature map corresponding to that position are extracted as the initial boundary coordinates.

[0069] Subsequently, based on the attention weight value corresponding to the position in the spatial dimension attention weight distribution matrix, the boundary coordinates of the radial sampling contours of adjacent cross-sectional image layers with the same sequence number are calculated by weighted averaging to obtain the constrained boundary coordinates of that position (i.e., achieving inter-layer geometric consistency constraint processing). The weighted average calculation formula is: constrained boundary coordinates = current layer initial boundary coordinates × current weight value + adjacent layer boundary coordinates × adjacent layer weight values, and the sum of all weight values ​​involved in the calculation is 1; for example: the current layer weight value is 0.6, the initial boundary coordinates are (100, 200), the adjacent layer weight value is 0.4, and the boundary coordinates are (110, 200), then the constrained boundary coordinates = 100 × 0.6 + 110 × 0.4 = 104, the y-coordinate remains unchanged, and the final value is (104, 200). Of course, in some embodiments, the weighted average calculation formula can also be: constrained boundary coordinates = Σ(boundary coordinates of layer i × normalized weight_i), where normalized weight_i = attention weight_i / Σ(attention weights of all layers involved in the calculation).

[0070] After that, as Figure 3 As shown, the bounded boundary coordinates corresponding to all radial sampling contour lines within the same cross-sectional image layer are sequentially connected to form the bounded skin closed contour of that layer. The internal region of the bounded skin closed contour is then filled (for example, using a scan line filling algorithm to scan the internal region of the closed contour line by line and assign a value of 1 to all pixels within the closed interval enclosed by the contour lines) to obtain the target skin mask sub-region of that layer (i.e., Figure 3 (The masking area of ​​this layer is shown). Finally, the target skin masking sub-regions corresponding to all cross-sectional image layers are stacked and integrated along the layer tangent axis to obtain the three-dimensional target skin masking area.

[0071] Optionally, the step of performing isosurface evolution processing on the original pixel values ​​after suppression to generate an isosurface mesh, and performing patch validity screening processing on the isosurface mesh to generate skin extraction results, includes: Step S32: Perform isosurface evolution processing on the original pixel values ​​after suppression processing to obtain a continuous topological manifold mesh of skin tissue, and use the continuous topological manifold mesh of skin tissue as the isosurface mesh; Step S33: Based on the constructed vertex grayscale validity judgment criteria, perform patch validity screening on the continuous topological manifold mesh of the skin tissue to generate skin extraction results.

[0072] Step S3 also includes step S31, extracting the original pixel values ​​within the target skin masking area and suppressing their background values.

[0073] Preferably, in step S31, the processing objects are the preprocessed oral CBCT image to be processed obtained in step S21 and the target skin masking region obtained in step S25, and the original pixel value is the original HU value of the voxel corresponding to the target skin masking region in the oral CBCT image to be processed.

[0074] First, based on the one-to-one mapping relationship between downsampled pixels and original high-resolution pixels established by the volume data space segmentation operator in step S52, the three-dimensional target skin masking region is mapped back to the high-resolution space of the oral CBCT image to be processed, resulting in a high-resolution target skin masking region with the same size as the oral CBCT image to be processed. The one-to-one mapping relationship here records the coordinate range of 4 pixels in the original high-resolution image corresponding to each downsampled pixel, which can restore the low-resolution masking region to the high-resolution space without loss.

[0075] Subsequently, each voxel in the original oral CBCT image to be processed is traversed to determine whether it is located within the high-resolution target skin masking region. If the voxel is within the high-resolution target skin masking region, its original HU value is retained; if the voxel is outside the high-resolution target skin masking region, its HU value is set to a preset background value lower than the air tissue HU value (typically -1500 HU). Finally, after processing all voxels, the background-suppressed oral CBCT volume data is obtained.

[0076] Preferably, in step S32, the object of processing is the oral CBCT volume data to be processed after background value suppression obtained in step S31. The core processing action is non-cross-domain isosurface evolution processing. This processing is a scene-based optimization of the traditional Marching Cubes (MC) algorithm, a classic algorithm for extracting isosurfaces from three-dimensional volume data. The term "non-cross-domain" means that the isosurface evolution calculation is strictly limited to the target skin masking area and its boundaries obtained in step S25, without calculating the entire volume data space. This achieves scene-based optimization of the traditional Marching Cubes algorithm and significantly reduces the amount of computation.

[0077] First, refer to Figure 4 Using the skin cutting threshold obtained in step S13 as the isosurface threshold, the oral CBCT body data to be processed after background value suppression is divided into cubic units. Each cubic unit consists of 8 adjacent voxels and is the smallest processing unit of the MC algorithm.

[0078] Subsequently, all cube cells are traversed to determine whether the cube cell is completely located within the target skin masking area or intersects with the target skin masking area. This step is the core optimization of the traditional MC algorithm. The traditional MC algorithm traverses all cube cells of the entire volume data, while this application only processes the cells related to the target skin masking area, which greatly reduces the amount of computation.

[0079] Subsequently, if the cube element is completely outside the target skin masking area, it is skipped without isosurface calculation. If the cube element is within the target skin masking area or intersects with it, the Monte Carlo (MC) algorithm is used to calculate the isosurface for that cube element, generating corresponding triangular mesh patches. The core logic of the MC algorithm is as follows: based on the relationship between the HU values ​​of the eight vertices of the cube element and the isosurface threshold, the topological structure of the isosurface within the cube element is determined. Then, the coordinates of the isosurface vertices are calculated through linear interpolation, ultimately generating triangular mesh patches. This process achieves isosurface evolution. Finally, all generated triangular mesh patches are integrated to obtain a continuous topological manifold mesh of the skin tissue, which is then used as the isosurface mesh.

[0080] Preferably, in step S33, the object being processed is the continuous topological manifold mesh of skin tissue obtained in step S32, and the processing is based on a pre-constructed vertex grayscale validity determination criterion.

[0081] First, the specific content of the vertex grayscale validity determination criteria is clarified: For any triangular facet in the triangular mesh, if at least one of the two vertices of any edge of the triangular facet has an original voxel HU value lower than the preset background value (-1500HU) set in step S31, then the triangular facet is determined to be an invalid triangular facet; if the original voxel HU values ​​of the two vertices of all edges of the facet are higher than the preset background value, then the triangular facet is determined to be a valid triangular facet.

[0082] Subsequently, the three vertices of each triangular facet in the continuous topological manifold mesh of the skin tissue, along with the original voxel HU value corresponding to each vertex, are extracted. Then, based on the vertex grayscale validity criterion, all triangular faces are judged for validity, and all invalid triangular faces are marked. For example, if the HU values ​​of the two vertices of one edge of a triangular facet are -2000 and 0 respectively, where -2000 is lower than the preset background value of -1500, then this triangular facet is marked as invalid. Finally, all marked invalid triangular faces are deleted, and all valid triangular faces are retained. Topological consistency optimization is performed on the retained valid triangular faces. If an edge is shared by only one triangular facet, it is judged as a non-manifold edge and deleted; if all adjacent edges of a triangular facet have no other shared faces, it is judged as an isolated triangular facet and deleted. This process eliminates non-manifold edges and isolated triangular faces in the mesh. Through this facet validity screening process, the final 3D skin mesh model is obtained, which is the skin extraction result.

[0083] Optionally, step S4, the tissue-air interface grayscale anchoring mapping model includes an encoding component, a skin interface topology anchoring component, and a threshold regression component; Step S41: The encoding component is used to perform convolution processing on the prior anatomical features of skin tissue to obtain convolutional encoded features, and to perform global pooling processing on the convolutional encoded features to obtain skin interface encoded features. Step S42: The skin interface topology anchoring component is used to extract the radial sampling contour line of the cross-sectional image layer corresponding to the prior anatomical features of skin tissue and then determine its pixel sequence. Based on the preset asymmetric gradient template, the pixel sequence is subjected to point-by-point gradient response extraction processing to obtain a gradient response sequence. Peak localization processing is performed on the gradient response sequence to obtain the skin interface topology anchoring features. Step S43: The threshold regression component is used to perform splicing and fusion processing on the skin interface encoding features and skin interface topological anchoring features to obtain fused features, and to perform nonlinear regression fitting processing on the fused features based on the prior regular constraint of the anterior skin anatomical distribution of the oral face to obtain the skin cutting threshold.

[0084] Preferably, the specific implementation process of step S4 is as follows: This step is used to define the overall structure of the tissue-air interface grayscale anchoring mapping model. The tissue-air interface grayscale anchoring mapping model is a dual-branch anchoring nonlinear regression model specifically used for the fusion registration of oral CBCT and facial 3D surface scanning. The model includes a parallel encoding component and a skin interface topology anchoring component, and also includes a threshold regression component that is connected to the outputs of the two components. The inputs of both the encoding component and the skin interface topology anchoring component receive prior anatomical features of skin tissue, and the outputs of both the encoding component and the skin interface topology anchoring component are connected to the input of the threshold regression component. The output of the threshold regression component outputs the skin cutting threshold.

[0085] Preferably, in step S41, the processing object is the prior skin tissue anatomy features obtained in step S123. First, the encoding component performs multi-scale one-dimensional convolution processing on the input prior skin tissue anatomy features, setting up three parallel one-dimensional convolutional layers with the input dimension consistent with the dimension of the prior skin tissue anatomy features. The kernel sizes are 3, 5, and 7, the stride is 1, and the padding method is equal-length padding, capturing global information in feature sequences of different lengths. Then, the outputs of the three convolutional layers are concatenated along the channel dimension to extract deep global features from the prior skin tissue anatomy features, obtaining the initial convolutional encoded features. Here, the multi-scale one-dimensional convolution uses three one-dimensional convolutional layers with kernel sizes of 3, 5, and 7 in parallel to capture global information in feature sequences of different lengths. The outputs of the three convolutional layers are then concatenated to obtain the initial convolutional encoded features. Subsequently, the encoding component performs batch normalization and ReLU nonlinear activation processing on the initial convolutional encoded features to alleviate the gradient vanishing problem during model training, obtaining optimized convolutional encoded features. Next, the encoding component performs global pooling (or global average pooling in this embodiment) on the convolutionally encoded features to eliminate dimensional differences in the features, resulting in fixed-dimensional skin interface encoded features. Global average pooling calculates the average of all elements in the one-dimensional feature sequence, converting the variable-length feature sequence into a single fixed-length value. For example, if the one-dimensional feature sequence is [1,2,3,4,5], the value after global average pooling is (1+2+3+4+5) / 5=3. Finally, the encoding component outputs the skin interface encoded features to the threshold regression component.

[0086] Preferably, in step S42, the processing object is the radial sampling contour line of the cross-sectional image layer corresponding to the skin tissue anatomy prior features obtained in step S123. First, the skin interface topology anchoring component extracts the radial sampling contour line of the cross-sectional image layer corresponding to the skin tissue anatomy prior features, and determines the pixel HU value sequence (i.e., pixel sequence) corresponding to each radial sampling contour line. Then, the skin interface topology anchoring component uses a preset [1,1,0,-1,-1] asymmetric gradient template to perform point-by-point gradient response extraction processing on the pixel HU value sequence to obtain the gradient response sequence corresponding to each radial sampling contour line; the specific calculation rules are completely consistent with the asymmetric fifth-order gradient template calculation rules in step S231, ensuring the consistency of feature extraction logic between the model training stage and the inference stage. After that, the skin interface topology anchoring component performs peak localization processing on the gradient response sequence, extracts the position and pixel features corresponding to the gradient response peak, and obtains the skin interface topology anchoring features. Finally, the skin interface topology anchoring component outputs the skin interface topology anchoring features to the threshold regression component.

[0087] Preferably, in the specific implementation of step S42, the processing benchmark is first determined, and the cross-sectional image layer corresponding to the prior anatomical features of skin tissue is extracted. The processing object is the prior anatomical features of skin tissue generated in step S123. First, the basis for the generation of the prior anatomical features of skin tissue is clarified: the prior anatomical features of skin tissue are extracted from historical oral images, uniformly extracted along the slicing axis from the top of the skull to the mandible, with 20-50 layers covering the entire facial region. Therefore, each set of prior anatomical features of skin tissue corresponds to a set of cross-sectional image layers of the same source and spatial location. This set of cross-sectional image layers is the benchmark object for all subsequent processing. Then, the set of cross-sectional image layers is sorted according to the spatial order of the slicing axis, resulting in an ordered set of cross-sectional image layers to be processed. Finally, for each cross-sectional image layer to be processed, the connected components of the head tissue are extracted by HU value threshold segmentation, and the centroid coordinates of the connected components of the head tissue are calculated. These centroid coordinates are used as the center point of the head tissue in the cross-sectional image layer for the generation of the subsequent radial sampling contour lines.

[0088] Then, radial sampling contour lines are extracted from the cross-sectional image layers. The processing objects are each cross-sectional image layer obtained above, and the head tissue center point of the corresponding cross-sectional image layer. First, taking the head tissue center point of the cross-sectional image layer as the origin, the front of the face in the cross-sectional image layer as the 0-degree reference direction, and the line connecting the two external auditory canals as the left and right boundaries, rays originating from the origin are generated at fixed angular intervals within the ±90-degree frontal facial skin distribution area. Each ray is a radial sampling contour line; the angular interval is fixed at 5 degrees, generating a total of 36 radial sampling contour lines to ensure full coverage of the frontal facial skin area, with no invalid sampling and no feature omissions. Subsequently, for each radial sampling contour line, its angle sequence number, start coordinates, and end coordinates in the cross-sectional image layer are recorded to establish a unique identifier for the contour line. Finally, the 36 radial sampling contour lines of the same cross-sectional image layer are sorted by angle sequence number to obtain the radial sampling contour line set of that cross-sectional image layer.

[0089] Oral CBCT skin extraction only serves the registration and fusion of the anterior and posterior facial scan data. There is no registration requirement for the posterior half of the skull. Therefore, sampling contour lines are generated only within ±90 degrees. This is different from the 360-degree full-circle sampling of general image segmentation, which reduces invalid calculations and fully focuses on the core target area.

[0090] Next, the corresponding pixel sequence is determined based on the radial sampling contour lines. The processing objects are each of the radial sampling contour lines obtained above and the corresponding cross-sectional image layer. First, for a single radial sampling contour line, all pixels that the ray passes through on the cross-sectional image layer are extracted with pixel-level precision. All pixels are linearly sorted according to the order from the center point of the head tissue (the starting point of the contour line) to the outer side of the face (the ending point of the contour line). Then, the HU value corresponding to each sorted pixel is read sequentially to form a one-dimensional numerical sequence. This sequence is the pixel sequence corresponding to the radial sampling contour line. The arrangement order of the elements in the pixel sequence is completely consistent with the spatial order of the pixels on the contour line, and the value of each element is the HU value of the corresponding pixel. Then, when the radial sampling contour line extends to the image edge of the cross-sectional image layer, pixel extraction is immediately stopped to ensure that the pixel sequence has no invalid background pixels. Finally, the above operation is performed on all radial sampling contour lines of the same cross-sectional image layer to obtain the pixel sequence set corresponding to that cross-sectional image layer.

[0091] For example: A cross-sectional image layer has a resolution of 256×256, and the center point of the head tissue is (128, 128). A radial sampling contour line in the 0-degree direction starts from (128, 128) and extends along the positive y-axis in front of the face. The pixel coordinates it passes through are (128, 128), (128, 129), (128, 130), (128, 131), (128, 132), (128, 133), (128, 134), and (128, 135). The HU values ​​of the corresponding pixels are 200, 180, 190, 210, 10, -80, -150, and -200, respectively. The one-dimensional sequence [200, 180, 190, 210, 10, -80, -150, -200] is the pixel sequence corresponding to the radial sampling contour line.

[0092] Next, point-by-point gradient response extraction is performed on the pixel sequence based on the preset asymmetric gradient template to obtain the gradient response sequence. The object is each pixel sequence obtained above, and the processing tool is the preset [1,1,0,-1,-1] asymmetric fifth-order gradient template. All processing rules in this step are completely consistent with step S231 to ensure the consistency of features between model training and inference.

[0093] The first step is to clarify the design principles and parameter definitions of the core tool: The preset asymmetric gradient template used in this step is the [1,1,0,-1,-1] asymmetric fifth-order gradient template, a customized template adapted to the unidirectional gray-level step characteristics of the skin-air interface in oral CBCT, which differs from general symmetric gradient templates such as Sobel and Prewitt. The template has 5 weight bits, with weights from left to right as [1,1,0,-1,-1], where: The 0-weighted bit in the 3rd position at the center: is used for precise anchor positioning of the skin-air interface to lock the pixel position of the interface; The first and second weighted bits on the left side of the center: correspond to the gray-scale accumulation area on the skin tissue side, used to capture high HU value features inside the skin tissue; The 4th and 5th weighted bits on the right side of the center correspond to the grayscale attenuation area on the air side, which is used to capture the low HU value characteristics of the air region. This template only produces a strong response to the unidirectional grayscale step of "high HU value of skin tissue → low HU value of air", and has no response to the reverse grayscale change, which can greatly reduce the invalid interference caused by metal artifacts in the oral cavity and skull tissue.

[0094] The second step is to clarify the calculation rules for point-by-point sliding convolution: Align the center 0 weight bit of the asymmetric fifth-order gradient template with each pixel in the pixel sequence, and perform the following calculation for each aligned position: Each weight bit of the template is multiplied by the HU value of the corresponding pixel in the pixel sequence; Add up all the results of the multiplications to obtain the gradient response value at that alignment position; Boundary completion rule: When the template slides to the beginning and end of the pixel sequence, and some weight positions of the template do not have corresponding pixel values, the HU values ​​of the endpoints of the beginning and end of the pixel sequence are used to fill in the missing pixel values, ensuring that a valid gradient response value can be calculated for each alignment position.

[0095] The third step is to generate a gradient response sequence: according to the original order of the pixel sequence, the gradient response values ​​calculated at each alignment position are arranged sequentially to form a set of one-dimensional numerical sequences, which is the gradient response sequence corresponding to the radial sampling contour line; the above calculation is performed on all pixel sequences of the same cross-sectional image layer to obtain the set of gradient response sequences corresponding to the cross-sectional image layer.

[0096] For example, taking the pixel sequence [200,180,190,210,10,-80,-150,-200] above as an example, the complete calculation process is as follows: The first pixel of the template 0 weight bit alignment sequence (HU=200): There are no corresponding pixels for the two weight bits on the left side of the template. The first pixel HU value of 200 is used to fill in the missing pixels. The gradient response value = 200×1+200×1+200×0+180×(-1)+190×(-1)=30. The second pixel of the template 0 weight bit alignment sequence (HU=180): There is no corresponding pixel for the left weight bit of the template. It is filled with the HU value of the first pixel 200. The gradient response value = 200×1+180×1+180×0+190×(-1)+210×(-1)=-20. The third pixel of the template 0 weight bit alignment sequence (HU=190): All weight bits of the template have corresponding pixels, and the gradient response value = 180×1+190×1+190×0+210×(-1)+10×(-1)=150; The 4th pixel of the template 0 weight bit alignment sequence (HU=210): All weight bits of the template have corresponding pixels, and the gradient response value = 190×1+210×1+210×0+10×(-1)+(-80)×(-1)=470; The 5th pixel of the template 0 weight bit alignment sequence (HU=10): All weight bits of the template have corresponding pixels, and the gradient response value = 210×1+10×1+10×0+(-80)×(-1)+(-150)×(-1)=450; The 6th pixel of the template 0 weight bit alignment sequence (HU=-80): All weight bits of the template have corresponding pixels, and the gradient response value = 10×1+(-80)×1+(-80)×0+(-150)×(-1)+(-200)×(-1)=280; The 7th pixel of the template 0 weight bit alignment sequence (HU=-150): There is no corresponding pixel for the 1 weight bit on the right side of the template. It is filled with the HU value of the last pixel -200. The gradient response value = (-80)×1+(-150)×1+(-150)×0+(-200)×(-1)+(-200)×(-1)=170; The template 0 weight bit alignment sequence of the 8th pixel (HU=-200): There are no corresponding pixels for the two weight bits on the right side of the template. The last pixel HU value of -200 is used to fill in the missing pixels. The gradient response value = (-150)×1+(-200)×1+(-200)×0+(-200)×(-1)+(-200)×(-1)=50; The gradient response sequence corresponding to this pixel sequence is finally obtained as [30,-20,150,470,450,280,170,50].

[0097] Then, peak localization and invalid peak removal are performed on the gradient response sequences. The processing objects are each gradient response sequence obtained above and the corresponding radial sampling contour line. First, for a single gradient response sequence, the maximum value in the sequence is extracted as the gradient response peak of the sequence; at the same time, the index position of the gradient response peak in the gradient response sequence is located, which corresponds to the anchor pixel position of the skin-air interface in the pixel sequence. Then, the HU value corresponding to the anchor pixel position is extracted as the interface anchor gray value of the radial sampling contour line; at the same time, the two-dimensional coordinates of the anchor pixel in the cross-sectional image layer are recorded as the interface anchor coordinates of the contour line. Afterwards, based on the anatomical prior of continuous smoothness of oral and facial skin, invalid peak results of all contour lines are removed: ① If the gradient response peak of a certain contour line is lower than the preset response threshold (usually 50), it is judged as a noise response and the result is removed; ② If the Euclidean distance between the interface anchor coordinates of a certain contour line and the interface anchor coordinates of the adjacent angular contour lines exceeds the preset distance threshold (usually 10 pixels), it is judged as an abnormal jump and the result is removed. Finally, the set of effective interface anchor grayscale values ​​and the set of effective interface anchor coordinates are obtained after filtering.

[0098] For example: The maximum value of the gradient response sequence [30,-20,150,470,450,280,170,50] is 470, which is the gradient response peak value. It corresponds to the 4th position in the sequence and the 4th pixel in the pixel sequence. The HU value is 210, and the two-dimensional coordinates are (128,131). Therefore, the interface anchor gray value of this contour line is 210, and the interface anchor coordinates are (128,131). The peak value of 470 is greater than the response threshold of 50, and the distance between the peak value and the anchor coordinates of the adjacent contour line is less than 10 pixels. Therefore, it is determined to be a valid result.

[0099] Finally, the skin interface topological anchoring features are integrated and generated, processing the aforementioned set of effective interface anchoring grayscale values ​​and effective interface anchoring coordinates. First, the effective interface anchoring grayscale values ​​of a single cross-sectional image layer are sorted according to the angle indices of the radial sampling contour lines, forming a grayscale feature subsequence. The effective interface anchoring coordinates are then subjected to min-max normalization and sorted according to the angle indices of the radial sampling contour lines, forming a coordinate feature subsequence. Subsequently, the grayscale feature subsequence and the coordinate feature subsequence are concatenated end-to-end to form a one-dimensional feature vector corresponding to that cross-sectional image layer. Next, for all cross-sectional image layers corresponding to the prior anatomical features of skin tissue, the above feature generation operation is performed, concatenating the one-dimensional feature vectors corresponding to all cross-sectional image layers in order of the layer tangent axis direction to obtain a fixed-dimensional one-dimensional vector. Finally, this one-dimensional vector is the final skin interface topological anchoring feature, which is output to the threshold regression component for subsequent nonlinear regression fitting processing.

[0100] Preferably, in step S43, the processing objects are the skin interface encoding features obtained in step S41 and the skin interface topological anchoring features obtained in step S42. First, the threshold regression component performs channel-dimensional concatenation and fusion processing on the skin interface encoding features and the skin interface topological anchoring features to obtain fused features; for example, the skin interface encoding features are 16-dimensional vectors, and the skin interface topological anchoring features are 16-dimensional vectors, which are concatenated to obtain 32-dimensional fused features. Subsequently, the threshold regression component constructs a dual-constraint loss function, which includes a threshold fitting main loss term and a prior regularization constraint term for the anterior anatomical distribution of the oral and facial skin; wherein the threshold fitting main loss term adopts the mean squared error loss function, which is used to calculate the error between the skin cutting threshold predicted by the model and the optimal skin segmentation threshold labeled by the gold standard; the prior regularization constraint term for the anterior anatomical distribution of the oral and facial skin is used to penalize the result that the predicted threshold deviates from the reasonable HU value range (-50 to 50) of the skin anatomy, the farther the predicted threshold deviates from the reasonable range, the larger the penalty value, and the higher the overall value of the loss function. Next, the threshold regression component, based on a dual-constraint loss function, performs nonlinear regression fitting on the fused features through two fully connected layers. The first fully connected layer has the input dimension of the fused features and an output dimension of 16, with the ReLU activation function. The second fully connected layer has an input dimension of 16 and an output dimension of 1, with a linear activation function to ensure that the output threshold is a continuous grayscale value. Finally, the threshold regression component outputs the fitted skin cutting threshold.

[0101] In one embodiment, the step of training the tissue-air interface grayscale anchoring mapping model specifically includes: constructing a training sample set: acquiring historical oral CBCT three-dimensional volume data of the same or similar origin as in step S11, and generating corresponding skin tissue anatomical prior features for each historical oral CBCT three-dimensional volume data according to the methods in steps S11, S121 to S123; simultaneously, labeling each historical oral CBCT three-dimensional volume data with a gold standard skin cutting threshold, wherein the gold standard skin cutting threshold is the optimal segmentation threshold determined on the sample by manual segmentation or semi-automatic segmentation methods; model training: using the skin tissue anatomical prior features as input and the gold standard skin cutting threshold as a supervision signal, inputting them into the tissue-air interface grayscale anchoring mapping model constructed in steps S4, S41-S42, and iteratively optimizing the model parameters through the backpropagation algorithm and the double-constraint loss function defined in step S43 until the model converges, thereby obtaining the trained tissue-air interface grayscale anchoring mapping model.

[0102] Optionally, step S5, the volume data space segmentation operator includes a coordinate system alignment component, a downsampling component, a facial front ROI pre-definition component, a sampling contour line inter-layer association component, and a cross-section layer-by-layer segmentation component connected in sequence. Step S51: The coordinate system alignment component is used to extract the inherent anatomical landmarks of the skull that match the distribution area of ​​the anterior half of the facial skin in the oral CBCT image to be processed; based on the inherent anatomical landmarks of the skull, the oral CBCT image to be processed is processed to calculate a three-dimensional rigid transformation matrix to obtain a registration rigid transformation matrix; based on the registration rigid transformation matrix, the oral CBCT image to be processed is processed to perform three-dimensional spatial transformation to obtain CBCT volume data; the CBCT volume data is then processed to perform coordinate system alignment to obtain facial orientation CBCT volume data. Step S52: The downsampling component is used to perform edge-preserving bilateral filtering on the face orientation CBCT volume data in the plane of the cross-sectional image to obtain filtered CBCT volume data; perform downsampling on the filtered CBCT volume data to obtain initial downsampled volume data; and perform coordinate mapping on the initial downsampled volume data to obtain downsampled CBCT volume data. Step S53: The facial anterior ROI pre-definition component is used to perform effective region definition processing on the downsampled CBCT volume data based on the distribution area of ​​facial skin on the anterior half of the head, to obtain facial ROI defined volume data. Step S54: The sampling contour line interlayer association component is used to extract the centroid of the head tissue connected region in the cross-sectional image layer of the facial ROI-defined volume data, and use the centroid of the head tissue connected region as the head tissue center point; perform radial sampling contour line generation processing on the facial ROI-defined volume data with the head tissue center point as the origin to obtain the radial sampling contour line in front of the face; perform inter-layer topological association processing on the radial sampling contour line in front of the face to obtain volume data with interlayer association labels; Step S55: The cross-sectional layer-by-layer segmentation component is used to perform seamless layer-by-layer segmentation processing on the volume data with inter-layer association markers along the layer-slicing axis to obtain single-layer surface voxel slices; the single-layer surface voxel slices are subjected to voxel-pixel conversion processing to obtain each cross-sectional image layer.

[0103] Preferably, step S5 is used to define the overall structure of the volume data space segmentation operator, which is a directional segmentation operator with interlayer topological pre-association for facial skin extraction in oral CBCT. This operator includes a coordinate system alignment component, a downsampling component, a facial anterior ROI pre-definition component, a sampling contour line interlayer association component, and a cross-sectional layer-by-layer segmentation component connected sequentially; wherein the output of the preceding component is connected to the input of the following component, the input of the coordinate system alignment component receives the oral CBCT image to be processed, and the output of the cross-sectional layer-by-layer segmentation component outputs all cross-sectional image layers.

[0104] Preferably, in one scenario, in step S51, the object being processed is the oral CBCT image to be processed. First, the coordinate system alignment component extracts the inherent anatomical landmarks of the skull that match the distribution area of ​​the anterior half of the facial skin in the oral CBCT image to be processed. These inherent anatomical landmarks include the nasal root point, the superior margins of both external auditory canals, and the chin apex. These landmarks are fixed in position within the skull and are not affected by changes in the morphology of facial soft tissues, resulting in a set of inherent anatomical landmarks of the skull. For example, the extracted nasal root point has three-dimensional coordinates of (0, 100, 50), the superior margin of the right external auditory canal has coordinates of (-50, 0, 50), the superior margin of the left external auditory canal has coordinates of (50, 0, 50), and the chin apex has coordinates of (0, -100, 0). Subsequently, the coordinate system alignment component performs a three-dimensional rigid transformation matrix calculation on the oral CBCT image to be processed based on the set of inherent anatomical landmarks of the skull, obtaining a registration rigid transformation matrix. The core technology is explained in detail here: Definition of 3D rigid transformation: 3D rigid transformation includes only rotation and translation operations, and does not change the shape and size of the object. The transformation formula is: P'=R P+T, where P is the three-dimensional coordinate of the marker point in the original coordinate system, R is a 3×3 rotation matrix, T is a 3×1 translation vector, and P' is the three-dimensional coordinate of the marker point in the standard coordinate system.

[0105] The specific process of matrix calculation is as follows: First, take the line connecting the upper margins of the bilateral external auditory canals as the X-axis, with the direction from the right external auditory canal to the left external auditory canal; Second, take the direction from the root of the nose to the front of the face as the Y-axis; Third, take the axis passing through the midpoint of the line connecting the upper margins of the bilateral external auditory canals and perpendicular to the plane of the skull base as the Z-axis, and construct a standard facial orientation anatomical coordinate system; Fourth, based on the coordinates of the marker points in the original coordinate system and the coordinates of the marker points in the standard coordinate system, calculate the rotation matrix R and the translation vector T through singular value decomposition. The two together form the registration rigid transformation matrix (or calculate the centroid of the original set of marker points and the centroid of the set of marker points in the standard coordinate system, align the two centroids to obtain the translation vector T; construct a covariance matrix for the centered original marker point coordinates and the standard marker point coordinates, and perform singular value decomposition on the covariance matrix to obtain the rotation matrix R. The two together form the registration rigid transformation matrix). Subsequently, the coordinate system alignment component performs a three-dimensional spatial transformation on the oral CBCT image to be processed, based on the registration rigid transformation matrix, transforming the volume data in the original coordinate system to the standard coordinate system, thus obtaining the spatially registered CBCT volume data. Finally, the coordinate system alignment component performs coordinate system alignment processing on the spatially registered CBCT volume data, perfectly aligning the coordinate axes of the volume data with the coordinate axes of the standard facial orientation anatomical coordinate system, thus obtaining the facial orientation CBCT volume data.

[0106] Preferably, in step S52, the object being processed is the face-oriented CBCT volume data obtained in step S51. First, the downsampling component performs edge-preserving bilateral filtering on the face-oriented CBCT volume data within the plane of the cross-sectional image layer, suppressing image noise while preserving the edge features of the skin-air interface, resulting in filtered CBCT volume data. Here, the core technology is explained in detail: Bilateral filtering definition: Bilateral filtering is a nonlinear filtering method that considers both pixel spatial distance and pixel value similarity. Unlike ordinary Gaussian filtering, which only considers spatial distance, it can smooth noise while preserving the edge features of the image.

[0107] The core formula for filtering is: Filtered pixel value = (Pixel value of each pixel in the neighborhood × Spatial weight × Range weight) / (Sum of spatial weight × Range weight); where spatial weight measures the spatial distance between two pixels, the closer the distance, the greater the weight; and range weight measures the similarity of the HU values ​​of two pixels, the closer the HU values, the greater the weight.

[0108] For example, a 3×3 neighborhood bilateral filter is applied to a pixel. The HU value of the center pixel in the neighborhood is 0, and the HU values ​​of adjacent pixels are 5, -5, 100, and -100, respectively. Pixels with closer spatial distances have greater spatial weights, and pixels with HU values ​​closer to the center pixel have greater domain weights. The final filtered pixel values ​​will be closer to 0, 5, and -5, while being minimally affected by 100 and -100, thus smoothing noise while preserving the edges of the skin-air interface. Subsequently, the downsampling component performs a 50% downsampling process on the cross-sectional image layer of the filtered CBCT volume data, reducing the resolution of the cross-sectional image to half of its original value, resulting in the initial downsampled volume data.

[0109] Next, the downsampling component performs coordinate mapping processing on the initial downsampled volume data. Here, we first clarify the images and core definitions of two types of pixels: Original high-resolution pixels are the smallest imaging units of the native resolution cross-sectional image layer of the preprocessed oral CBCT image to be processed in step S21. Each original high-resolution pixel corresponds one-to-one with a three-dimensional voxel in the three-dimensional volume data of the oral CBCT to be processed, and the pixel's grayscale value is equal to the HU value of the corresponding voxel. Downsampled pixels are the smallest imaging units of the low-resolution cross-sectional image layer generated after performing a 50% downsampling process on the original high-resolution cross-sectional image layer in this step, belonging to the downsampled CBCT volume data output in this step. A one-to-one mapping relationship between downsampled pixels and original high-resolution pixels is established, and the coordinate range of the original high-resolution pixel corresponding to each downsampled pixel is recorded to obtain the downsampled CBCT volume data. For example, the downsampled pixel (0,0) corresponds to the four pixels (0,0), (0,1), (1,0), and (1,1) in the original high-resolution image. Finally, the downsampling component outputs the downsampled CBCT volume data to the predefined ROI component on the front of the face.

[0110] In this application, the downsampling operation is performed only within the cross-sectional plane. The original layer thickness and layer number are preserved throughout the entire slice axis direction without any downsampling. Therefore, the original high-resolution cross-sectional image layer and the downsampled cross-sectional image layer correspond one-to-one according to the layer number (the Nth original high-resolution cross-sectional image layer corresponds to the Nth downsampled cross-sectional image layer). The mapping rule between the two is as follows: For each set of cross-sectional image layers with corresponding layer numbers, in the original high-resolution cross-sectional image layer, a 2×2 adjacent block of original high-resolution pixels uniquely corresponds to a downsampled pixel in the downsampled cross-sectional image layer. The coordinate range of the original high-resolution pixel corresponding to each downsampled pixel is the row and column coordinate range of the above 2×2 original high-resolution pixel block in the original high-resolution cross-sectional image layer. For example, the original high-resolution cross-sectional image layer has a resolution of 512×512. The 2×2 original high-resolution pixel blocks with row number 0-1 and column number 0-1 correspond to one downsampled pixel with row number 0 and column number 0 in the downsampled 256×256 resolution cross-sectional image layer. The coordinate range of the original high-resolution pixel corresponding to this downsampled pixel is the interval between row 0-1 and column 0-1 of the original image.

[0111] Preferably, in step S53, the processed object is the downsampled CBCT volumetric data obtained in step S52. First, the anterior facial ROI pre-definition component determines the effective facial region range of each cross-sectional image layer in the downsampled CBCT volumetric data based on the anatomical prior of the anterior half of the head's facial skin distribution area. The anatomical prior is: the anterior half of the head is the facial skin distribution area, and the posterior half is an invalid area without skin extraction requirements, with the line connecting the upper edges of the bilateral external auditory canals as the boundary. Subsequently, the anterior facial ROI pre-definition component performs effective region definition processing on the downsampled CBCT volumetric data based on the effective facial region range, retaining the facial region data of the anterior half of the head. Afterward, the anterior facial ROI pre-definition component performs invalid region removal processing on the volumetric data after the effective region definition, deleting invalid tissue region data such as the posterior skull and cervical spine without skin extraction requirements. Finally, the facial ROI-defined volumetric data is obtained, and the anterior facial ROI pre-definition component outputs the facial ROI-defined volumetric data to the sampling contour line inter-layer association component.

[0112] Preferably, in step S54, the processing object is the facial ROI bounding volume data obtained in step S53. First, the sampling contour line inter-layer association component extracts the head tissue connected domain of each cross-sectional image layer in the facial ROI bounding volume data, calculates the centroid coordinates of the head tissue connected domain, and uses the centroid of the head tissue connected domain as the center point of the head tissue; the calculation formula for the centroid coordinates is: centroid x-coordinate = average of the x-coordinates of all pixels in the connected domain, centroid y-coordinate = average of the y-coordinates of all pixels in the connected domain. Subsequently, the sampling contour line inter-layer association component uses the center point of the head tissue as the origin and performs radial sampling contour line generation processing on each cross-sectional image layer of the facial ROI bounding volume data. For example, within a range of ±90 degrees in front of the face, a ray is generated every 5 degrees, generating a total of 36 radial sampling contour lines facing the front of the face, thus obtaining the radial sampling contour lines in front of the face; the generation rules are completely consistent with the radial sampling contour line generation rules in step S8, ensuring that the contour line definitions in the segmentation stage and the boundary extraction stage are unified. Subsequently, the sampling contour line interlayer association component performs interlayer topological association processing on the radial sampling contour lines in front of the face, establishing a one-to-one interlayer topological mapping relationship for radial sampling contour lines at the same angle in adjacent cross-sectional image layers, ensuring that sampling contour lines at the same angle are within the same anatomical longitudinal plane along the layer cutting axis. Finally, volume data with interlayer association labels is obtained, and the sampling contour line interlayer association component outputs the volume data with interlayer association labels to the cross-sectional layer-by-layer segmentation component.

[0113] Preferably, in step S55, the object of processing is the volume data with interlayer correlation markers obtained in step S54. First, the cross-sectional layer-by-layer segmentation component determines the slice axis (Z-axis) of the standard facial orientation anatomical coordinate system, using the slice thickness of the original oral CBCT image to be processed as the segmentation interval, ensuring that the sliced ​​cross-sectional image layers completely match the slice thickness of the original CBCT data. Subsequently, the cross-sectional layer-by-layer segmentation component performs non-overlapping and non-interval layer-by-layer segmentation processing on the volume data with interlayer correlation markers along the slice axis direction, segmenting the three-dimensional volume data into multiple independent two-dimensional tomographic slices, obtaining single-layer facial voxel slices; each single-layer facial voxel slice corresponds to a slice of the original CBCT volume data, containing the HU values ​​and interlayer correlation markers of all voxels within that slice. Subsequently, the cross-sectional layer-by-layer segmentation component, based on the inherent correspondence between the original 3D voxels and the original high-resolution 2D pixels (original high-resolution pixels) defined in step S52, performs voxel-to-pixel conversion processing on the single-layer partial voxel slices, mapping the HU value of the 3D voxel to the pixel grayscale value of the 2D image, thus obtaining the 2D image data corresponding to each cross-sectional image layer. The inherent correspondence here is: each voxel in the CBCT 3D volume data corresponds to a pixel in the 2D cross-sectional image, and the HU value of the voxel is the grayscale value of the pixel. Finally, all cross-sectional image layers are obtained, and the cross-sectional layer-by-layer segmentation component outputs all cross-sectional image layers for subsequent boundary feature extraction processing.

[0114] Furthermore, in one embodiment of this application, a skin extraction system is also provided, comprising: The threshold generation module is used to obtain the skin cutting threshold by performing nonlinear regression fitting on the prior anatomical features of skin tissue in historical oral images. The masking region generation module is used to acquire the oral CBCT image to be processed, extract the boundary features of the cross-sectional image layer of the oral CBCT image to be processed based on the skin cutting threshold to obtain a dynamic candidate masking region feature map, and perform inter-layer geometric consistency constraints on the dynamic candidate masking region feature map to generate the target skin masking region. The skin result generation module is used to extract the original pixel values ​​within the target skin masking area and suppress their background values. The suppressed original pixel values ​​are then subjected to isosurface evolution processing to generate an isosurface mesh. Finally, the isosurface mesh is subjected to patch validity screening to generate the skin extraction result.

[0115] The threshold generation module performs nonlinear regression fitting on the prior anatomical features of skin tissue in historical oral images to obtain the skin cutting threshold. The masking region generation module acquires the oral CBCT image to be processed, and extracts the boundary features of the cross-sectional image layer of the oral CBCT image to be processed based on the skin cutting threshold to obtain a dynamic candidate masking region feature map. The dynamic candidate masking region feature map is subjected to inter-layer geometric consistency constraints to generate the target skin masking region. The skin result generation module extracts the original pixel values ​​in the target skin masking region and suppresses their background values. The suppressed original pixel values ​​are subjected to isosurface evolution processing to generate isosurface grids. The isosurface grids are subjected to patch validity screening processing to generate the skin extraction result.

[0116] It should be understood that this application is not limited to the processes and structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.

Claims

1. A skin extraction method, characterized in that, include: Step S1: Perform nonlinear regression fitting on the prior anatomical features of skin tissue in historical oral images to obtain the skin cutting threshold; Step S2: Obtain the oral CBCT image to be processed. Based on the skin cutting threshold, extract the boundary features of the cross-sectional image layer of the oral CBCT image to be processed to obtain a dynamic candidate masking region feature map. Apply inter-layer geometric consistency constraints to the dynamic candidate masking region feature map to generate the target skin masking region. Step S3: Extract the original pixel values ​​within the target skin masking area and suppress their background values. Perform isosurface evolution processing on the suppressed original pixel values ​​to generate an isosurface mesh. Perform patch validity screening processing on the isosurface mesh to generate the skin extraction result.

2. The skin extraction method according to claim 1, characterized in that, Step S1 includes: The historical oral cavity image is acquired, and multi-scale gradient field mapping processing is performed on the historical oral cavity image to construct a gray-scale gradient evolution map of skin tissue. Based on the gray-level gradient evolution map of the skin tissue, multidimensional morphological feature points of the skin tissue are determined, and prior anatomical features of the skin tissue are generated accordingly. The generated prior anatomical features of the skin tissue are input into a preset tissue-air interface grayscale anchoring mapping model for nonlinear regression fitting to obtain the skin cutting threshold.

3. The skin extraction method according to claim 1, characterized in that, The process of extracting boundary features from the cross-sectional image layer of the oral CBCT image to be processed based on the skin cutting threshold to obtain a dynamic candidate masking region feature map includes: By constructing a volume data space segmentation operator, all cross-sectional image layers are extracted from the oral CBCT image to be processed; Based on the skin cutting threshold, boundary features are extracted for each cross-sectional image layer to determine the spatial topological correlation between adjacent cross-sectional image layers. Based on the spatial topological correlation, the dynamic candidate masking region feature map is generated.

4. The skin extraction method according to claim 1, characterized in that, The process of performing isosurface evolution processing on the original pixel values ​​after suppression to generate an isosurface mesh, and then performing patch validity screening on the isosurface mesh to generate skin extraction results includes: The original pixel values ​​after suppression are subjected to isosurface evolution processing to obtain a continuous topological manifold mesh of skin tissue, which is then used as the isosurface mesh. Based on the constructed vertex grayscale validity criterion, the continuous topological manifold mesh of the skin tissue is subjected to patch validity screening to generate skin extraction results.

5. The skin extraction method according to claim 2, characterized in that, The process of determining multidimensional morphological feature points of skin tissue based on the gray-level gradient evolution map of the skin tissue, and generating prior anatomical features of the skin tissue accordingly, includes: Based on the gray-scale gradient evolution map of the skin tissue, multidimensional morphological feature points of the skin tissue were determined; The multidimensional morphological feature points of the skin tissue are subjected to cross-scale anatomical topology nesting processing to obtain a cross-scale nested feature set. Based on the cross-scale nested feature set, the prior anatomical features of the skin tissue are generated.

6. The skin extraction method according to claim 3, characterized in that, The generation of dynamic candidate masking region feature maps based on spatial topological correlation includes: Asymmetric gradient transition kernel excitation response processing is performed on each cross-sectional image layer to extract the corresponding discrete boundary candidate response points. The discrete boundary candidate response points are subjected to dilatational morphological topology enhancement processing. Based on spatial topological correlation, the enhanced discrete boundary candidate response points are subjected to maximum connected component extraction processing to generate the dynamic candidate masking region feature map.

7. The skin extraction method according to claim 1, characterized in that, The step of applying inter-layer geometric consistency constraints to the feature map of the dynamic candidate masking region to generate the target skin masking region includes: Calculate the spatial dimension attention weight distribution matrix among the feature maps of the dynamic candidate masking regions corresponding to different cross-sectional image layers; Based on the spatial dimension attention weight distribution matrix, inter-layer geometric consistency constraint processing is performed on the feature maps of dynamic candidate masking regions corresponding to all cross-sectional image layers to obtain the target skin masking region.

8. The skin extraction method according to claim 2, characterized in that, The tissue-air interface grayscale anchoring mapping model includes an encoding component, a skin interface topology anchoring component, and a threshold regression component; The encoding component is used to perform convolution processing on the prior anatomical features of skin tissue to obtain convolutional encoded features, and to perform global pooling processing on the convolutional encoded features to obtain skin interface encoded features. The skin interface topology anchoring component is used to extract the radial sampling contour line of the cross-sectional image layer corresponding to the prior anatomical features of the skin tissue and then determine its pixel sequence. Based on the preset asymmetric gradient template, the pixel sequence is subjected to point-by-point gradient response extraction processing to obtain a gradient response sequence. Peak localization processing is performed on the gradient response sequence to obtain the skin interface topology anchoring features. The threshold regression component is used to perform splicing and fusion processing on the skin interface encoding features and the skin interface topological anchoring features to obtain fused features, and to perform nonlinear regression fitting processing on the fused features based on the prior regularization constraint of the anterior skin anatomical distribution of the oral face to obtain the skin cutting threshold.

9. The skin extraction method according to claim 3, characterized in that, The volume data space segmentation operator includes a coordinate system alignment component, a downsampling component, a facial front ROI pre-definition component, a sampling contour line inter-layer association component, and a cross-section layer-by-layer segmentation component connected in sequence. The coordinate system alignment component is used to extract the inherent anatomical landmarks of the skull that match the distribution area of ​​the anterior half of the facial skin in the oral CBCT image to be processed; and to perform three-dimensional rigid transformation matrix calculation on the oral CBCT image to be processed based on the inherent anatomical landmarks of the skull to obtain the registration rigid transformation matrix. Based on the registration rigid transformation matrix, the oral CBCT image to be processed is subjected to three-dimensional spatial transformation processing to obtain CBCT volume data; the CBCT volume data is then subjected to coordinate system alignment processing to obtain facial orientation CBCT volume data; The downsampling component is used to perform edge-preserving bilateral filtering on the face orientation CBCT body data in the plane of the cross-sectional image to obtain filtered CBCT body data; The filtered CBCT volume data is downsampled to obtain initial downsampled volume data; the initial downsampled volume data is then subjected to coordinate mapping to obtain downsampled CBCT volume data. The facial anterior ROI pre-definition component is used to perform effective region definition processing on the downsampled CBCT volume data based on the distribution area of ​​the facial skin on the anterior half of the head, to obtain facial ROI defined volume data. The sampling contour line interlayer association component is used to extract the centroid of the head tissue connected domain in the cross-sectional image layer of the facial ROI bounded volume data, and use the centroid of the head tissue connected domain as the head tissue center point. Radial sampling contour line generation processing is performed on the facial ROI bounded volume data with the center point of the head tissue as the origin to obtain the radial sampling contour line in front of the face. Perform inter-layer topological correlation processing on the radial sampling contour lines in front of the face to obtain volume data with inter-layer correlation markers; The cross-sectional layer-by-layer segmentation component is used to perform seamless layer-by-layer segmentation processing on the volume data with inter-layer association markers along the layer-slicing axis to obtain single-layer surface voxel slices; the single-layer surface voxel slices are then subjected to voxel-to-pixel conversion processing to obtain each cross-sectional image layer.

10. A skin extraction system, characterized in that, include: The threshold generation module is used to obtain the skin cutting threshold by performing nonlinear regression fitting on the prior anatomical features of skin tissue in historical oral images. The masking region generation module is used to acquire the oral CBCT image to be processed, extract the boundary features of the cross-sectional image layer of the oral CBCT image to be processed based on the skin cutting threshold to obtain a dynamic candidate masking region feature map, and perform inter-layer geometric consistency constraints on the dynamic candidate masking region feature map to generate the target skin masking region. The skin result generation module is used to extract the original pixel values ​​within the target skin masking area and suppress their background values. The suppressed original pixel values ​​are then subjected to isosurface evolution processing to generate an isosurface mesh. Finally, the isosurface mesh is subjected to patch validity screening to generate the skin extraction result.

Citation Information

Patent Citations

  • Medical image reconstruction method

    CN115082585A

  • Facial data processing method and device, electronic equipment and storage medium

    CN121259223A

  • Tooth three-dimensional precision correction method and system based on multi-source oral cavity data fusion

    CN121639531A

  • Method for facial image registration

    KR102097396B1