A mandible anatomical partition segmentation method, device, storage medium and program product

CN122780313APending Publication Date: 2026-09-18PEKING UNIV SCHOOL OF STOMATOLOGY +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610794565.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-03
Publication Date
2026-09-18

AI Technical Summary

Technical Problem

[0005]近年来,基于深度学习的医学图像分割方法取得了显著进展,然而相关技术大多仅能实现下颌骨整体的分割,无法精确地将下颌骨划分为具有临床意义的各解剖亚区

Benefits of technology

[0006] The purpose of this application is to provide a method, device, storage medium, and program product for segmenting the anatomical regions of the mandible, so as to achieve fine segmentation from raw medical images to various anatomical subregions of the mandible.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122780313A_ABST
    Figure CN122780313A_ABST
Patent Text Reader

Abstract

This application provides a method, apparatus, storage medium, and program product for mandibular anatomical segmentation. The method includes: acquiring a medical image containing a mandibular region; locating the mandibular region based on the medical image to obtain a location result; cropping a local mandibular image from the medical image based on the location result; obtaining a mandibular mask, mandibular tooth segmentation results, and coordinate information of preset anatomical landmarks based on the local mandibular image; and segmenting the mandibular mask according to preset anatomical segmentation rules based on the coordinate information of the preset anatomical landmarks and the mandibular tooth segmentation results to obtain segmentation results for each anatomical subregion of the mandible. This achieves end-to-end automatic segmentation of each anatomical subregion of the mandible, improving segmentation efficiency and consistency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical image processing technology, and more specifically, to a method, device, storage medium, and program product for segmenting the anatomical regions of the mandible. Background Technology

[0002] The mandible is the only movable bone in the maxillofacial region. It is horseshoe-shaped and consists of the mandibular body and mandibular ramus. The mandibular body is horizontal, and the mandibular ramus is vertical; the junction of the body and ramus is the angle of the mandible. The outer layer of the mandible is cortical bone, while the inner layer is cancellous bone. Cortical bone has high density and hardness, appearing as a bright area with high grayscale values ​​in CT images; cancellous bone has lower density and a looser structure, appearing as a dark area with low grayscale values ​​in CT images.

[0003] The mandible is a common bony structure in the maxillofacial region, and accurate classification and diagnosis of its lesions are a prerequisite for clinical diagnosis and treatment in oral and maxillofacial surgery. In diagnostic processes based on imaging techniques such as CBCT and CT, precise segmentation of the mandibular region is fundamental to achieving standardized classification and diagnosis of lesions. Precise segmentation clearly defines the boundaries between various anatomical subregions of the mandible and diseased tissues, providing accurate anatomical localization and quantitative evidence of the extent of invasion for fracture classification, tumor staging, and deformity classification. Simultaneously, it allows for the extraction of repeatable quantitative imaging features, providing objective indicators for the differential diagnosis of different types of lesions, effectively reducing discrepancies in image interpretation among physicians. It is also a crucial step in building AI-assisted diagnostic systems and improving diagnostic accuracy and standardization, possessing irreplaceable value for the precise diagnosis and treatment of mandibular diseases.

[0004] In the clinical diagnosis and treatment of mandibular fractures, doctors typically classify and assess the fracture site according to anatomical zoning standards. For example, the commonly used nine-zoning classification system for mandibular fractures divides the mandible into several anatomical subregions, including the condylar region, coronoid region, mandibular angle and ramus region, mandibular body region, and chin and paramental region. Weak points of the mandible include the condylar neck, median symphysis, mental foramen, and mandibular angle; these areas are more prone to fracture when subjected to external forces.

[0005] In recent years, deep learning-based medical image segmentation methods have made significant progress. However, most of these technologies can only segment the mandible as a whole and cannot accurately divide the mandible into clinically significant anatomical subregions. Summary of the Invention

[0006] The purpose of this application is to provide a method, device, storage medium, and program product for segmenting the anatomical regions of the mandible, so as to achieve fine segmentation from raw medical images to various anatomical subregions of the mandible.

[0007] The first aspect of this application provides a method for anatomical partitioning of the mandible, the method comprising: The acquired imaging content includes medical images of the mandibular region; The mandibular region is located based on the medical images, and the location results are obtained; A partial image of the mandible is obtained by cropping from the medical image based on the positioning results; Based on the local image of the mandible, the mandibular mask, the segmentation results of the lower teeth, and the coordinate information of the preset anatomical landmarks are obtained; Based on the coordinate information of the preset anatomical landmarks and the segmentation results of the lower teeth, the mandibular mask is divided into segments according to the preset anatomical zoning rules to obtain the segmentation results of each anatomical subregion of the mandible.

[0008] In the above process, end-to-end automatic segmentation of each anatomical subregion of the mandible was achieved, improving the efficiency and consistency of zoning and providing an anatomical basis for regional analysis and surgical planning of mandibular-related diseases.

[0009] Further, the localization of the mandibular region based on the medical images to obtain the localization result includes: The medical images are processed using a three-dimensional encoding and decoding segmentation network to output a probability map of the mandibular region; A binary mask is obtained by threshold segmentation of the probability map, and the bounding box of the binary mask is calculated. The bounding box is then determined as the positioning result.

[0010] In the above implementation process, a three-dimensional encoding and decoding segmentation network is used to perform preliminary localization of medical images. The spatial range of the mandible is automatically determined by using the probability map output by the network, after thresholding and bounding box calculation. This can effectively filter out interference from surrounding non-target structures such as the skull base and cervical spine, thus improving the robustness of localization.

[0011] Furthermore, the preset anatomical landmarks include the lowest point of the sigmoid notch and the angle of the mandible; Based on the coordinate information of the preset anatomical landmarks and the segmentation results of the lower teeth, the mandibular mask is divided into sections according to preset anatomical partitioning rules, including: Based on the coordinates of the lowest point of the sigmoid notch and the coordinates of the mandibular angle point, the boundaries of each anatomical subregion in the mandibular mask are determined.

[0012] In the above implementation process, the lowest point of the sigmoid notch and the angle of the mandible are used as the reference feature points for partitioning. By utilizing the characteristic that the two constitute the boundary reference between the mandibular ascending ramus and the mandibular body in anatomy, the division of the boundaries of each anatomical subregion has a clear skeletal morphological basis, and the partitioning results are more reasonable in terms of anatomical structure.

[0013] Furthermore, the preset anatomical zoning rules include a nine-zoning standard for mandibular fracture classification; The determination of the boundaries of each anatomical subregion within the mandibular mask based on the coordinates of the lowest point of the sigmoid notch and the coordinates of the angle of the mandible includes: Based on the coordinates of the lowest point of the sigmoid notch, the coordinates of the mandibular angle point, and the position information indicated by the tooth number label corresponding to the target tooth position in the lower tooth segmentation result, the boundaries of each anatomical subregion are determined to divide the mandibular mask into the condylar region, the coronoid region, the mandibular angle and ramus region, the mandibular body region, the chin and parachin region, and the transition zone of each anatomical subregion of the mandible.

[0014] In the above implementation process, by integrating the three-dimensional position information indicated by the lowest point of the sigmoid notch, the angle of the mandible, and the tooth number label of the target tooth position, the sub-regions are divided according to the clinically accepted nine-zone standard for mandibular fracture classification. The mandibular mask is divided into the condylar region, the coronoid region, the angle of the mandible and the ascending ramus region, the mandibular body region, the chin and parachin region, and the transition zone. The zoning results not only conform to the standardized anatomical definition, but also accurately define the transition boundaries between each sub-region by introducing tooth position spatial reference. The zoning results can be directly adapted to the clinical diagnostic needs such as fracture classification assessment.

[0015] Furthermore, the lower tooth segmentation result is obtained through the following steps: The local image of the mandible is input into the tooth segmentation sub-network to obtain the segmentation result of the lower teeth and the tooth number label corresponding to each tooth in the segmentation result of the lower teeth; the tooth segmentation sub-network adopts a three-dimensional encoding and decoding segmentation network structure; the tooth segmentation sub-network is trained using a hybrid loss function that combines the Dess loss function and the cross-entropy loss function.

[0016] In the above implementation process, a hybrid loss function combining Descein loss and cross-entropy loss is used to simultaneously achieve independent instance segmentation of each tooth and automatic identification of the corresponding tooth number label in the three-dimensional spatial context. This solves the problem of boundary adhesion between different teeth in the three-dimensional dental arch and the problem of missing segmentation caused by the imbalance between the foreground tooth body and the background bone tissue categories. The output segmentation results with tooth number labels can provide spatial position reference for each tooth for subsequent partitioning, making the partition boundary delineation more accurate.

[0017] Furthermore, the mandibular mask is obtained through the following steps: The local image of the mandible is input into a fine segmentation subnetwork to obtain the mandibular mask.

[0018] In the above implementation process, by introducing a fine segmentation subnetwork on the local image of the mandible, and performing secondary fine semantic segmentation within the initially located local area, the complete outline of the mandible can be separated more accurately from the surrounding soft tissues and bony structures, obtaining a high-quality mandibular mask, and providing a reliable spatial range benchmark and boundary constraint for the subsequent division of each anatomical subregion.

[0019] Furthermore, the coordinate information of the preset anatomical landmarks is obtained through the following steps: The local image of the mandible is input into the key point detection sub-network to obtain a three-dimensional Gaussian heatmap for each preset anatomical landmark. Based on the voxel coordinates in the three-dimensional Gaussian heat map whose probability values ​​meet the preset requirements, the coordinate information of the corresponding anatomical landmarks is determined.

[0020] In the above implementation process, a key point detection sub-network based on heatmap regression is adopted to output a three-dimensional Gaussian heatmap for each preset anatomical landmark, and its three-dimensional coordinates are automatically extracted through a probability peak localization strategy, thereby achieving high-precision synchronous localization of multiple anatomical landmarks.

[0021] A second aspect of this application provides a mandibular anatomical partitioning device, the device comprising: The acquisition module is used to acquire medical images containing images of the mandibular region. The positioning module is used to locate the mandibular region based on the medical images and obtain the positioning results; A cropping module is used to crop a local image of the mandible from the medical image based on the positioning result; The processing module is used to obtain the mandibular mask, the segmentation result of the lower teeth, and the coordinate information of preset anatomical landmarks based on the local image of the mandible. The partitioning module is used to partition the mandibular mask according to the preset anatomical landmark coordinate information and the mandibular tooth segmentation results, and obtain the segmentation results of each anatomical subregion of the mandible.

[0022] A third aspect of this application provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of any of the methods described in the first aspect.

[0023] A fourth aspect of this application provides a computer program product, the computer program product including a computer program, which, when executed by a processor, implements any of the methods described in the first aspect. Attached Figure Description

[0024] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 A schematic diagram of an overall process provided for an embodiment of this application; Figure 2 A flowchart illustrating a method for segmenting the anatomical regions of the mandible, provided in an embodiment of this application; Figure 3 This is a schematic diagram of a mandibular anatomical partitioning device provided in an embodiment of this application. Detailed Implementation

[0026] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.

[0027] It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures. Furthermore, in the description of this application, terms such as "first," "second," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0028] In related technologies, automated segmentation of the mandibular region in maxillofacial CT images is mainly achieved using convolutional neural network-based semantic segmentation algorithms. The core solutions fall into two categories: First, there are R-CNN-based region segmentation schemes. Early models used Region CNN for mandibular region segmentation, but these models require a large number of images with consistent format for training, making them unsuitable for the clinical needs of medical image segmentation. Subsequent improved algorithms such as Fast R-CNN, Faster R-CNN, and YOLO improve training speed by adding region-of-interest pooling layers and merging them with the neural network, but the segmentation accuracy in the midface and mandibular regions still cannot meet the clinical requirements for fracture diagnosis. Second, there are U-Net-based semantic segmentation schemes. These often use the U-Net model to achieve intelligent segmentation of maxillofacial bone tissue regions such as the frontal bone and mandible. This model integrates deep and shallow features through a jump-connected U-shaped network structure, which better matches the characteristics of medical image data. However, the related technologies directly generate maximum intensity projection (MIP) images based on the original CT data. Because the target area of ​​the mandible accounts for a small proportion in the whole-cranial CT, the generated MIP images have poor contrast, resulting in large errors in the segmentation results. At the same time, the related U-Net model is not good at extracting the boundary features of the subdivided anatomical regions such as the condyle, ascending ramus, and body of the mandible, and the segmentation accuracy of the midface bone tissue is poor, which cannot provide a stable and high-precision segmentation basis for subsequent mandibular fracture identification and classification diagnosis.

[0029] Currently, the internationally accepted AOCMF classification system provides a clinical consensus basis for the standardized description of mandibular fracture regions. This system divides the mandible into nine anatomical regions, including the chin and paramental region, the mandibular body region, the mandibular angle and ramus, the condyle, and the coronoid process. Regarding the delineation of regional boundaries, the system precisely defines the dividing lines between each anatomical region: the chin and paramental region is bounded by the roots of the two mandibular canines; the mandibular body region is the area between the mandibular canines and the third molars; the mandibular angle and ramus region is an irregular pentagonal area, with its anterior boundary being a vertical line posterior to the third molar, its lateral boundary extending to the lower posterior border of the mandibular angle, and its superior boundary including the mandibular notch and the base of the condyle and coronoid process. The condyle is defined by the line connecting the mandibular notch and the concave part of the coronoid process, while the coronoid process is defined by extending from the lowest point of the mandibular notch to the masseter tuberosity. Furthermore, the system defines several transition zones between the various regions to address situations where fractures are located precisely on or across region boundaries, ensuring the accuracy and consistency of fracture localization. While this zoning standard is valuable in clinical diagnosis and surgical planning, there is currently a lack of methods to automate its implementation. Clinically, doctors still primarily rely on manually defining region boundaries, which is time-consuming, labor-intensive, and has limited interpretability. Related automatic segmentation algorithms often treat the mandible as a single target for overall segmentation, failing to directly obtain sub-region results conforming to the nine-zone standard.

[0030] For any of the questions raised above, refer to Figure 1 This application proposes to first acquire oral medical images and then process them using a pre-defined automatic segmentation algorithm and key point detection algorithm to obtain lower tooth segmentation results, mandibular segmentation results, and coordinates of multiple anatomical key points. Specifically, the automatic segmentation algorithm adopts a multi-stage cascaded network architecture. The first step is a mandibular localization network, used to locate the entire mandibular region in the oral medical images through image segmentation. Based on the first step, the second step uses three parallel sub-networks to perform fine segmentation, tooth segmentation, and key point detection on the localized mandibular image, respectively. The fine segmentation sub-network outputs a complete mandibular mask, the tooth segmentation sub-network outputs the lower tooth segmentation results, and the key point detection sub-network outputs the coordinates of the lowest point of the sigmoid notch and the angle of the mandible. This solves the problems of disconnect between segmentation and detection, low processing efficiency, poor consistency of manual segmentation, strong subjectivity, and inability to standardize automatic partitioning in related technologies.

[0031] The following is a detailed description of the mandibular localization network in the first step. Regarding dataset construction, the complete cortical bone boundary and internal region of the mandible in each image are labeled layer by layer to generate a complete mandibular mask as the gold standard. In the data preprocessing stage, all original images are first resampled to unify their voxel spacing, eliminating parameter differences between different acquisition devices; subsequently, Z-score normalization is performed on the images. Regarding algorithm model construction, 3D U-Net is used as the basic network structure. In the model training stage, the preprocessed images and the corresponding gold standard are downsampled to a preset size. The coarse segmentation result of the mandible is used as the network output, and online data augmentation strategies are introduced, including random scaling, random rotation, random translation, and contrast enhancement. The training process uses the Dice loss function, selects the Adam optimizer, and combines a learning rate decay strategy until the loss function converges. In the inference phase: The oral medical image to be processed is input into the trained mandibular localization network, which outputs a probability map of the mandibular region. A binary mask for the mandible is obtained through threshold segmentation, and the minimum bounding box is calculated based on this mask. Finally, the original image is cropped according to the bounding box to obtain a localized image of the mandible. After this localization step, the subsequent processing range is narrowed down to the area where the mandible is located, effectively reducing the computational load and eliminating interference from irrelevant structures.

[0032] After obtaining the mandibular image, the second step employs three parallel sub-networks to perform fine segmentation, tooth segmentation, and keypoint detection on the localized mandibular image, respectively. Regarding dataset construction, the first step has already obtained the gold standard for the mandible. Closed annotations are performed on the mandibular teeth in each image, and corresponding tooth numbers are assigned based on their positions. In 3D keypoint annotation, the lowest point of the sigmoid notch and the mandibular angle are located on each axial section, and a circular mask with a radius of 2 pixels is drawn centered on these points. Through the stacking of masks on continuous sections, a 3D spherical mask is constructed, and its geometric centroid serves as the reference coordinate for the annotation points.

[0033] The fine segmentation of the mandible uses 3D U-Net as the basic network structure. During model training, fixed-size voxel blocks are randomly cropped from the localized mandibular image and used as training data input, optimized using the Dice loss function. During model inference, a block-first, then-merging prediction method is employed. The local mandibular image is divided into multiple overlapping fixed-size voxel blocks, which are sequentially input into the trained fine segmentation model. The segmentation results of all voxel blocks are then voted on and merged to form a complete fine segmentation result.

[0034] The tooth segmentation process is as follows: First, the localized mandibular image and tooth annotations are downsampled to a preset size. Then, the gold standard tooth annotations are converted into a 17-channel one-hot encoded form according to 16 categories (31-38, 41-48), with the background treated as a separate channel. A 3D U-Net is used as the segmentation network. During the training phase, a strategy combining Dice loss and cross-entropy loss is employed to address the tooth category imbalance problem, as shown in the formula below: ; Where N is the number of samples, specifically the number of voxels output by the segmentation network; C is the number of categories, and in this embodiment N is 17; y ic It is a sample i The actual label, using one-hot encoding; p ic It is a sample i The predicted probability of belonging to category c is obtained from the segmentation model output; ε is a smoothing factor to prevent the denominator from being 0; during inference, the localized mandibular image is downsampled to a preset size and input into the trained tooth segmentation network. After obtaining the probability map, the tooth category of each voxel is obtained through argmax. Finally, the segmentation result is upsampled to restore the original resolution to obtain the fine segmentation result and corresponding tooth number of each tooth.

[0035] Keypoint detection employs a heatmap regression method, using 3D U-Net as the base network structure. During model training, the reference coordinates of each keypoint are converted into a 3D Gaussian heatmap as a training supervision signal. Specifically, for each keypoint, a 3D heatmap tensor of the same size as the input image patch is constructed, centered on its reference coordinates. The value of each voxel position in the heatmap is defined as the Gaussian response intensity from that voxel to the center of the keypoint, calculated using an isotropic 3D Gaussian distribution function. ; in, The reference coordinates of the key points The standard deviation of the Gaussian kernel controls the diffusion range of the response region in the heatmap. This function ensures that the response is centered at key points. Place As the voxel moves further away from the center, the response value gradually decays and approaches 0 according to a Gaussian distribution, with values ​​in the edge regions approaching 0. The network output layer is designed with two channels, each corresponding to a 3D heatmap of a keypoint. The size of the heatmap is consistent with the spatial size of the input image patch. The last layer of the network uses the Sigmoid activation function to restrict the output value to the [0,1] interval, and uses the mean squared error loss function to regress and predict the probability of each voxel value containing a keypoint. During the inference stage, the coordinates of the voxel with the highest probability value are the preliminary predicted position of the corresponding keypoint.

[0036] This application first constructs a mandibular bone localization network using 3D U-Net to complete the overall coarse segmentation of the mandible and the clipping of the region of interest, thereby reducing the computational range and eliminating interference from irrelevant tissues. Then, through three parallel 3D U-Net sub-networks, it simultaneously achieves fine segmentation of the mandible, classification and segmentation of the 16 lower teeth, and 3D key point heatmap regression detection of the lowest point of the sigmoid notch and the angle of the mandible. It outputs all the basic data required for partitioning in one stop, solving the problems of disconnect between segmentation and detection and low processing efficiency in related technologies.

[0037] Meanwhile, based on the fitted lower dental arch curve, this application constructs a cutting plane related to the dental arch using the centroid normal vector of the target tooth position and the Z-axis; and constructs cutting planes related to the ascending ramus, condyle, and coronoid process using a combination of anatomical key points, thereby achieving full quantification and repeatable definition of the partition boundaries, solving the pain points of poor consistency and strong subjectivity in manual division.

[0038] In addition, this application automatically outputs precise segmentation results of nine mandibular anatomical regions that conform to international clinical consensus through spatial enclosure and truncation processing of multiple cutting planes, realizing the full-process automation of the clinical gold standard for zoning and filling the gap in the lack of standardized automatic zoning methods in related technologies.

[0039] Based on this, the embodiments of this application provide a method for anatomical segmentation of the mandible, referring to... Figure 2 , Figure 2 This is a flowchart illustrating a method for segmenting the anatomical regions of the mandible, as provided in an embodiment of this application.

[0040] In this embodiment, the method includes: Step S10: Acquire medical images including the mandibular region; It should be noted that medical images can be CT (Computed Tomography) images or CBCT (Cone Beam Computed Tomography) images. The mandible appears as a high-grayscale area in CT images, with the outer cortical bone having high density and a high grayscale value, while the inner cancellous bone has lower density and a relatively lower grayscale value. This density difference provides the radiological basis for subsequent segmentation processing.

[0041] The medical images can be raw CT images, raw CBCT images, or standardized images that have undergone preprocessing.

[0042] For example, the process of acquiring medical images includes: first, acquiring raw medical images, which may be CT scan data in DICOM format; then, resampling the raw medical images to convert them into image data with a preset spatial resolution, for example, resampling anisotropic voxel spacing to isotropic voxel spacing (e.g., 1mm × 1mm × 1mm); then, performing standard deviation normalization on the resampled images, that is, normalizing the gray values ​​of the images to a distribution with a mean of 0 and a standard deviation of 1, in order to eliminate gray value differences caused by different scanning devices or scanning parameters. After the above preprocessing, the medical image is obtained.

[0043] Step S20: Based on the medical images, locate the mandibular region to obtain the localization result; Step S30: Based on the positioning results, a local image of the mandible is obtained by cropping from the medical image; It should be noted that the step of locating the mandibular region based on the medical images to obtain the location results is to find the approximate spatial location of the mandible in the original images.

[0044] The localization result refers to information that describes the spatial extent of the mandible in the image. It can be a three-dimensional bounding box, a centroid coordinate plus size range, or a coarse binary mask.

[0045] A partial image of the mandible refers to a sub-volume image extracted from medical images that contains only the mandible and its immediate surrounding area.

[0046] In medical images, the mandible is initially located to obtain its spatial extent (such as the bounding box), and a region of interest image containing the mandible is cropped from the original image accordingly. This is done to reduce subsequent computation and remove redundant background interference.

[0047] Step S40: Based on the local image of the mandible, obtain the mandibular mask, the segmentation results of the lower teeth, and the coordinate information of the preset anatomical landmarks; It should be noted that the mandibular mask is a binary image in which voxels belonging to the mandible are labeled as 1 (or true) and the rest as 0 (or false), depicting the morphology of the mandible.

[0048] The results of mandibular tooth segmentation include two parts: an independent binary mask for each mandibular tooth and a tooth position number (tooth number label) for each tooth, such as teeth 31-38 in the FDI tooth position representation.

[0049] Predefined coordinate information of anatomical landmarks: the coordinates of predefined anatomical points that are key to the partition in three-dimensional space.

[0050] Specifically, based on the cropped local image of the mandible, three parallel processing tasks are executed synchronously or asynchronously, and the outputs are: a fine mandibular mask, the segmentation results of the lower teeth (which include the mask of each tooth and its corresponding tooth number label), and the three-dimensional coordinate information of a series of preset anatomical landmarks.

[0051] Step S50: Based on the coordinate information of the preset anatomical landmarks and the segmentation results of the mandible, the mandibular mask is divided into segments according to the preset anatomical zoning rules to obtain the segmentation results of each anatomical subregion of the mandible.

[0052] It should be noted that the pre-defined anatomical zoning rules are based on spatial division logic derived from anatomical knowledge. For example, the upper boundary of the mandibular body region is defined as the horizontal section above the mental foramen, the lower boundary as the lower border of the mandible, the posterior boundary as the anterior notch of the angle of the mandible, and the anterior boundary as the distal section of the mental foramen, etc.

[0053] Anatomical subregion segmentation results: The results after staining or marking the mandibular mask. Each voxel not only indicates that it belongs to the mandible, but also indicates which specific anatomical subregion it belongs to (such as the condylar region, mandibular angle region, etc.).

[0054] Specifically, the extracted anatomical landmark coordinates and mandibular segmentation results (especially the position information of the tooth number label) are used as spatial reference benchmarks. The mandibular mask is spatially divided according to preset rules that conform to anatomical definitions, and finally the segmentation results of each anatomical subregion of the mandible are obtained.

[0055] For example, the segmentation results of each anatomical subregion of the mandible are obtained as follows: Images are acquired; the medical images are registered with a standard mandibular atlas template to obtain a spatial transformation matrix. This transformation matrix is ​​applied to the mandibular bounding box of the atlas template to obtain the localization result on the patient image; local images are then cropped; these local images are input into a multi-task deep learning network that shares a single encoder but has three independent decoder heads, simultaneously outputting a mandibular mask, mandibular tooth segmentation results, and landmark coordinates; finally, partitioning is performed based on rules.

[0056] For example, the segmentation results of each anatomical subregion of the mandible can also be obtained as follows: acquire images; use a 3D encoding and decoding segmentation network to obtain a probability map and threshold it to obtain the bounding box of the localization result as a binary mask; crop to obtain local images; use three independent, serial or parallel networks to obtain the mandibular mask, the segmentation results of the lower teeth, and the coordinates of the landmark points, respectively. Finally, perform partitioning.

[0057] For example, the segmentation results of each anatomical subregion of the mandible can also be obtained as follows: acquire images; integrate a spatial transformation network into an end-to-end network. This module automatically learns spatial transformation parameters and directly performs affine transformation and cropping on the input image to obtain a standardized local image of the mandible. The transformation parameters in this process are the implicit localization results; the subsequent network directly processes the standardized image to obtain masks, teeth, and landmarks; finally, partitioning is performed based on rules.

[0058] Optionally, the acquired image content includes medical images of the mandibular region, including: Acquire raw medical images; The original medical images are resampled. The resampled images are normalized to their standard deviation to obtain the medical images.

[0059] It should be noted that original medical images refer to images such as CT images or CBCT images.

[0060] The original image is resampled to achieve isotropic spatial resolution, thus eliminating differences in scanning parameters between different devices. Then, the resampled image is normalized (e.g., Z-score normalization) to convert the image grayscale values ​​into a distribution with a mean of 0 and a standard deviation of 1, which helps improve the training stability and generalization ability of deep learning models.

[0061] In this embodiment, end-to-end automatic segmentation of each anatomical subregion of the mandible was achieved, improving the efficiency and consistency of zoning and providing an anatomical basis for regional analysis and surgical planning of mandibular-related diseases.

[0062] Based on any of the above embodiments, the step of locating the mandibular region based on the medical image to obtain the location result includes: The medical images are processed using a three-dimensional encoding and decoding segmentation network to output a probability map of the mandibular region; A binary mask is obtained by threshold segmentation of the probability map, and the bounding box of the binary mask is calculated. The bounding box is then determined as the positioning result.

[0063] It should be noted that the term "3D encoder-decoder segmentation network" refers to any convolutional neural network that uses an encoder-decoder architecture for semantic segmentation of 3D images.

[0064] The outer bounding box, also known as the axis-aligned minimum bounding box, is used to describe the spatial extent of an object.

[0065] Specifically, a deep learning model, such as 3D U-Net or V-Net, is used. Its structure includes an encoding path (progressively extracting high-level features and reducing resolution) and a decoding path (progressively restoring resolution and performing pixel-level predictions). The output is a 3D probability map of the same size as the input image, where the value of each voxel represents the probability (e.g., between 0 and 1) that the voxel belongs to the mandibular region. A threshold (e.g., 0.5) is set, and the probability map is converted into a binary mask (values ​​greater than the threshold are 1, otherwise 0). Then, a minimal cuboid with sides parallel to the coordinate axes is found in 3D space that completely contains all voxels with a value of 1. The spatial coordinate range of this cuboid constitutes the localization result.

[0066] In this embodiment, a three-dimensional encoding and decoding segmentation network is used to perform preliminary localization of medical images. The spatial range of the mandible is automatically determined by using the probability map output by the network, after thresholding and bounding box calculation. This can effectively filter out interference from surrounding non-target structures such as the skull base and cervical spine, thus improving the robustness of localization.

[0067] Based on any of the above embodiments, the preset anatomical landmarks include the lowest point of the sigmoid notch and the angle of the mandible; Based on the coordinate information of the preset anatomical landmarks and the segmentation results of the lower teeth, the mandibular mask is divided into sections according to preset anatomical partitioning rules, including: Based on the coordinates of the lowest point of the sigmoid notch and the coordinates of the mandibular angle point, the boundaries of each anatomical subregion in the mandibular mask are determined.

[0068] It should be noted that this embodiment describes the specific types of preset anatomical landmarks and their roles in partitioning.

[0069] This embodiment determines the boundaries of the mandibular ramus, mandibular body, and other regions by detecting the positions of two key anatomical structures: the lowest point of the sigmoid notch and the angle of the mandible.

[0070] The lowest point of the sigmoid notch: the lowest point of the arc-shaped notch between the condyle and coronoid process, located at the upper border of the ascending ramus of the mandible.

[0071] Mandibular angle point: The most prominent point or geometric angle point in the mandibular angle region formed by the intersection of the lower border of the mandibular body and the posterior border of the mandibular ramus.

[0072] In this embodiment, the lowest point of the sigmoid notch and the angle of the mandible are used as the reference feature points for partitioning. By utilizing the characteristic that the two constitute the boundary reference between the mandibular ascending ramus and the mandibular body in anatomical terms, the division of the boundaries of each anatomical subregion has a clear skeletal morphological basis, and the partitioning results are more reasonable in terms of anatomical structure.

[0073] Based on any of the above embodiments, the preset anatomical zoning rules include the nine-zoning standard for mandibular fracture classification; The determination of the boundaries of each anatomical subregion within the mandibular mask based on the coordinates of the lowest point of the sigmoid notch and the coordinates of the angle of the mandible includes: Based on the coordinates of the lowest point of the sigmoid notch, the coordinates of the mandibular angle point, and the position information indicated by the tooth number label corresponding to the target tooth position in the lower tooth segmentation result, the boundaries of each anatomical subregion are determined to divide the mandibular mask into the condylar region, the coronoid region, the mandibular angle and ramus region, the mandibular body region, the chin and parachin region, and the transition zone of each anatomical subregion of the mandible.

[0074] It should be noted that this embodiment uses the lowest point of the sigmoid notch and the mandibular angle point, combined with the position information indicated by the tooth number label corresponding to the target tooth position in the lower tooth segmentation result (e.g., the position of the canine and the first molar) to define the boundary plane or region of each subregion.

[0075] The nine-zone classification standard for mandibular fractures is a commonly used fracture classification system in oral and maxillofacial surgery traumatology. It divides the mandible into nine zones: condylar zone, coronoid zone, mandibular angle and ramus zone, mandibular body zone, chin and parachin zone, and transitional zones between these zones.

[0076] The target tooth position refers to a specific tooth that serves as an anatomical boundary in the zoning rules. For example, the boundary between the mandibular body region and the paramental region can be defined as the distal surface of the canine. This embodiment does not impose any restrictions on the target tooth position.

[0077] In this embodiment, by integrating the three-dimensional position information indicated by the lowest point of the sigmoid notch, the angle of the mandible, and the tooth number label of the target tooth position, the mandibular bone is divided into subregions according to the clinically accepted nine-zone classification standard for mandibular fractures. The mandibular mask is divided into the condylar region, the coronoid region, the angle of the mandible and the ascending ramus region, the mandibular body region, the chin and parachin region, and the transition region. This ensures that the zoning results not only conform to the standardized anatomical definition, but also accurately define the transition boundaries between each subregion by introducing tooth position spatial reference. The zoning results can be directly adapted to the clinical diagnostic needs such as fracture classification assessment.

[0078] Based on any of the above embodiments, the lower tooth segmentation result is obtained through the following steps: The local image of the mandible is input into the tooth segmentation sub-network to obtain the segmentation result of the lower teeth and the tooth number label corresponding to each tooth in the segmentation result of the lower teeth; the tooth segmentation sub-network adopts a three-dimensional encoding and decoding segmentation network structure; the tooth segmentation sub-network is trained using a hybrid loss function that combines the Dess loss function and the cross-entropy loss function.

[0079] It should be noted that this embodiment aims to obtain accurate tooth segmentation results with tooth number labels, that is, to segment each individual tooth and identify its position.

[0080] The tooth segmentation sub-network refers to a deep learning-based image segmentation model used to perform voxel-by-voxel category prediction on the input three-dimensional mandibular local image, classifying each voxel in the image as either background or a tooth region with a specific tooth number. This embodiment does not limit the type of the tooth segmentation sub-network, as long as it can simultaneously complete the tasks of tooth instance segmentation and tooth position recognition from the three-dimensional image.

[0081] The tooth segmentation sub-network is used to receive the local image of the mandible as input and output the segmentation result of the lower teeth with tooth number labels; wherein, the segmentation result of the lower teeth is a three-dimensional labeled image with the same spatial resolution as the input image, each foreground voxel has a unique label value, which corresponds to the tooth position number of the tooth to which the voxel belongs, and different teeth have different label values, so that each tooth can be segmented and identified independently.

[0082] The 3D encoder-decoder segmentation network structure refers to an end-to-end 3D convolutional neural network with a symmetrical architecture, consisting of an encoding path and a decoding path. The encoding path extracts multi-scale spatial semantic features of the input 3D image through multi-level 3D convolution and downsampling operations. The decoding path gradually restores the spatial resolution of the feature maps through upsampling operations and fuses them with the feature maps of corresponding layers in the encoding path via skip connections to compensate for the lost spatial details during downsampling. The final output is a voxel-by-voxel category prediction result with the same size as the original input image. For example, the 3D encoder-decoder segmentation network structure can be implemented based on V-Net, 3D U-Net, or variants thereof.

[0083] A hybrid loss function combining the Descein loss function and the cross-entropy loss function is an objective function used for network training. Its total loss is a weighted combination of the Descein loss term and the cross-entropy loss term. The Descein loss function is calculated based on the Descein similarity coefficient between the predicted segmentation result and the ground truth label, measuring the overlap between the two regions and focusing on optimizing the overall shape and overlap of the segmented region. The cross-entropy loss function calculates the difference between the predicted class probability distribution and the ground truth class label on a voxel-by-voxel basis, focusing on optimizing the classification accuracy of each voxel. This hybrid loss function effectively improves the segmentation bias caused by the significant difference in the number of voxels between the foreground and background tooth regions in 3D tooth images by leveraging the robustness of the Descein loss to class imbalance and the high sensitivity of the cross-entropy loss to pixel-by-pixel classification accuracy.

[0084] For example, the tooth segmentation sub-network is trained using a hybrid loss function combining the Descein loss function and the cross-entropy loss function through the following steps: A training sample set is obtained, comprising multiple local images of the mandible and the corresponding ground truth values ​​for tooth number annotations; the training samples are input into the initial tooth segmentation sub-network to obtain predicted tooth number segmentation results; based on the predicted tooth number segmentation results and the ground truth values ​​for tooth number annotations, the Descein loss and cross-entropy loss are calculated respectively, and summed to obtain the hybrid loss; based on the hybrid loss, the network parameters of the tooth segmentation sub-network are updated using a backpropagation algorithm until the hybrid loss converges, resulting in the trained tooth segmentation sub-network.

[0085] The tooth number label refers to the numbered identifier used to identify teeth in different positions. In the dental annotation system, each tooth is assigned a specific number according to its anatomical position in the dental arch. In the segmentation results output by the tooth segmentation sub-network, voxels belonging to teeth in different positions are assigned different label values. These label values ​​correspond one-to-one with the tooth number, so that each tooth region in the segmentation results can be directly identified as the tooth with the corresponding tooth number based on its label value, achieving simultaneous completion of tooth segmentation and tooth position identification.

[0086] In this embodiment, a hybrid loss function combining Descein loss and cross-entropy loss is used to simultaneously achieve independent instance segmentation of each tooth and automatic identification of the corresponding tooth number label in the three-dimensional spatial context. This solves the problem of boundary adhesion between different teeth in the three-dimensional dental arch and the problem of missed segmentation caused by the imbalance between the foreground tooth body and the background bone tissue categories. The output segmentation results with tooth number labels can provide spatial position references for each tooth for subsequent partitioning, making the partition boundary delineation more accurate.

[0087] Based on any of the above embodiments, the mandibular mask is obtained through the following steps: The local image of the mandible is input into a fine segmentation subnetwork to obtain the mandibular mask.

[0088] It should be noted that this embodiment aims to perform binary semantic segmentation on the mandibular region to obtain a high-quality mandibular mask, providing accurate skeletal boundary constraints and region of interest definition for subsequent mandibular tooth segmentation tasks.

[0089] Since the density of the mandible in CT images typically ranges from 700 to 3000 Henle units (HU), and its structure consists of a dense outer cortical bone layer and a loose inner cancellous bone layer, by introducing a deep learning-based fine segmentation subnetwork for pixel-level prediction, the complete mandibular region can be separated more accurately from the surrounding soft tissue structure, effectively overcoming the limitations of traditional fixed threshold segmentation methods in handling noise, artifacts, and areas with uneven density.

[0090] The fine segmentation sub-network refers to a deep learning network model used for voxel-by-voxel binary classification prediction of input 3D medical images. It can label voxels belonging to the mandibular tissue as foreground and other tissues and background as background, thereby generating a corresponding mandibular mask. This embodiment does not limit the type of fine segmentation sub-network, as long as it can achieve high-precision extraction of the complete spatial morphology of the mandible from local mandibular images.

[0091] The fine segmentation sub-network receives the local image of the mandible as input and outputs a mandible mask with the same spatial dimension as the input image. The mandible mask is a binary three-dimensional image, in which voxels predicted to be mandibular regions are assigned a first label value (e.g., 1), and the remaining voxels are assigned a second label value (e.g., 0). This mask can provide an accurate anatomical domain for subsequent tooth segmentation and tooth position recognition tasks, that is, it limits the search and inference space of the tooth segmentation sub-network, thereby improving the efficiency of the overall processing flow.

[0092] In this embodiment, by introducing a fine segmentation subnetwork on the local image of the mandible, secondary fine semantic segmentation is performed within the initially located local area. This enables more accurate separation of the complete outline of the mandible from the surrounding soft tissues and bony structures, resulting in a high-quality mandibular mask. This provides a reliable spatial range benchmark and boundary constraints for the subsequent division of each anatomical subregion.

[0093] Based on any of the above embodiments, the coordinate information of the preset anatomical landmarks is obtained through the following steps: The local image of the mandible is input into the key point detection sub-network to obtain a three-dimensional Gaussian heatmap for each preset anatomical landmark. Based on the voxel coordinates in the three-dimensional Gaussian heat map whose probability values ​​meet the preset requirements, the coordinate information of the corresponding anatomical landmarks is determined.

[0094] It should be noted that this embodiment aims to achieve automatic high-precision positioning of anatomical landmarks of the mandible, in order to overcome the shortcomings of traditional manual marking methods, such as long time consumption and significant differences between observers, and to provide spatial reference information for subsequent tasks such as tooth segmentation, arch curve fitting, orthodontic treatment planning or implant placement planning.

[0095] The preset anatomical landmarks refer to a set of mandibular feature points with clear anatomical definitions pre-selected according to oral clinical anatomy standards. For example, the preset anatomical landmarks may include at least one or more of the following: condylar apex, coracoid apex, angle of the mandible, mental foramen, anterior border of the mandibular ramus, premental point, and submental point. This embodiment does not limit the specific number or selection of preset anatomical landmarks and can configure them according to specific clinical application needs.

[0096] The keypoint detection subnetwork refers to a deep learning model for locating predetermined anatomical landmarks from 3D medical images, employing a keypoint detection strategy based on heatmap regression. This subnetwork receives the local image of the mandible as input and outputs a 3D heatmap for each preset anatomical landmark. The 3D Gaussian heatmap is a 3D probability distribution map with the same spatial dimensions as the input image. The value of each voxel in the map represents the probability that the voxel belongs to the corresponding anatomical landmark location. The probability distribution follows a 3D Gaussian distribution centered on the actual location of the landmark. Regions with higher probability values ​​in the heatmap are more likely to be the location of the corresponding anatomical landmark.

[0097] The probability value meeting the preset requirement refers to the situation where the voxel probability value in the three-dimensional Gaussian heatmap meets the predetermined judgment condition. For example, the coordinates of the voxel with the highest probability value in the heatmap can be determined as the coordinate information of the corresponding anatomical landmark. Alternatively, the centroid coordinates of the voxel region with a probability value exceeding a preset threshold (e.g., 0.5) can be determined as the coordinate information of the corresponding anatomical landmark. This embodiment does not limit the preset requirement, as long as the preset requirement setting can uniquely determine the coordinate position of each anatomical landmark from the three-dimensional Gaussian heatmap.

[0098] This embodiment utilizes a keypoint detection subnetwork to perform end-to-end processing of local mandibular images, simultaneously outputting 3D Gaussian heatmaps of multiple preset anatomical landmarks, and quickly obtaining the coordinate information of each landmark through post-processing. Compared with methods based on average models and non-rigid registration, this embodiment does not rely on the construction of an accurate template library, exhibiting better robustness to normal variations in mandibular morphology between individuals; and compared with manual annotation methods, this embodiment significantly shortens the landmark localization time and achieves higher localization accuracy.

[0099] In this embodiment, a key point detection sub-network based on heatmap regression is used to output a three-dimensional Gaussian heatmap for each preset anatomical landmark, and its three-dimensional coordinates are automatically extracted through a probability peak localization strategy, thereby achieving high-precision synchronous localization of multiple anatomical landmarks.

[0100] This application also provides a mandibular anatomical partitioning device, see reference. Figure 3 The device includes: The acquisition module 301 is used to acquire medical images including images of the mandibular region. The positioning module 302 is used to locate the mandibular region based on the medical image and obtain the positioning result; The cropping module 303 is used to crop a local image of the mandible from the medical image based on the positioning result; Processing module 304 is used to obtain mandibular mask, lower tooth segmentation results, and coordinate information of preset anatomical landmarks based on the local image of the mandible. The partitioning module 305 is used to partition the mandibular mask according to the preset anatomical landmark coordinate information and the mandibular tooth segmentation result, and obtain the segmentation results of each anatomical subregion of the mandible.

[0101] In one embodiment, the positioning module 302 is further configured to process the medical image using a three-dimensional encoding and decoding segmentation network, output a probability map of the mandibular region; perform threshold segmentation on the probability map to obtain a binary mask, calculate the bounding box of the binary mask, and determine the bounding box as the positioning result.

[0102] In one embodiment, the preset anatomical landmarks include the lowest point of the sigmoid notch and the angle of the mandible; the partitioning module 305 is also used to determine the boundaries of each anatomical subregion in the mandibular mask based on the coordinates of the lowest point of the sigmoid notch and the coordinates of the angle of the mandible.

[0103] In one embodiment, the preset anatomical partitioning rules include a nine-part classification standard for mandibular fractures; the partitioning module 305 is also used to determine the boundaries of each anatomical subregion based on the coordinates of the lowest point of the sigmoid notch, the coordinates of the mandibular angle point, and the position information indicated by the tooth number label corresponding to the target tooth position in the lower tooth segmentation result, so as to divide the mandibular mask into the condylar region, the coronoid region, the mandibular angle and ramus region, the mandibular body region, the chin and parachin region, and the transition zone of each anatomical subregion of the mandible.

[0104] In one embodiment, the lower tooth segmentation result is obtained through the following steps: inputting the local image of the mandible into the tooth segmentation sub-network to obtain the lower tooth segmentation result and the tooth number label corresponding to each tooth in the lower tooth segmentation result; the tooth segmentation sub-network adopts a three-dimensional encoding and decoding segmentation network structure; the tooth segmentation sub-network is trained using a hybrid loss function that combines the Dess loss function and the cross-entropy loss function.

[0105] In one embodiment, the mandibular mask is obtained by inputting a local image of the mandible into a fine segmentation subnetwork to obtain the mandibular mask.

[0106] In one embodiment, the coordinate information of the preset anatomical landmarks is obtained through the following steps: inputting the local image of the mandible into the key point detection subnetwork to obtain a three-dimensional Gaussian heatmap for each preset anatomical landmark; and determining the coordinate information of the corresponding anatomical landmark based on the voxel coordinates in the three-dimensional Gaussian heatmap whose probability values ​​meet the preset requirements.

[0107] The aforementioned mandibular anatomical subdivision device enables end-to-end automatic subdivision of each anatomical subregion of the mandible, improving subdivision efficiency and consistency.

[0108] Based on the methods described in any of the above embodiments, this application also provides a computer storage medium storing a computer program, which, when executed by a processor, can be used to perform the methods described in any of the above embodiments.

[0109] Based on the methods described in any of the above embodiments, this application also provides a computer program product, which includes one or more computer programs or instructions. The computer program or instructions may be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. When executed by a processor, the computer program implements the methods described in any of the above embodiments.

[0110] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can also be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of apparatus, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than those marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram and / or flowchart, and combinations of blocks in block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or action, or using a combination of dedicated hardware and computer instructions.

[0111] In addition, the functional modules in the various embodiments of this application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0112] If the aforementioned functions are implemented as software functional modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0113] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of protection of this application. It should be noted that similar reference numerals and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0114] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0115] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

Claims

1. A method for anatomical partitioning of the mandible, characterized in that, The method includes: The acquired imaging content includes medical images of the mandibular region; The mandibular region is located based on the medical images, and the location results are obtained; A partial image of the mandible is obtained by cropping from the medical image based on the positioning results; Based on the local image of the mandible, the mandibular mask, the segmentation results of the lower teeth, and the coordinate information of the preset anatomical landmarks are obtained; Based on the coordinate information of the preset anatomical landmarks and the segmentation results of the lower teeth, the mandibular mask is divided into segments according to the preset anatomical zoning rules to obtain the segmentation results of each anatomical subregion of the mandible.

2. The method according to claim 1, characterized in that, The process of locating the mandibular region based on the medical images to obtain the localization result includes: The medical images are processed using a three-dimensional encoding and decoding segmentation network to output a probability map of the mandibular region; A binary mask is obtained by threshold segmentation of the probability map, and the bounding box of the binary mask is calculated. The bounding box is then determined as the positioning result.

3. The method according to claim 1, characterized in that, The pre-defined anatomical landmarks include the lowest point of the sigmoid notch and the angle of the mandible; Based on the coordinate information of the preset anatomical landmarks and the segmentation results of the lower teeth, the mandibular mask is divided into sections according to preset anatomical partitioning rules, including: Based on the coordinates of the lowest point of the sigmoid notch and the coordinates of the mandibular angle point, the boundaries of each anatomical subregion in the mandibular mask are determined.

4. The method according to claim 3, characterized in that, The preset anatomical zoning rules include the nine-zoning standard for mandibular fracture classification; The determination of the boundaries of each anatomical subregion within the mandibular mask based on the coordinates of the lowest point of the sigmoid notch and the coordinates of the angle of the mandible includes: Based on the coordinates of the lowest point of the sigmoid notch, the coordinates of the mandibular angle point, and the position information indicated by the tooth number label corresponding to the target tooth position in the lower tooth segmentation result, the boundaries of each anatomical subregion are determined to divide the mandibular mask into the condylar region, the coronoid region, the mandibular angle and ramus region, the mandibular body region, the chin and parachin region, and the transition zone of each anatomical subregion of the mandible.

5. The method according to claim 1, characterized in that, The lower tooth segmentation result is obtained through the following steps: The local image of the mandible is input into the tooth segmentation sub-network to obtain the segmentation result of the lower teeth and the tooth number label corresponding to each tooth in the segmentation result of the lower teeth; the tooth segmentation sub-network adopts a three-dimensional encoding and decoding segmentation network structure; the tooth segmentation sub-network is trained using a hybrid loss function that combines the Dess loss function and the cross-entropy loss function.

6. The method according to claim 1, characterized in that, The mandibular mask is obtained through the following steps: The local image of the mandible is input into a fine segmentation subnetwork to obtain the mandibular mask.

7. The method according to claim 1, characterized in that, The coordinate information of the preset anatomical landmarks is obtained through the following steps: The local image of the mandible is input into the key point detection sub-network to obtain a three-dimensional Gaussian heatmap for each preset anatomical landmark. Based on the voxel coordinates in the three-dimensional Gaussian heat map whose probability values ​​meet the preset requirements, the coordinate information of the corresponding anatomical landmarks is determined.

8. A mandibular anatomical partitioning device, characterized in that, The device includes: The acquisition module is used to acquire medical images containing images of the mandibular region. The positioning module is used to locate the mandibular region based on the medical images and obtain the positioning results; A cropping module is used to crop a local image of the mandible from the medical image based on the positioning result; The processing module is used to obtain the mandibular mask, the segmentation result of the lower teeth, and the coordinate information of preset anatomical landmarks based on the local image of the mandible. The partitioning module is used to partition the mandibular mask according to the preset anatomical landmark coordinate information and the mandibular tooth segmentation results, and obtain the segmentation results of each anatomical subregion of the mandible.

9. A computer-readable storage medium, characterized in that, It stores computer instructions that, when executed by a processor, implement the steps of any of the methods described in claims 1-7.

10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the method described in any one of claims 1-7.