Deep learning-based periodontal identification system

The deep learning-based peridental recognition system utilizes alignment and registration, texture guidance, and noise attention generation modules to solve the problem of automatic recognition of dental radiographs under complex conditions, achieving stable and reliable recognition and segmentation of peridental areas.

CN121236345BActive Publication Date: 2026-03-03北京优尔康健口腔门诊部有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511549097.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-28
Publication Date
2026-03-03
Estimated Expiration
2045-10-28

AI Technical Summary

Technical Problem

Under complex clinical acquisition conditions, dental radiographs are affected by factors such as differences in scale and posture, confusion between left and right sides, film noise and superposition of labels, low contrast between periodontal ligament and alveolar bone plate, and limited annotation, making it difficult for the peridental region of the tooth root to be automatically identified under unified coordinates, stable semantics and correct topology.

Method used

A deep learning-based peridental recognition system is adopted, including an alignment and registration module, a symmetry and lateral constraint module, a texture guidance construction module, a recognition network training module, and a noise attention generation module. By generating alignment reference maps, lateral masks, texture response maps, and noise masks, and combining encoder-multiscale decoder structures and topological constraint maps, stable recognition of peridental areas is achieved.

Benefits of technology

It eliminates scale and pose differences caused by multi-source acquisition, avoids left-right confusion, reduces the risk of boundary breakage and adhesion, alleviates instability caused by scarce annotations, and ensures the stability and reproducibility of recognition results across devices and scenes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121236345B_ABST
    Figure CN121236345B_ABST
Patent Text Reader

Abstract

The application discloses a periodontal identification system based on deep learning and belongs to the technical field of visual identification, and specifically comprises the following steps: generating an alignment reference image from a dental film image, registering the dental film image to a unified dental arch coordinate system, and retaining the original gray scale and the alignment reference image as basic input; obtaining a dental arch center line based on the alignment reference image, generating a side mask, limiting a candidate area band of the periodontal region, and recording a direction index from a dental crown to a root tip; applying a multi-direction filter on the alignment reference image to generate a periodontal membrane texture response graph and an alveolar bone plate texture response graph, and stacking the two graphs and the original gray scale as a guide channel; constructing a periodontal identification network, outputting a periodontal segmentation graph, and introducing a topological constraint graph for auxiliary supervision; obtaining a noise mask through self-supervised reconstruction, and using the noise mask as an attention graph inside the periodontal identification network to shield stripe and identification character textures, and outputting a periodontal segmentation graph and a boundary polyline.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of visual recognition technology, and more specifically to a deep learning-based system for identifying the periodontal portion of the tooth root. Background Technology

[0002] In recent years, automatic identification of the peridental region has evolved from traditional image processing based on thresholding and morphology to active contouring, graph cuts, and conditional random fields combined with shape priors, and further to end-to-end segmentation based on deep learning. Common approaches include directly segmenting dental radiographs using fully convolutional networks or encoder-decoder networks, or performing fine segmentation locally after tooth localization. Some methods supplement this with multi-scale filtering, texture features, and post-processing constraints to mitigate interference from imaging differences. Weakly supervised and semi-supervised strategies have also been used to address the problem of insufficient annotation, but these often rely on global contrast or empirical post-processing, making it difficult to incorporate information about the dental arch structure and orientation, and lacking robustness to film noise and scale markings.

[0003] However, under complex clinical acquisition conditions, dental radiographs are affected by factors such as scale and pose differences, left-right confusion, film noise and superimposed label characters, low contrast between the periodontal ligament and alveolar bone plate, and limited annotation, making it difficult to automatically identify the periroot region under unified coordinates, stable semantics, and correct topology. Essentially, this lacks a comprehensive mechanism that integrates spatial standardization, anatomical focus, and structural consistency. This mechanism should not only align and define the target domain at the input layer, highlight key textures at the feature layer, but also constrain connectivity and suppress non-target textures at the output layer, thereby accumulating a standardized periroot region representation that can be directly used for subsequent analysis. Summary of the Invention

[0004] The purpose of this invention is to provide a deep learning-based peridental recognition system to solve the problems in the background art:

[0005] The objective of this invention can be achieved through the following technical solutions:

[0006] A deep learning-based system for identifying the periodontal portion of the tooth root includes:

[0007] The alignment and registration module is used to generate an alignment reference map from dental radiographs, register the dental radiographs to a unified dental arch coordinate system, and retain the original grayscale and alignment reference map as the basic input.

[0008] The symmetry and lateralization definition module is used to obtain the dental midline based on the alignment reference map, generate a lateralization mask, define the candidate area zone of the periphery of the tooth root, and record the direction index from the crown to the apex;

[0009] The texture guidance building module is used to apply multi-directional filters on the alignment reference map to generate periodontal ligament texture response map and alveolar bone plate texture response map, and stack the two maps with the original grayscale to form a guide channel;

[0010] The recognition network training module is used to construct a periodontal tooth root recognition network. It adopts an encoder-multi-scale decoder structure, receives basic input and guide channels, outputs a periodontal tooth root segmentation map, and introduces a topological constraint graph for auxiliary supervision.

[0011] The noise attention generation module is used to train a film noise dictionary on unlabeled dental radiographs, obtain a noise mask through self-supervised reconstruction, and use it as a masking stripe and character texture for attention maps within the peri-root recognition network.

[0012] The inference execution module is used to sequentially complete the generation of alignment, lateral mask, texture response map and noise mask, and then feed the alignment reference map, lateral mask, texture response map, noise mask and original grayscale into the peri-root recognition network to output the peri-root segmentation map and boundary polyline.

[0013] As a further aspect of the present invention: in the alignment and registration module, the process of generating an alignment reference image from the dental radiograph, registering the dental radiograph to a unified dental arch coordinate system, and retaining the original grayscale and the alignment reference image as basic input is as follows:

[0014] Borders are removed and grayscale equalization is performed on dental radiographs. The dental arch localization network outputs initial estimates of dental arch control points and dental arch midlines, which serve as input data for dental arch modeling and coordinate system determination.

[0015] Based on the initial estimate of the dental midline, the outer edge points of the dental arch are selected, the trajectory of the dental arch center is fitted using Bézier curves, and the arc length direction from the crown to the root apex is determined.

[0016] Based on the center trajectory of the dental arch, an arc length-radius coordinate grid is constructed. The dental radiograph is unfolded according to the grid to generate an alignment reference map, and the arc length index and radial index are labeled for each pixel.

[0017] Shape-guided thin-plate spline registration is performed to map the dental radiographs to a unified dental arch coordinate system, preserving the original grayscale and outputting an alignment reference image.

[0018] As a further aspect of the present invention: in the symmetry and lateralization limiting module, the process of obtaining the dental midline based on the alignment reference diagram, generating a lateralization mask, limiting the candidate region zone of the periapical area of ​​the tooth root, and recording the directional index from the crown to the apex is as follows:

[0019] On the alignment reference map, the symmetry axis of the tooth row is extracted, and the initial curve of the tooth row midline is generated by combining grayscale ridge line tracking, and the search area is limited to the dental arch.

[0020] Based on the initial curve of the dental arch centerline, the Bezier curve is used for smoothing and the endpoints are regularized. The centerline of the interdental space is obtained by iteratively locating the center of the interdental space.

[0021] Using the midline of the dentition as the boundary, a left mask and a right mask are generated along the normal direction, and a band-shaped boundary is set on the inner and outer sides of the alveolar ridge to construct a candidate area zone for the peri-root.

[0022] A direction index is established based on the tangent direction of the tooth midline, with the direction from the crown to the apex set as the positive direction, and the direction index of each pixel is recorded in the candidate area zone of the tooth root periphery.

[0023] As a further aspect of the present invention: in the texture guidance construction module, the process of applying a multi-directional filter to the alignment reference map to generate a periodontal ligament texture response map and an alveolar bone plate texture response map, and stacking the two maps with the original grayscale to form a guidance channel is as follows:

[0024] On the alignment reference map, a multi-directional filter bank is constructed based on the discrete values ​​of the direction index, and direction-selective texture response operators are defined in the tangent direction and the normal direction, respectively, to characterize the periodontal ligament texture and the alveolar bone plate texture.

[0025] Perform orientation-selective convolution and bandpass filtering on the alignment reference map to generate periodontal ligament texture response map and alveolar bone plate texture response map, and align them pixel by pixel with the orientation index.

[0026] The periodontal ligament texture response map, alveolar bone plate texture response map, and original grayscale are stacked in a fixed channel order to form a guide channel.

[0027] As a further aspect of the present invention: the specific content of performing direction-selective convolution and bandpass filtering operations on the alignment reference map to generate a periodontal ligament texture response map and an alveolar bone plate texture response map, and aligning them pixel-by-pixel with the direction index, is as follows:

[0028] Based on the discrete values ​​of the orientation index, an orientation-selective convolution kernel is selected for each pixel of the alignment reference map, and convolution is performed in the tangent direction and the normal direction respectively to obtain a primary texture response map pair;

[0029] The primary texture response map is bandpass filtered in the frequency domain, and the tangential and normal components are aggregated to generate the periodontal ligament texture response map and the alveolar bone plate texture response map.

[0030] Based on the orientation index, pixel-level registration is performed on the two types of texture response maps to ensure that they correspond one-to-one with the orientation index in terms of size and coordinates, and the channel order is fixed.

[0031] As a further aspect of the present invention: in the recognition network training module, the process of constructing the peri-root recognition network, employing an encoder-multi-scale decoder structure, receiving basic input and guiding channels, outputting a peri-root segmentation map, and introducing a topological constraint graph for auxiliary supervision, is as follows:

[0032] An input layer is established, and the basic input and the guiding channel are aligned and normalized in a fixed order to form a multi-channel tensor, which serves as the unified input of the encoder-multi-scale decoder structure.

[0033] Construct an encoder-decoder structure. The encoder extracts texture and boundary features layer by layer, and the decoder samples, reconstructs, and fuses skip connection features at multiple scales, while maintaining spatial coordinate correspondence with the alignment reference map.

[0034] At the decoding end, a segmented output branch is set up, which is mapped to a pixel category map through convolution and activation function, while maintaining the same size and coordinates as the basic input, to obtain the peri-root segmentation map;

[0035] Training objectives are set, and a topological constraint graph is generated based on the annotations. This graph, along with the peri-root segmentation graph, participates in supervision. Pixel classification loss is used to constrain semantic consistency, and topological constraint loss is used to constrain structural relationships.

[0036] As a further aspect of the present invention: in the noise attention generation module, the process of training a film noise dictionary on unlabeled dental radiographs, obtaining a noise mask through self-supervised reconstruction, and using it as an attention map masking stripes and character textures within the peri-root recognition network is as follows:

[0037] Low-rank sparse decomposition was performed on unlabeled dental radiographs to separate structural components from noise residuals, and noise residuals were aggregated to construct a noise sample set.

[0038] Using a set of noise samples as training data, a film noise dictionary is obtained by self-supervised dictionary learning and sparse coding, and the noise reconstruction map is reconstructed and binarized to generate a noise mask.

[0039] Within the peridental recognition network, a noise mask is used as an attention map to act on early features and skip connection features of the encoder to suppress stripes and marker character textures.

[0040] As a further aspect of the present invention: in the inference execution module, the process of sequentially completing the generation of alignment, lateral mask, texture response map, and noise mask, and then feeding the alignment reference map, lateral mask, texture response map, noise mask, and original grayscale into the peri-root recognition network to output the peri-root segmentation map and boundary polyline is as follows:

[0041] Input dental X-ray images, complete registration and coordinate mapping based on alignment reference images, and retain the original grayscale and alignment reference images as the basic input for the inference stage;

[0042] Based on the dental midline and direction index, a left and right mask are generated along the normal direction, and a candidate area zone for the peri-root is defined on the inner and outer sides of the alveolar ridge.

[0043] The periodontal ligament texture response map and alveolar bone plate texture response map are calculated on the alignment reference map. The noise reconstruction map is reconstructed according to the film noise dictionary and generated by binarization.

[0044] The alignment reference image, lateral mask, periodontal ligament texture response image, alveolar bone plate texture response image, noise mask, and original grayscale are fed into the peri-root recognition network in a fixed channel order, and the peri-root segmentation map and boundary polyline are output.

[0045] The beneficial effects of this invention are:

[0046] This invention eliminates scale and pose differences caused by multi-source acquisition by registering dental radiographs to a unified dental arch coordinate system under the guidance of an alignment reference map. It establishes stable left-right segmentation and directional coding from crown to apex using the dental midline, lateral mask, and direction index, avoiding left-right confusion and providing a consistent reference for subsequent feature extraction. It generates periodontal ligament texture response map and alveolar bone plate texture response map on the alignment reference map, and stacks them with the original grayscale as a guide channel, so that weak contrast boundaries are structurally strengthened on the input side, reducing the risk of boundary breakage and adhesion caused by low contrast. At the model level, an encoder-multi-scale decoder structure is used to fuse basic input and guidance channels. A topological constraint graph is introduced for supervision to maintain the slender and continuous structure of the peri-root and output peri-root segmentation map and boundary polylines, which facilitates subsequent quantitative analysis and report generation. The film noise dictionary is trained using unlabeled dental radiographs, and a noise mask is obtained through self-supervised reconstruction. This mask is used as an attention map within the network to suppress stripes and pseudo-textures of identifier characters, thus mitigating the instability caused by scarce annotations. The inference sequence is fixed, and the alignment, lateralization, texture, and noise information flow are unified to ensure that the recognition results remain stable and reproducible across devices and scenes. Attached Figure Description

[0047] The invention will now be further described with reference to the accompanying drawings.

[0048] Figure 1 This is a schematic diagram of the module of the deep learning-based periodontal tooth root recognition system of the present invention. Detailed Implementation

[0049] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0050] Please see Figure 1 As shown, this invention is a deep learning-based periodontal root recognition system, comprising:

[0051] The alignment and registration module is used to generate an alignment reference map from dental radiographs, register the dental radiographs to a unified dental arch coordinate system, and retain the original grayscale and alignment reference map as the basic input.

[0052] In the alignment and registration module, the process of generating an alignment reference image from the dental radiograph, registering the dental radiograph to a unified dental arch coordinate system, and retaining the original grayscale and alignment reference image as basic input is as follows:

[0053] After acquiring dental radiographs, the images are first border-removed and grayscale equalized to eliminate interference from imaging boundaries and exposure differences in subsequent calculations. Then, the processed radiographs are input into the dental arch localization network. This network is a deep convolutional neural network designed specifically for dental radiographs. It employs an encoder-decoder structure to model the overall morphology of the dental arch, outputting initial estimates of arch control points and the midline. During training, manually labeled arch key points and the midline are used as supervision targets, ensuring the network provides stable geometric cues even under complex imaging conditions. Here, arch control points constrain the arch shape, and the initial estimate of the midline serves as the starting trajectory for subsequent curve fitting and coordinate system determination. For ease of understanding, the dental arch localization network can be viewed as an automated tool for identifying the arch skeleton and key anchor points from the entire image, with its output directly contributing to the arch modeling and coordinate system determination process.

[0054] Based on an initial estimate of the dental midline, a strip scan is performed along the normal direction of the curve to extract a set of stable outer edge points of the dental arch; this set of outer edge points defines the geometric boundary of the dental arch. On this basis, control points of the dental arch are selected as constraints, and a Bézier curve is used for fitting and smoothing to obtain a continuous and differentiable dental arch center trajectory. To unify the direction of subsequent parameters, the arc length direction from the crown to the apex is determined based on anatomical orientation and used as the positive direction of subsequent coordinate parameters. For example, in the incisor region, the positive arc length direction is consistent with the direction from the crown margin to the apex in the image; in the molar region, the positive arc length direction extends continuously along the dental arch center trajectory and does not reverse due to changes in image posture. Through the above processing, the dental arch center trajectory inherits the boundary information of the outer edge points and also possesses smooth properties suitable for parameterization.

[0055] Using the dental arch center trajectory as a reference, an arc-length to radius coordinate grid is constructed: the arc-length axis is unfolded along the tangent direction of the dental arch center trajectory, and the radius axis points towards the alveolar bone plate along the normal direction. Subsequently, the dental radiograph is unfolded and transformed using this coordinate grid to generate an alignment reference map, making the originally curved dental arch band appear as a regular rectangular region in the reference map. To support pixel-by-pixel indexing management in subsequent algorithms, each pixel is labeled with an arc-length index and a radial index on the alignment reference map. This process can be analogized to unfolding a curved road onto a linear coordinate system using kilometer markers, where the arc-length index is similar to the kilometer position, and the radial index is similar to the lateral distance from the center of the road to both sides. Through this unfolding, the relative positional relationship between the periodontal ligament and the alveolar bone plate on the reference map becomes clearer, and subsequent feature extraction and mask generation can be directly addressed and associated by index, reducing the uncertainty caused by the curved geometry.

[0056] Based on the aforementioned geometric parameterization, shape-guided thin-plate spline registration is performed. Using the arch control points and the arch center trajectory as shape-guided constraints, a continuous deformation field is established from the original dental radiograph to a unified arch coordinate system. The thin-plate spline is strictly aligned at the control points and propagates deformation smoothly in the remaining areas, ensuring consistent spatial correspondence of the same tooth position in different images. After registration, the dental radiograph is mapped to the unified arch coordinate system to obtain an alignment reference as described above. Figure 1 The spatial arrangement is consistent while preserving the original grayscale information as the basis for subsequent processing. For example, images of the same patient taken by different devices maintain consistent arc length and radial indices for specified tooth positions within a unified dental arch coordinate system, ensuring that subsequent segmentation and measurement are based on unified spatial semantics. Through these steps, a complete construction from the original dental radiograph to the aligned reference image and the unified dental arch coordinate system is achieved.

[0057] The symmetry and lateralization definition module is used to obtain the dental midline based on the alignment reference map, generate a lateralization mask, define the candidate area zone of the periphery of the tooth root, and record the direction index from the crown to the apex;

[0058] In the symmetry and lateralization definition module, the process of obtaining the dental midline based on the alignment reference map, generating a lateralization mask, defining the candidate region zone around the root, and recording the directional index from the crown to the apex is as follows:

[0059] On the alignment reference image, the dentition symmetry axis is first extracted based on the similarity of the left and right structures of the dental arch. Specifically, left and right reflection matching is performed on the image, the grayscale similarity and edge consistency between corresponding columns are calculated, and the dentition symmetry axis is established on the trajectory with the highest consistency. Then, guided by the dentition symmetry axis, an initial curve for the dentition midline is generated using grayscale ridge tracking. Grayscale ridge tracking selects continuous and grayscale-stable ridges along the dental arch direction as the path, ensuring the initial midline curve closely follows the direction of the dentition center. Simultaneously, the search area is limited to the dental arch to avoid interference from background structures or soft tissue shadows on the path continuity. For ease of understanding, this process can be analogized to finding the centerline in the middle of a road with similar shapes on both sides: first, find the symmetrical midline of the road, then draw an initial centerline along the continuous ridges of the road's brightness and texture, thus obtaining a robust initial midline shape.

[0060] After obtaining the initial curve of the dentition midline, a Bézier curve is used for smoothing, and the two ends of the curve are regularized to ensure that the tangent direction of the curve at the endpoints is consistent with the dental arch boundary. To improve the fit of the curve to the actual tooth position, a strip scan is performed along the normal direction of the initial midline curve. The center position of the interdental space is iteratively located using local extrema and structural texture stability criteria. The center of the interdental space reflects the geometric valley point separating adjacent tooth positions, and its location result is used to fine-tune the midline direction and node distribution to ensure that the midline remains continuous and consistent in sequence among all tooth positions. Taking the incisor to canine region as an example, the center of the interdental space roughly falls in the gray valley area at the junction of the tooth crowns; after updating the Bézier control points using this position as the anchor point, the midline smoothly passes through this point and remains consistent with the overall direction of the dental arch. After the above smoothing and positioning, the dentition midline that runs through the dental arch is obtained, laying the geometric benchmark for subsequent lateral division and region construction.

[0061] Using the midline of the dental arch as a boundary, left and right masks are generated along the normal direction of the midline. Specifically, a unit normal is calculated for each sampling point on the midline, and the normal is extended to both sides, assigning the area to the left and right masks respectively. To define the candidate area zone around the root, strip-shaped boundaries are set on the inner and outer sides of the alveolar ridge: the inner boundary is close to the root surface where the periodontal ligament is located, and the outer boundary is close to the outer edge of the alveolar bone plate. The strip-shaped boundary extends continuously along the midline, making the candidate area zone appear as a long strip-shaped region distributed around the midline throughout the dental arch. For example, the midline can be regarded as the central trajectory of the extended zone, the left and right masks respectively restrict the assigned side of the extended zone, and the inner and outer boundaries of the alveolar ridge define the near-root surface and near-bone plate range of the extended zone. Through the above division and boundary setting, the candidate area zone around the root maintains a spatial correspondence with the anatomical structure, covering the target structure while avoiding crossing into non-target areas.

[0062] A direction index is established based on the tangent direction of the dental midline, with the direction from the crown to the apex set as positive. The direction index describes the directed position parameters along the dental midline; essentially, it's a direction index map of the same size as the image, where each pixel stores its corresponding directed arc length index value. Specifically, for each pixel within the candidate region of the tooth root periphery, it is projected onto the dental midline along the normal direction. The arc length parameter of the projection point on the midline is read, and a direction index value is assigned according to the positive convention. Thus, the pixel itself doesn't possess a direction, but is labeled with a directed position index derived from the midline. For ease of understanding, the dental midline can be likened to a road with markers, where the marker starting point facing the apex is positive. Any pixel within the candidate region, after being projected onto the road centerline by the normal, can have its corresponding marker position read; this marker position is the pixel's direction index value. This index is used in subsequent processing to maintain consistency in the order across tooth positions, lateral sides, and images, and to provide a unified arc length reference for mask generation, texture convergence, and structural constraints.

[0063] The texture guidance building module is used to apply multi-directional filters on the alignment reference map to generate periodontal ligament texture response map and alveolar bone plate texture response map, and stack the two maps with the original grayscale to form a guide channel;

[0064] In the texture guidance construction module, a multi-directional filter is applied to the alignment reference map to generate a periodontal ligament texture response map and an alveolar bone plate texture response map. The process of stacking the two maps with the original grayscale to form a guidance channel is as follows:

[0065] On the alignment reference map, a multi-directional filter bank is constructed based on the discrete values ​​of the direction index, and direction-selective texture response operators are defined in the tangential and normal directions, respectively. The direction index is derived from the directional reference of the aforementioned dental midline and is used to indicate the local dominant direction at each location; its discrete values ​​constitute a finite set of discrete directions used to select the orientation of the convolution kernel. The operator in the tangential direction is more sensitive to elongated textures along the dental arch, while the operator in the normal direction is more sensitive to layered textures pointing towards the alveolar lamina. Through this setting, each location in the alignment reference map can be matched with a filter consistent with its local orientation, achieving targeted characterization of periodontal ligament texture and alveolar lamina texture. For ease of understanding, the direction index can be understood as arranging the orientation of observation for each location, and the filter bank is like a magnifying glass for observing in different directions.

[0066] Direction-selective convolution and bandpass filtering are performed on the alignment reference map to obtain periodontal ligament texture response maps and alveolar bone plate texture response maps, which are then aligned pixel-by-pixel with the direction index. Specifically, direction-selective convolution first selects convolution kernels with consistent orientation from the filter bank based on the local direction index, generating a response selectively to the texture direction; then, bandpass filtering is applied to retain frequency components that match the anatomical scale, making the response maps more focused on the layering features of the periodontal ligament and alveolar bone plate. Both types of response maps are spatially the same size and coordinate as the alignment reference map, and maintain pixel-by-pixel consistency with the direction index as a reference, providing a unified alignment benchmark for subsequent index-based addressing and convergence. For example, if the direction index at a certain location points towards the apex, the convolution kernel at that location has the same orientation, and the pixel values ​​in the two response maps are ultimately related to that orientation.

[0067] The periodontal ligament texture response map, alveolar bone plate texture response map, and original grayscale are stacked in a fixed channel order to form a guide channel. This fixed channel order maintains a consistent mapping between feature sources during training and inference, preventing semantic misinterpretation by the network due to channel substitution. The original grayscale preserves the image's density information, while the two types of texture response maps provide directional cues related to anatomical levels. All three are stacked in the same coordinate system and pixel grid, forming the basic input for subsequent models. For example, the network interprets the first channel as the periodontal ligament texture response map, the second channel as the alveolar bone plate texture response map, and the third channel as the original grayscale, thus ensuring a stable feature-semantic binding.

[0068] In a further implementation of this embodiment, the specific content of performing direction-selective convolution and bandpass filtering operations on the alignment reference map to obtain the periodontal ligament texture response map and the alveolar bone plate texture response map, and aligning them pixel-by-pixel with the direction index, is as follows:

[0069] Based on the discrete values ​​of the orientation index, a direction-selective convolution kernel is selected for each pixel of the alignment reference map, and convolution is performed in both the tangential and normal directions to obtain primary texture response map pairs. This process completes the orientation matching of the kernels at the pixel level: when the orientation index of a pixel represents the dominant local orientation, the convolution kernel in the tangential direction extracts strip and fibrous signals along that orientation, while the convolution kernel in the normal direction extracts layered signals in the direction perpendicular to it. In this way, two complementary primary response maps can be obtained simultaneously at the same location, carrying texture information along the arc length direction and the radial direction, respectively. For example, near the root apex, the tangential direction response is more likely to show fine lines extending along the dental arch, while the normal direction response is more likely to reveal textures pointing towards the alveolar bone plate.

[0070] A bandpass filter is applied to the primary texture response map in the frequency domain, aggregating the tangential and normal components to generate the periodontal ligament texture response map and the alveolar lamina texture response map. The bandpass filter suppresses irrelevant, slow background changes and sharp discrete noise, highlighting the texture layering within the structural scale. Subsequently, aggregation is performed based on the orientation of the directional components: components related to the arc length direction are merged into the periodontal ligament texture response map, and components related to the radial direction are merged into the alveolar lamina texture response map. After this processing, the two maps form physically separable channels: the former is biased towards the thin-film texture along the tooth row, while the latter is biased towards the lamellar texture pointing towards the lamina. Both share the same coordinates and grid, facilitating subsequent stacking with the original grayscale values.

[0071] Based on the orientation index, pixel-level registration is performed on the two types of texture response maps to ensure a one-to-one correspondence between them and the orientation index in terms of size and coordinates. The channel order is fixed to facilitate subsequent processing. Here, the orientation index of each pixel refers to an index map of the same size as the alignment reference map. Each pixel position in the index map stores its directed position parameter (directed arc length index) relative to the tooth centerline, with the direction from the crown to the root defined as positive. The implementation is as follows: for any pixel within the candidate region, it is projected onto the tooth centerline along the normal direction. The arc length parameter of the projection point on the centerline is read and written into the index map at the same coordinate position. This can be analogized to a road marker system: the tooth centerline is the central road, and the marker length increases from the crown to the root. Any pixel projected onto the road via the normal yields a corresponding marker length value, which is the orientation index of that pixel. In this way, a stable pixel-level correspondence is established between the texture response map and the orientation index, facilitating subsequent feature addressing, region constraints, and anatomical consistency checks based on the index.

[0072] The recognition network training module is used to construct a periodontal tooth root recognition network. It adopts an encoder-multi-scale decoder structure, receives basic input and guide channels, outputs a periodontal tooth root segmentation map, and introduces a topological constraint graph for auxiliary supervision.

[0073] In the recognition network training module, the process of constructing a periodontal root recognition network, using an encoder-multi-scale decoder structure, receiving basic input and guidance channels, outputting a periodontal root segmentation map, and introducing a topological constraint graph for auxiliary supervision is as follows:

[0074] In this embodiment, an input layer is first established. The basic input and guiding channels are aligned and normalized according to a fixed channel order, and combined into a multi-channel tensor, which serves as the unified input for the encoder-multi-scale decoder structure. The basic input refers to the alignment reference map and the original grayscale, carrying geometric alignment and density information; the guiding channels refer to the periodontal ligament texture response map and the alveolar bone plate texture response map, carrying directional texture cues. To ensure the comparability of images from different batches, the channel order remains consistent during the training and inference phases, and the normalization method and coordinate system are consistent with the alignment reference. Figure 1 This ensures that the network receives data under the same spatial reference. For example, dental images acquired by different devices and in different batches, after registration and channel normalization, have the following structure: the first channel is the periodontal ligament texture response map, the second channel is the alveolar bone plate texture response map, the third channel is the original grayscale, and the fourth channel is the alignment reference map. This fixed order ensures that subsequent convolutional kernels read the same semantic input at the same position, avoiding feature-semantic misalignment caused by channel substitution.

[0075] Based on a unified input, an encoder-decoder structure is constructed. The encoder employs layer-by-layer downsampling and convolutional extraction to aggregate texture and boundary features through progressively deeper receptive fields, while preserving skip connection features to transmit local details. The decoder performs multi-scale upsampling and reconstruction, fusing the features with corresponding skip connection features at each scale to gradually restore spatial resolution. To ensure a one-to-one correspondence between the output and the alignment reference image, the decoder maintains the spatial coordinate correspondence with the alignment reference image during each upsampling stage, avoiding coordinate drift caused by interpolation and boundary clipping. For example, when the encoder extracts global cues of the peri-root texture and bone plate boundary at a deeper layer, the corresponding shallow skip connections still retain fine line information of the thin film-like boundary. By fusing both, the decoder ensures that the final feature includes both the overall layout of the structural hierarchy and maintains pixel-level boundary continuity, thus providing a stable and detailed representation for subsequent segmentation.

[0076] A segmentation output branch is set up at the decoding end. Through convolution and activation functions, the fused features are mapped to a pixel category map, outputting a peridental segmentation map that maintains consistency with the size and coordinates of the base input. Pixel category division is based on task requirements, typically including a peridental target category and a non-target background category. If necessary, auxiliary boundary categories are added to enhance the separability of transition zones. To avoid offset caused by inconsistent sizes, the segmentation output branch uses a convolution kernel with the same resolution as the last-stage decoded features and is strictly aligned with the coordinate grid of the alignment reference map. For example, if a pixel corresponds to a narrow area at the junction of the periodontal ligament and alveolar bone plate on the alignment reference map, the segmentation output should provide a matching category label for that pixel location. If the input resolution or cropping window changes, coordinate alignment and output of the same size ensure that the prediction for that pixel still falls in the same spatial location, thus maintaining the traceability and comparability of the segmentation results.

[0077] During the training phase, a joint loss is defined, and a topological constraint graph is generated based on the annotations. This graph, along with the periapical segmentation map, participates in supervision. Pixel classification loss is used to constrain semantic consistency, while topological constraint loss is used to constrain structural relationships. Their meanings are as follows: Pixel classification loss aims to compare each pixel against the labeled category, measuring the consistency between the predicted category and the true category for each pixel. This ensures stable semantic annotations for the same anatomical region under different images and imaging conditions, thus forming semantically consistent pixel-level predictions. Topological constraint loss aims to constrain the structural relationships described by the topological constraint graph, limiting connectivity, the number of pores, the continuity of thin bands, and the number of branches, ensuring that the predicted results structurally conform to the anatomical morphology of the periapical region. For example, in noisy images, pixel classification loss might cause the model to classify high-confidence regions as periapical, but breaks or isolated islands may still appear at the boundaries. In this case, topological constraint loss penalizes locations in the topological constraint graph that are connected but broken, and corrects regions that correspond to thin bands but are locally swollen, guiding the network to correct structural defects without changing the overall semantics. For example, in certain areas, the periodontal ligament should geometrically appear as a long, narrow, connected band. Topological constraint loss restricts invalid bifurcations and meaningless cavities within this band, thus maintaining consistency with anatomical relationships. Ultimately, the two types of losses work synergistically in the same objective function: the former handles pixel-semantic alignment, and the latter handles structure-morphology alignment. Through the above training strategy, the network obtains segmentation results consistent with anatomical topology while maintaining semantic accuracy; this result is directly output as a periapical segmentation map during the inference phase.

[0078] The noise attention generation module is used to train a film noise dictionary on unlabeled dental radiographs, obtain a noise mask through self-supervised reconstruction, and use it as a masking stripe and character texture for attention maps within the peri-root recognition network.

[0079] In the noise attention generation module, the process of training a film noise dictionary on unlabeled dental radiographs, obtaining a noise mask through self-supervised reconstruction, and using it as an attention map masking stripes and character textures within the peridental recognition network is as follows:

[0080] To distinguish between anatomical structures and imaging noise in unlabeled dental radiographs, this embodiment performs low-rank sparse decomposition on each radiograph. The structural components in the decomposition result exhibit slowly changing, continuous anatomical morphology, mainly including the grayscale structures of the tooth root, periodontal ligament band, and alveolar bone plate; the noise residuals are mainly represented by stripes, markings, scratches, and indentations. To form a sample source that can be used for learning, the noise residuals are cropped and screened, removing abnormal blocky areas caused by metal artifacts or overexposure, retaining patches with stroke, stripe, and fine line features, and recording their positional relationship in the original image. By repeating the above process on multiple unlabeled dental radiographs, a noise sample set is obtained, which covers common film stripes, printed character strokes, positioning marks, and border indentations. For example, the printed area containing the hospital name usually appears as fine, high-contrast strokes, which clearly fall into the noise residuals after decomposition, while the tooth root and bone plate contours are still retained in the structural components.

[0081] Using a set of noise samples as training data, a film noise dictionary is obtained through self-supervised dictionary learning and sparse coding. Self-supervised learning here refers to learning dictionary atoms that can represent noise patterns by minimizing reconstruction errors and sparsity constraints without relying on manual annotation, making the dictionary representative in terms of stripe periodicity, character stroke direction, and edge abrupt changes. After obtaining the film noise dictionary, sparse coefficients are calculated for any dental radiograph under this dictionary, and a noise reconstruction map is reconstructed. Subsequently, the noise reconstruction map is binarized to generate a noise mask. The binarization stage combines connectivity and local contrast to suppress isolated points and weak pseudo-responses, while preserving continuous strokes and bundled stripes. The resulting noise mask is spatially identical in size and coordinates to the original image, clearly marking the locations of stripes and identifying character textures. For example, if a date character exists in the lower left corner of the image, the noise reconstruction map shows a continuous high response at that location, forming a connected region on the noise mask after binarization that matches the character strokes; while the narrow anatomical textures around the tooth root have a weaker response because they do not conform to the shape of the dictionary atoms, and are therefore not covered by the mask.

[0082] Within the peri-root recognition network, a noise mask is used as an attention map, applied to early encoder features and skip connection features. Specifically, at locations marked as noise by the noise mask, the corresponding feature responses are suppressed with pixel-level weights; at unmarked locations, the original response intensity and texture details are preserved. Since the noise mask is the same size and coordinates as the original image, this pixel-level effect does not alter the spatial correspondence; suppression is limited to the marked stripes and character textures, while other anatomical structures remain stably transmitted. This process prevents high-frequency interference caused by stripes and characters from being transmitted to the decoder along skip connections, thus avoiding misjudgment boundaries during segmentation. For example, if a transverse film stripe crosses the peri-root candidate region, the attention map suppresses stripe-related responses in the early encoder features at that stripe location, allowing the decoder to restore the boundary direction consistent with the anatomy based on the guide channel and basic input, rather than deviating along the stripe direction. Similarly, since the strokes of character markers often do not align with the anatomical texture direction, after the mask explicitly masks them, the network's edge detection and region determination are no longer affected by the character strokes. In this way, the noise mask functions stably within the network in the form of an attention map, without disrupting the original coordinate system or affecting the transmission of anatomical textures in non-noise areas.

[0083] The inference execution module is used to sequentially complete the generation of alignment, lateral mask, texture response map and noise mask, and then feed the alignment reference map, lateral mask, texture response map, noise mask and original grayscale into the peri-root recognition network to output the peri-root segmentation map and boundary polyline.

[0084] In the inference execution module, alignment, lateral masking, texture response map, and noise mask generation are completed sequentially. The alignment reference map, lateral masking, texture response map, noise mask, and original grayscale are then fed into the peri-root recognition network to output the peri-root segmentation map and boundary polyline.

[0085] The inference phase begins with inputting the dental radiograph to be processed. Based on the constructed alignment reference map, the radiograph is registered and its coordinates mapped to ensure it resides in a unified dental arch coordinate system. After mapping, the system retains both the original grayscale and the alignment reference map as the foundational input for the inference phase. The original grayscale carries density and detail information, while the alignment reference map provides geometric references and unfolding relationships. To illustrate this, the process can be likened to unifying images of the same road segment taken by different cameras onto the same map grid: regardless of changes in the original viewpoint, the position of the same tooth within the reference system remains relatively stable after coordinate mapping.

[0086] In a unified coordinate system, based on the dental midline and direction index, left and right masks are generated along the normal direction of the midline to distinguish the lateral affiliation of the dental arch. Subsequently, using the inner and outer sides of the alveolar ridge as constraints, a band-shaped boundary is defined to construct candidate regions in the peridental region. These candidate regions continuously expand around the dental midline to limit the scope of subsequent features and segmentation, avoiding crossing non-target anatomical areas. For example, in the left molar region, the left mask restricts features and predictions to the left dental arch, and the band-shaped boundary focuses processing on the strip region near the periodontal ligament and alveolar bone plate, thus maintaining a spatial correspondence with the anatomical location.

[0087] On the aligned reference map, two types of directional textures are calculated according to a predetermined direction selection mechanism: one is the periodontal ligament texture response map, used to represent the thin film texture along the dental arch; the other is the alveolar bone plate texture response map, used to represent the lamellar texture pointing towards the bone plate. Simultaneously, the input image is reconstructed based on the film noise dictionary to obtain a noise reconstruction map, which is then binarized to generate a noise mask. The noise mask has the same size and coordinates as the original image and is used to identify the locations of stripes and characters. For example, if date characters appear at the corner of the image, the noise reconstruction map shows a continuous response at that location, and the corresponding noise mask after binarization forms a connected region that matches the stroke; while the narrow structures around the root are excluded from the noise region by the mask because their shape is inconsistent with the dictionary atoms. All three types of maps (the two texture response maps and the noise mask) are pixel-by-pixel aligned with the direction index for easy subsequent unified processing.

[0088] The alignment reference map, lateral mask, periodontal ligament texture response map, alveolar bone plate texture response map, noise mask, and original grayscale are fed into the peridental recognition network in a fixed channel order. This fixed channel order ensures that the same semantic meaning corresponds to the same channel position during training and inference, avoiding semantic confusion caused by channel substitution. The recognition network performs fusion inference on the above multi-channel inputs under a unified dental arch coordinate system, outputting a peridental segmentation map and extracting boundary polylines from the segmentation results. Boundary polylines can be obtained using contour tracking or refinement-skeletonization and polyline fitting: the former extracts closed or open contour lines at a given threshold in the segmentation probability map, while the latter simplifies boundary pixels in binary segmentation and connects them into polylines. For example, when a narrow gap exists between adjacent teeth within a candidate region, the segmentation map presents a corresponding shape at that location, and the boundary polyline is continuously drawn along this shape, corresponding one-to-one with the coordinate grid of the alignment reference map. At this point, the inference stage sequentially completes four steps: registration and coordinate mapping, lateral mask and candidate region construction, directional texture and noise mask generation, and fixed channel input and result output, forming a complete process from input to periapical segmentation map and boundary polyline.

[0089] The foregoing has provided a detailed description of one embodiment of the present invention, but this description is merely a preferred embodiment and should not be construed as limiting the scope of the invention. All equivalent variations and modifications made within the scope of the claims of this invention should still fall within the patent coverage of this invention.

Claims

1. A deep learning-based periodontal root recognition system, characterized in that, include: The alignment and registration module is used to generate an alignment reference map from dental radiographs, register the dental radiographs to a unified dental arch coordinate system, and retain the original grayscale and alignment reference map as the basic input. The symmetry and lateralization definition module is used to obtain the dental midline based on the alignment reference map, generate a lateralization mask, define the candidate area zone of the periphery of the tooth root, and record the direction index from the crown to the apex; The texture guidance building module is used to apply multi-directional filters on the alignment reference map to generate periodontal ligament texture response map and alveolar bone plate texture response map, and stack the two maps with the original grayscale to form a guide channel; The recognition network training module is used to construct a periodontal tooth root recognition network. It adopts an encoder-multi-scale decoder structure, receives basic input and guide channels, outputs a periodontal tooth root segmentation map, and introduces a topological constraint graph for auxiliary supervision. The noise attention generation module is used to train a film noise dictionary on unlabeled dental radiographs, obtain a noise mask through self-supervised reconstruction, and use it as a masking stripe and character texture for attention maps within the peri-root recognition network. The inference execution module is used to sequentially complete the generation of alignment, lateral mask, texture response map and noise mask, and then feed the alignment reference map, lateral mask, texture response map, noise mask and original grayscale into the peri-root recognition network to output the peri-root segmentation map and boundary polyline.

2. The deep learning-based periodontal root identification system according to claim 1, characterized in that, In the alignment and registration module, the process of generating an alignment reference image from the dental radiograph, registering the dental radiograph to a unified dental arch coordinate system, and retaining the original grayscale and alignment reference image as basic input is as follows: Borders are removed and grayscale equalization is performed on dental radiographs. The dental arch localization network outputs initial estimates of dental arch control points and dental arch midlines, which serve as input data for dental arch modeling and coordinate system determination. Based on the initial estimate of the dental midline, the outer edge points of the dental arch are selected, the trajectory of the dental arch center is fitted using Bézier curves, and the arc length direction from the crown to the root apex is determined. Based on the center trajectory of the dental arch, an arc length-radius coordinate grid is constructed. The dental radiograph is unfolded according to the grid to generate an alignment reference map, and the arc length index and radial index are labeled for each pixel. Shape-guided thin-plate spline registration is performed to map the dental radiographs to a unified dental arch coordinate system, preserving the original grayscale and outputting an alignment reference image.

3. The deep learning-based periodontal root identification system according to claim 1, characterized in that, In the symmetry and lateralization definition module, the process of obtaining the dental midline based on the alignment reference map, generating a lateralization mask, defining the candidate region zone around the root, and recording the directional index from the crown to the apex is as follows: On the alignment reference map, the symmetry axis of the tooth row is extracted, and the initial curve of the tooth row midline is generated by combining grayscale ridge line tracking, and the search area is limited to the dental arch. Based on the initial curve of the dental arch centerline, the Bezier curve is used for smoothing and the endpoints are regularized. The centerline of the interdental space is obtained by iteratively locating the center of the interdental space. Using the midline of the dentition as the boundary, a left mask and a right mask are generated along the normal direction, and a band-shaped boundary is set on the inner and outer sides of the alveolar ridge to construct a candidate area zone for the peri-root. A direction index is established based on the tangent direction of the tooth midline, with the direction from the crown to the apex set as the positive direction, and the direction index of each pixel is recorded in the candidate area zone of the tooth root periphery.

4. The deep learning-based periodontal root identification system according to claim 1, characterized in that, In the texture guidance construction module, the process of applying a multi-directional filter to the alignment reference map to generate a periodontal ligament texture response map and an alveolar bone plate texture response map, and stacking the two maps with the original grayscale to form a guidance channel, is as follows: On the alignment reference map, a multi-directional filter bank is constructed based on the discrete values ​​of the direction index, and direction-selective texture response operators are defined in the tangent direction and the normal direction, respectively, to characterize the periodontal ligament texture and the alveolar bone plate texture. Perform orientation-selective convolution and bandpass filtering on the alignment reference map to generate periodontal ligament texture response map and alveolar bone plate texture response map, and align them pixel by pixel with the orientation index. The periodontal ligament texture response map, alveolar bone plate texture response map, and original grayscale are stacked in a fixed channel order to form a guide channel.

5. The deep learning-based periodontal root identification system according to claim 4, characterized in that, The specific content of performing direction-selective convolution and bandpass filtering operations on the alignment reference map to generate periodontal ligament texture response map and alveolar bone plate texture response map, and aligning them pixel-by-pixel with the direction index, is as follows: Based on the discrete values ​​of the orientation index, an orientation-selective convolution kernel is selected for each pixel of the alignment reference map, and convolution is performed in the tangent direction and the normal direction respectively to obtain a primary texture response map pair; The primary texture response map is bandpass filtered in the frequency domain, and the tangential and normal components are aggregated to generate the periodontal ligament texture response map and the alveolar bone plate texture response map. Based on the orientation index, pixel-level registration is performed on the two types of texture response maps to ensure that they correspond one-to-one with the orientation index in terms of size and coordinates, and the channel order is fixed.

6. The deep learning-based periodontal root identification system according to claim 1, characterized in that, In the aforementioned recognition network training module, the process of constructing the peri-root recognition network, employing an encoder-multi-scale decoder structure, receiving basic input and guiding channels, outputting a peri-root segmentation map, and introducing a topological constraint graph for auxiliary supervision, is as follows: An input layer is established, and the basic input and the guiding channel are aligned and normalized in a fixed order to form a multi-channel tensor, which serves as the unified input of the encoder-multi-scale decoder structure. Construct an encoder-decoder structure. The encoder extracts texture and boundary features layer by layer, and the decoder samples, reconstructs, and fuses skip connection features at multiple scales, while maintaining spatial coordinate correspondence with the alignment reference map. At the decoding end, a segmented output branch is set up, which is mapped to a pixel category map through convolution and activation function, while maintaining the same size and coordinates as the basic input, to obtain the peri-root segmentation map; Training objectives are set, and a topological constraint graph is generated based on the annotations. This graph, along with the peri-root segmentation graph, participates in supervision. Pixel classification loss is used to constrain semantic consistency, and topological constraint loss is used to constrain structural relationships.

7. The deep learning-based periodontal root identification system according to claim 1, characterized in that, In the noise attention generation module, the process of training a film noise dictionary on unlabeled dental radiographs, obtaining a noise mask through self-supervised reconstruction, and using it as an attention map to mask stripes and character textures within the peri-root recognition network is as follows: Low-rank sparse decomposition was performed on unlabeled dental radiographs to separate structural components from noise residuals, and noise residuals were aggregated to construct a noise sample set. Using a set of noise samples as training data, a film noise dictionary is obtained by self-supervised dictionary learning and sparse coding, and the noise reconstruction map is reconstructed and binarized to generate a noise mask. Within the peridental recognition network, a noise mask is used as an attention map to act on early features and skip connection features of the encoder to suppress stripes and marker character textures.

8. The deep learning-based periodontal root identification system according to claim 1, characterized in that, In the inference execution module, the process of sequentially completing alignment, lateral mask generation, texture response map generation, and noise mask generation, and then feeding the alignment reference map, lateral mask, texture response map, noise mask, and original grayscale values ​​into the peri-root recognition network to output the peri-root segmentation map and boundary polyline is as follows: Input dental X-ray images, complete registration and coordinate mapping based on alignment reference images, and retain the original grayscale and alignment reference images as the basic input for the inference stage; Based on the dental midline and direction index, a left and right mask are generated along the normal direction, and a candidate area zone for the peri-root is defined on the inner and outer sides of the alveolar ridge. The periodontal ligament texture response map and alveolar bone plate texture response map are calculated on the alignment reference map. The noise reconstruction map is reconstructed according to the film noise dictionary and generated by binarization. The alignment reference image, lateral mask, periodontal ligament texture response image, alveolar bone plate texture response image, noise mask, and original grayscale are fed into the peri-root recognition network in a fixed channel order, and the peri-root segmentation map and boundary polyline are output.

Citation Information

Patent Citations

  • WTNet pediatric lower jaw wisdom tooth embryo segmentation network architecture method

    CN118038057A

  • Dental maxillofacial image recognition method and system based on deep learning

    CN119006899A