Malocclusion nonradiative primary screening system and method based on deep learning
The generation of virtual skull lateral films through deep learning technology solves the radiation risks and cost problems in screening for malignant jaw deformities, and achieves low-cost and high-accuracy radiation-free screening, reducing labor costs and time costs.
Patent Information
- Application Number
- CN202510304991.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-08-08
AI Technical Summary
The prior art has radiation risks and cost barriers in screening of malignant deformities, making it difficult to obtain reliable image diagnostic basis through low-cost equipment, and lacks automated diagnostic methods.
A cross-modal image generation model based on deep learning and an extracted marker point model are adopted to generate virtual skull lateral films through ordinary photos, and an attention module and an adaptive lighting normalization module are used to improve image accuracy. Combined with a convolutional neural network and a graph convolutional network, an automatic identification of bone marker points is generated to generate diagnostic suggestions.
A low-cost, radiation-free initial screening of malformed jaw deformities is achieved, and high-precision virtual skull lateral films are generated, which reduces labor costs and time costs, and improves the accuracy and efficiency of screening.
Smart Images

Figure CN120451036A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent orthodontics, and in particular relates to a radiation-free initial screening system and method for malocclusion based on deep learning. Background Art
[0002] Generally speaking, accurate diagnosis of malocclusion relies on lateral cephalograms, which accurately reflect the bony relationship between the jaw and teeth. However, existing technologies have two core contradictions: 1. Radiation risk and ethical restrictions: The effective radiation dose of a single lateral cephalogram is approximately 2 to 5 μSv, and the effective radiation dose of a single wide-field oral cone-beam computed tomography scan is approximately 100 to 200 μSv. Children and pregnant women should use these devices with caution, and repeated scans increase the cumulative risk. 2. Cost barriers, equipment dependence, and professional dependence: The purchase cost of X-ray equipment ranges from 100,000 to 1 million yuan, and the functional penetration rate is approximately 65% to 80% depending on the region. Although some institutions are equipped with DR equipment, they may not fully utilize this function due to insufficient technical training or outdated equipment, and the proportion of institutions that actually conduct examinations may be even lower.
[0003] Due to radiation risks and cost barriers, the prevalence of malocclusion among 12-year-old children in my country is 72%, but less than 40% actually receive screening; about 80% of adults have varying degrees of occlusion problems, but only 12% actively seek screening or consultation.
[0004] Therefore, a low-cost, radiation-free malocclusion screening system or method is needed to address the radiation risks and cost barriers associated with malocclusion screening. Existing technologies need to address the following technical issues: 1. How to perform screening using images acquired with low-cost equipment while ensuring diagnostic reliability; 2. How to convert two-dimensional images across modalities to lateral cephalograms to provide a basis for screening and diagnosis; and 3. How to automatically identify the bone structure of lateral cephalograms and generate diagnostic recommendations, reducing both labor and time costs. Summary of the Invention
[0005] In order to solve the above technical problems, the present invention provides a radiation-free initial screening system and method for malocclusion based on deep learning.
[0006] The present invention is achieved through the following technical solutions.
[0007] The present invention provides a deep learning-based non-radiation initial screening system for malocclusion, comprising:
[0008] Cross-modal image generation model: used to extract features of the input image, calculate the attention weights of the features, obtain feature maps based on the attention weights, and generate virtual cephalograms based on the feature maps;
[0009] Landmark extraction model: used to extract landmarks from virtual lateral skull radiographs, obtain bony landmarks, and generate diagnostic recommendations based on the bony landmarks;
[0010] The attributes of the input image include: RGB format, high resolution, and patient posture;
[0011] The cross-modal image generation model includes: an attention module, an adaptive illumination normalization module, an encoder, a decoder, a skip connection module and a loss function.
[0012] Preferably, the RGB format attributes of the input image can be used to obtain skin and mucous membrane color features;
[0013] The high resolution of the input image is used to capture subtle anatomical structures;
[0014] The subtle anatomical structures include the mandibular margin, nasal tip, and soft tissue nasion;
[0015] The input image has specific requirements for the patient's posture attributes: the patient maintains a natural head position, and the line connecting the tragus and infraorbital point is parallel to the ground;
[0016] The line connecting the tragus and the infraorbital point is used as a reference plane, and the reference plane is aligned to ensure spatial consistency between the model input and the training data.
[0017] Preferably, the attention module is used to calculate attention weight;
[0018] The adaptive illumination normalization module is used to eliminate the influence of ambient light differences on the input image;
[0019] The encoder is used to extract input image features and obtain a feature map based on attention weights;
[0020] The decoder is used to restore the feature map to the size of the input image and generate a virtual cephalogram;
[0021] The skip connection module is used to connect each layer of the encoder to the corresponding layer of the decoder;
[0022] The loss function is used to determine the accuracy of the virtual skull lateral film and update the model;
[0023] Anatomical features are embedded in the loss function.
[0024] Preferably, the attention module includes:
[0025] Multi-level feature alignment unit: used to align spatial positions layer by layer;
[0026] Dynamic weight allocation unit: used to dynamically allocate attention weights based on space and channels;
[0027] Decoder guided generation unit: used to further modify the distribution of attention weights according to the loss function.
[0028] Preferably, the dynamic weight allocation unit includes a spatial dynamic allocation submodule and a channel dynamic allocation submodule;
[0029] The spatial dynamic allocation submodule specifically includes: locating the area corresponding to the bony landmark in the input image and assigning weights;
[0030] The channel dynamic allocation submodule specifically includes: assigning weights to different channel features;
[0031] The channel dynamic allocation submodule includes: strengthening channels related to bone density in the feature library.
[0032] Preferably, the adaptive illumination normalization module includes:
[0033] Lighting estimation unit: used to learn the lighting distribution parameters of the input image;
[0034] Normalization calculation unit: used to adjust the normalized intensity according to the illumination distribution parameters.
[0035] Preferably, extracting the landmark point model includes:
[0036] Convolutional neural network: used to detect key points of virtual skull lateral radiographs and obtain bony landmarks;
[0037] Graph convolutional network: Modeling the spatial relationship of bony landmarks and calculating clinical parameters;
[0038] Diagnostic suggestion module: Generates diagnostic suggestions based on clinical parameters.
[0039] A deep learning-based non-radiation initial screening method for malocclusion is implemented using the aforementioned non-radiation initial screening system for malocclusion, comprising the following steps:
[0040] S1. Get input image;
[0041] S2, the cross-modal image generation model generates a virtual lateral skull radiograph based on the input image;
[0042] S3, the landmark extraction model extracts bony landmarks based on the virtual skull lateral radiograph and generates diagnostic suggestions based on the bony landmarks;
[0043] S4. Upload the virtual skull lateral film and bony landmarks, and update the cross-modal image generation model and landmark extraction model.
[0044] Preferably, the step S2, wherein the cross-modal image generation model generates a virtual cephalogram according to the input image, comprises the following steps:
[0045] S21, the adaptive illumination module eliminates the ambient light difference of the input image;
[0046] S22, the attention module calculates the attention weight;
[0047] S23, the encoder extracts the features of the input image according to the attention weight and obtains a feature map;
[0048] S24. The decoder restores the feature map to the size of the input image and generates a virtual lateral skull film.
[0049] Preferably, the step S3, extracting the landmark point model, extracting the bony landmark points based on the virtual lateral skull film, and generating a diagnosis suggestion based on the bony landmark points comprises the following steps:
[0050] S31, convolutional neural network extracts bony landmarks based on virtual lateral skull radiographs;
[0051] S32, graph convolutional network models the spatial relationship of bony landmarks and calculates clinical parameters;
[0052] S33. The diagnosis suggestion module generates diagnosis suggestions based on clinical parameters.
[0053] The beneficial effects of the present invention are:
[0054] 1. Generate high-precision cephalographic films from ordinary photos, breaking through the bottleneck of non-radiation malocclusion diagnosis and achieving low-cost and high-accuracy preliminary screening for malocclusion;
[0055] 2. By embedding anatomical features in the loss function and introducing an attention module into the cross-modal image generation model, the accuracy and medical credibility of the generated images are improved;
[0056] 3. By extracting landmark point models to identify virtual lateral skull radiographs, clinical parameters and diagnostic evidence can be further obtained without physician judgment. This allows for fast recognition, further reducing labor and time costs.
[0057] 4. Through the cross-modal image generation model and the landmark point extraction model, the virtual skull lateral view, quantitative clinical parameters and diagnostic basis are output synchronously, which can increase the judgment basis and enhance the reliability during manual screening. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] Figure 1 Schematic diagram of the system structure provided by an embodiment of the present invention;
[0059] Figure 2 is a schematic diagram of the structure of a cross-modal image generation model provided by an embodiment of the present invention;
[0060] Figure 3is a flow chart of a method provided by an embodiment of the present invention;
[0061] Figure 4 is a schematic diagram of a virtual lateral skull radiograph provided by an embodiment of the present invention;
[0062] Figure 5 Schematic diagram of bony landmarks provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0063] The technical solution of the present invention is further described below, but the scope of protection claimed is not limited to the description.
[0064] The structure of a non-radiation initial screening system for malocclusion based on deep learning is as follows: Figure 1 As shown, it includes: a cross-modal image generation model and a landmark point extraction model;
[0065] Cross-modal image generation model: used to extract features of the input image, calculate the attention weights of the features, obtain feature maps based on the attention weights, and generate virtual cephalograms based on the feature maps;
[0066] The attributes of the input image include: RGB format, high resolution, and patient posture;
[0067] The attribute features of the input image in RGB format can be used to obtain the color features of skin and mucous membranes;
[0068] Among them, the present invention obtains the color characteristics of skin and mucous membranes based on the information of RGB channels; therefore, it is necessary to limit the image attributes to RGB format; further, the color information of the image is extracted through the RGB channel information, and the skin and mucous membrane colors are extracted, which can be used to learn the model to infer the implicit soft tissue structure clues therein and assist in inferring deep bony relationships; on the other hand, ordinary cameras output RGB images by default, and the use of RGB format lowers the user usage threshold, which is in line with the invention goal of "low cost and easy popularization".
[0069] The high resolution of the input image is used to capture subtle anatomical structures;
[0070] In this embodiment, the resolution is required to be ≥1920×1080;
[0071] High resolution is used to capture subtle anatomical structures, including features such as the mandibular edge, nasal tip, and soft tissue nasal root. Low-resolution images will lead to the loss of key details and reduce the reliability of generating virtual skull lateral films. The resolution is limited to ≥1920×1080 to ensure that the input image is high-resolution. The encoder can extract the edges of the maxillofacial soft tissues in the input image and reduce the upsampling error in the decoding stage.
[0072] Specifically, the encoder of the model can separate and enhance the color features related to the bone structure through the convolutional layer.
[0073] The input image has specific requirements for the patient's posture attributes: the patient maintains a natural head position, and the line connecting the tragus and infraorbital point is parallel to the ground.
[0074] The line connecting the tragus and the infraorbital point is used as a reference plane, and the reference plane is aligned to ensure spatial consistency between the model input and the training data.
[0075] Among them, the line connecting the tragus and the infraorbital point is the reference plane for clinical X-ray shooting. Aligning this plane can ensure the spatial consistency between the model input and the training data. The analysis and diagnosis of lateral skull radiographs rely on standardized shooting angles. If the patient's head is tilted, it will cause the jaw projection to deform, and the coordinate deviation of the landmark points in the generated virtual lateral skull radiograph will increase.
[0076] By requiring the patient's posture, the head angle of the input image is unified. The unified head angle enables the model to effectively learn the projective geometric relationship from two-dimensional photos to virtual skull lateral films. The cross-modal attention mechanism can more accurately associate visible light features with bone structure.
[0077] By limiting the format, resolution and patient posture requirements of the input image, it helps the encoder to extract features, thereby improving the accuracy of the virtual skull lateral film. Figure 4 The virtual lateral skull film shown is highly similar to the real lateral skull film obtained through X-ray imaging, ensuring the accuracy of the initial screening.
[0078] like Figure 2 As shown, the cross-modal image generation model in this embodiment adopts an improved U-Net++ structure and embeds a cross-modal attention mechanism.
[0079] The cross-modal image generation model includes:
[0080] Attention module: used to calculate attention weights;
[0081] In this embodiment, an attention module is added between each level of U-Net++. This module selectively emphasizes important features and suppresses unnecessary information by calculating the attention weights of feature maps.
[0082] The expression is as follows:
[0083]
[0084] Among them, Q is the feature vector from the visible light image, K / V is the key-value pair of the X-ray image feature library, which contains the anatomical medical prior knowledge of bone structure, is the scaling factor that keeps the gradient stable by controlling the variance.
[0085] The attention module includes:
[0086] Multi-level feature alignment unit: used to align spatial positions layer by layer;
[0087] Dynamic weight allocation unit: used to dynamically allocate attention weights based on space and channels;
[0088] The dynamic weight allocation unit includes a spatial dynamic allocation submodule and a channel dynamic allocation submodule;
[0089] The spatial dynamic allocation submodule specifically includes: locating the area corresponding to the bony landmark in the input image and assigning weights;
[0090] The spatial dynamic allocation submodule locates the areas corresponding to bony landmarks in the input photo, such as the mandibular angle, through the attention score matrix and assigns higher weights.
[0091] The channel dynamic allocation submodule specifically includes: assigning weights to different channel features;
[0092] The channel dynamic allocation submodule includes: strengthening channels related to bone density in the feature library.
[0093] The feature library refers to the X-ray image feature library, that is, the feature set extracted from the X-ray image during the model training process.
[0094] The channel dynamic allocation submodule selects different channel features, such as edges and textures, for example, strengthening the channels related to bone density in the X-ray image feature library.
[0095] Decoder guided generation unit: used to further modify the distribution of attention weights according to the loss function.
[0096] In this embodiment, a multi-level feature alignment unit is embedded in the cross-modal attention module at the skip connection of U-Net++ to align the spatial positions of visible light features and X-ray features layer by layer, such as the anatomical landmark corresponding to the line connecting the tragus and the infraorbital point.
[0097] The decoder-guided generation unit specifically refers to: in the final output layer, the weight distribution is further corrected through the anatomically constrained loss function, such as the L1 loss of the mandibular edge, to ensure that the generated image conforms to the clinical prior.
[0098] Adaptive illumination normalization module: used to eliminate the influence of ambient light differences in the input image;
[0099] In this embodiment, the adaptive illumination normalization module is directly embedded in the improved U-Net++ architecture and is one of the core components of the generative model. This layer is located at the front end of the U-Net++ encoder and serves as the first processing module after the input image enters the network.
[0100] The adaptive illumination normalization module includes:
[0101] Lighting estimation unit: used to learn the lighting distribution parameters of the input image;
[0102] The illumination estimation unit uses a lightweight neural network to learn the illumination distribution parameters of the input image, which are expressed as follows:
[0103] μ,σ=f estimator (X)
[0104] Among them, μ describes the brightness level of the area around each pixel; σ describes the magnitude of the illumination change around each pixel.
[0105] Through a learnable mapping function f estimator , directly generates two key lighting parameters μ and σ from the input image X.
[0106] Normalization calculation unit: used to adjust the normalized intensity according to the illumination distribution parameters.
[0107] The normalization calculation unit dynamically estimates the mean and standard deviation of each pixel, adjusts the normalized intensity according to the lighting conditions of different areas, preserves details, solves problems such as uneven exposure, noise and low contrast, and provides stable input for automated measurement and diagnosis.
[0108] Encoder: used to extract input image features and obtain feature maps based on attention weights;
[0109] Since the encoder reduces the image resolution when extracting image features, high-resolution input images are more conducive to feature extraction.
[0110] In this embodiment, the encoder includes a convolutional layer and a downsampling layer.
[0111] In the encoder, features of the input image are gradually extracted through a series of convolutional layers. Each convolution is usually followed by a ReLU activation function.
[0112] The core operation formula of the convolutional layer is as follows:
[0113]
[0114] Among them, I is the input feature map, K is the convolution kernel, (i, j) is the output coordinate, and (m, n) is the index of the convolution kernel.
[0115] Among them, feature extraction mainly includes edge feature detection and spatial relationship modeling. Edge feature detection is the primary convolution layer, and spatial relationship modeling is feature extraction based on the attention weights calculated by the attention module.
[0116] The ReLU function is used after each convolutional layer, and the expression is as follows:
[0117]
[0118] Among them, I is the input feature map, K is the convolution kernel, (i, j) is the output coordinate, and (m, n) is the index of the convolution kernel.
[0119] An element-by-element nonlinear transformation is applied to each position (i, j) of the convolution result, which enables the network to learn complex feature relationships, such as the mapping of soft tissue contours to bony structures, by setting negative values to zero and retaining positive values.
[0120] Furthermore, the soft tissue contour needs to be extracted through the color information of the input image; and the color information needs to be limited to the input image in RGB format and extracted through the RGB channel; therefore, it can be concluded that the RGB format is conducive to extracting such structural features and assisting in inferring deep bony relationships.
[0121] The downsampling layer is to use a pooling layer after each convolution block, such as maximum pooling, to downsample to reduce the spatial dimension and improve the abstract level of the features.
[0122] The pooling layer is a nonlinear downsampling operation in convolutional neural networks, which is used to gradually reduce the spatial dimension of the feature map while retaining key feature information. The expression is as follows:
[0123]
[0124] Among them, Output(i,j) is the value of the output feature map at (i,j), (p,q) is the coordinate offset of all positions in the pooling window, Input(i+p,j+q) is the value of the input feature map at (i+
[0125] The value at p,j+q).
[0126] Specifically, the expression process of the pooling layer is as follows: 1. Sliding window traversal: the pooling window slides on the input feature map according to the step size; 2. Local maximum extraction: the maximum value in each window is taken as the corresponding position value of the output feature map; 3. Feature map compression: the combination of maximum pooling with cross-modal attention mechanism and anatomical constraints supports high-precision mapping from ordinary photos to virtual skull lateral films, while reducing computational complexity.
[0127] Decoder: used to restore the feature map to the size of the input image and generate a virtual cephalogram;
[0128] In this embodiment, the decoder uses an upsampling layer to gradually restore the size of the feature map to the size of the input image. The upsampling layer usually uses a deconvolution method or an interpolation method.
[0129] The deconvolution method is expressed as follows:
[0130]
[0131] Among them, y(i,j) is the pixel value of the output feature map at the spatial position (i,j), corresponding to the output channel c out .
[0132] By reverse calculation and adjusting the step size, the position where sampling is required in the input feature map is determined, and the values of all channels of the input feature map within the kernel window are weighted and summed to finally output a high-resolution anatomical structure.
[0133] The interpolation expression is as follows:
[0134]
[0135] Among them, y(u,v) is the value of the output feature map at position (u,v), and x(i,j) is the value of the input feature map at position (i,j).
[0136] Interpolation can be used to infer values at unknown locations from known data points to generate a virtual cephalogram.
[0137] When generating virtual lateral skull radiographs, combining deconvolution (transposed convolution) and interpolation methods can balance generation efficiency and accuracy of anatomical details.
[0138] In this embodiment, the decoder also includes a splicing operation: in the decoding stage, the feature map from the encoder is combined into the decoded feature map through a splicing operation to ensure the transmission of high-resolution information.
[0139] Skip connection module: used to connect each layer of the encoder to the corresponding layer of the decoder;
[0140] The skip connection module allows the model to reuse low-level features in the decoding stage by directly connecting from each layer of the encoder to the corresponding layer of the decoder, thereby enhancing the ability to recover details.
[0141] Loss function: used to determine the accuracy of the virtual skull lateral radiograph and update the model;
[0142] Anatomical features are embedded in the loss function.
[0143] X-ray anatomical features, such as the sharpness of the mandibular edge and the contrast of the nasal spine, are introduced into the loss function, and the bone structure generation is enhanced through the weight matrix.
[0144] Landmark extraction model: used to extract landmark points from virtual lateral skull radiographs, obtain bony landmark points, and generate diagnostic suggestions based on the bony landmark points.
[0145] The landmark point extraction model includes:
[0146] In this embodiment, the training process of the landmark point extraction model includes the following steps: 1. Collecting an image data set containing lateral views, and annotating the lateral views, including the coordinates of the bony feature points; 2. Dividing the data into a training set, a validation set, and a test set. The commonly used format is image + annotation file, and each annotation file includes the coordinate information of the bony landmark points in the image; 3. Inputting the training set into the landmark point extraction model, and updating the parameters of the landmark point extraction model according to the validation set.
[0147] Convolutional neural network: used to detect key points of virtual skull lateral radiographs and obtain bony landmarks;
[0148] Graph convolutional network: Modeling the spatial relationship of bony landmarks and calculating clinical parameters;
[0149] Clinical parameters such as ANB angle, Wits value, etc.
[0150] Diagnostic suggestion module: Generates diagnostic suggestions based on clinical parameters.
[0151] Among them, clinical parameters are combined with Angle's classification rules and skeletal grade thresholds to output diagnostic recommendations, such as "skeletal class II, even angles."
[0152] like Figure 3 As shown, a deep learning-based non-radiation initial screening method for malocclusion is implemented by the above-mentioned non-radiation initial screening system for malocclusion, comprising the following steps:
[0153] S1. Get input image;
[0154] S2, the cross-modal image generation model generates a virtual lateral skull radiograph based on the input image;
[0155] The cross-modal image generation model generates a virtual cephalogram according to the input image in step S2, comprising the following steps:
[0156] S21, the adaptive illumination module eliminates the ambient light difference of the input image;
[0157] S22, the attention module calculates the attention weight;
[0158] S23, the encoder extracts the features of the input image according to the attention weight and obtains a feature map;
[0159] S24. The decoder restores the feature map to the size of the input image and generates a virtual lateral skull film.
[0160] The virtual cephalogram obtained by the present invention is compared with the real X-ray cephalogram, and the structural similarity index (SSIM) of the two is ≥ 0.91. When the test set N = 500, the coordinate error of the key landmarks is ≤ 0.5 mm.
[0161] S3, the landmark extraction model extracts bony landmarks based on the virtual skull lateral radiograph and generates diagnostic suggestions based on the bony landmarks;
[0162] The S3, extracting the landmark point model, extracting the bony landmark points based on the virtual lateral skull film, and generating a diagnosis suggestion based on the bony landmark points includes the following steps:
[0163] S31, convolutional neural network extracts bony landmarks based on virtual lateral skull radiographs;
[0164] like Figure 5 The figure shows a schematic diagram of bony landmark points obtained by modeling the spatial relationship of bony landmark points in this embodiment.
[0165] The bony landmarks obtained by the present invention can generate a thermal map to mark the contribution of key areas to the diagnostic results; the Kappa consistency coefficient of the initial screening diagnosis of bony landmarks is ≥0.85 with the gold standard, that is, the real X-ray lateral skull film; the single case analysis takes 1.2±0.3 minutes, while the traditional method takes 20 to 30 minutes.
[0166] S32, graph convolutional network models the spatial relationship of bony landmarks and calculates clinical parameters;
[0167] S33. The diagnosis suggestion module generates diagnosis suggestions based on clinical parameters.
[0168] S4. Upload the virtual skull lateral film and bony landmarks, and update the cross-modal image generation model and landmark extraction model.
[0169] In this embodiment, step S4 adopts a federated learning update model to protect patient privacy, specifically including the following steps: 1. Receiving encrypted local data; 2. Secure aggregation; 3. Updating the global model, updating progressively, and retaining historical versions; 4. Sending to the client.
[0170] Among them, updating the global model is specifically as follows:
[0171] By backpropagating the edge sharpness loss, the deconvolution layer weights of the decoder are updated, and the expression is as follows:
[0172]
[0173] in, is the difference between the generated image and the real image in terms of edge information, N is the total number of pixels in a single image, G(x i ) is the generated virtual skull lateral film, Y real For the real X-ray lateral skull film, is the edge detection factor;
[0174] Updated parameters include: the convolution kernel weights of the decoder's high-resolution reconstruction layer and the channel attention vector of the cross-modal attention module;
[0175] The node embedding matrix of the graph convolutional network is updated by the landmark point coordinate regression loss, which is expressed as follows:
[0176]
[0177] in, is the difference between the predicted landmark point and the true landmark point, N is the total number of samples, y i is the true value of the i-th sample, is the predicted value of the i-th sample;
[0178] The updated parameters include: the weight of the graph convolutional network adjacency matrix and the feature projection weight of the last layer of ResNet-50.
[0179] By updating the model parameters, the training set can be continuously increased, the model can be optimized, and the accuracy of the non-radiation screening system for malocclusion can be further improved.
[0180] By comparing the present invention with traditional X-ray diagnosis, the following technical comparison table can be obtained:
[0181] Traditional X-ray diagnosis The present invention Radiation risks 2~200μSv / time 0μSv / time Bone precision Gold standard (error 0mm) Error ≤ 0.5mm Equipment costs 100,000 to 1 million yuan <5000 yuan (ordinary camera or smartphone) Diagnosis time 20-30 minutes / case ≤2 minutes / case
[0182] The present invention generates high-precision lateral cephalograms from ordinary photos, breaking through the bottleneck of non-radiation malocclusion diagnosis and achieving low-cost and high-accuracy preliminary screening for malocclusion; by embedding anatomical features in the loss function and introducing an attention module in the cross-modal image generation model, the accuracy and medical credibility of the generated images are improved; by extracting a landmark point model to identify a virtual lateral cephalogram, clinical parameters and diagnostic basis can be further obtained without the need for physician judgment, with a fast recognition speed, further reducing labor and time costs; by synchronously outputting a virtual lateral cephalogram, quantified clinical parameters and diagnostic basis through a cross-modal image generation model and a landmark point extraction model, the judgment basis can be increased during manual screening, thereby enhancing reliability.
Claims
1. A deep learning-based non-radiation initial screening system for malocclusion, characterized by: include: Cross-modal image generation model: used to extract features of the input image, calculate the attention weights of the features, obtain feature maps based on the attention weights, and generate virtual cephalograms based on the feature maps; Landmark extraction model: used to extract landmarks from virtual lateral skull radiographs, obtain bony landmarks, and generate diagnostic recommendations based on the bony landmarks; The attributes of the input image include: RGB format, high resolution, and patient posture; The cross-modal image generation model includes: an attention module, an adaptive illumination normalization module, an encoder, a decoder, a skip connection module and a loss function.
2. The non-radiation primary screening system for malocclusion according to claim 1, characterized in that: The RGB format attributes of the input image can be used to obtain the color characteristics of skin and mucous membranes; The high resolution of the input image is used to capture subtle anatomical structures; The subtle anatomical structures include the mandibular margin, nasal tip, and soft tissue nasion; The input image has specific requirements for the patient's posture attributes: the patient maintains a natural head position, and the line connecting the tragus and infraorbital point is parallel to the ground; The line connecting the tragus and the infraorbital point is used as a reference plane, and the reference plane is aligned to ensure spatial consistency between the model input and the training data.
3. The non-radiation malocclusion screening system according to claim 1, characterized in that: The attention module is used to calculate the attention weight; The adaptive illumination normalization module is used to eliminate the influence of ambient light differences on the input image; The encoder is used to extract input image features and obtain a feature map based on attention weights; The decoder is used to restore the feature map to the size of the input image and generate a virtual cephalogram; The skip connection module is used to connect each layer of the encoder to the corresponding layer of the decoder; The loss function is used to determine the accuracy of the virtual skull lateral radiograph and update the model; Anatomical features are embedded in the loss function.
4. The non-radiation primary screening system for malocclusion according to claim 3, characterized in that: The attention module includes: Multi-level feature alignment unit: used to align spatial positions layer by layer; Dynamic weight allocation unit: used to dynamically allocate attention weights based on space and channels; Decoder guided generation unit: used to further modify the distribution of attention weights according to the loss function.
5. The non-radiation primary screening system for malocclusion according to claim 4, characterized in that: The dynamic weight allocation unit includes a spatial dynamic allocation submodule and a channel dynamic allocation submodule; The spatial dynamic allocation submodule specifically includes: locating the area corresponding to the bony landmark in the input image and assigning weights; The channel dynamic allocation submodule specifically includes: assigning weights to different channel features; The channel dynamic allocation submodule includes: strengthening channels related to bone density in the feature library.
6. The non-radiation primary screening system for malocclusion according to claim 3, characterized in that: The adaptive illumination normalization module includes: Lighting estimation unit: used to learn the lighting distribution parameters of the input image; Normalization calculation unit: used to adjust the normalized intensity according to the illumination distribution parameters.
7. The non-radiation primary screening system for malocclusion according to claim 1, characterized in that: The landmark point extraction model includes: Convolutional neural network: used to detect key points of virtual skull lateral radiographs and obtain bony landmarks; Graph convolutional network: Modeling the spatial relationship of bony landmarks and calculating clinical parameters; Diagnostic suggestion module: Generates diagnostic suggestions based on clinical parameters.
8. A deep learning-based non-radiation screening method for malocclusion, characterized by: The method is implemented by the non-radiation primary screening system for malocclusion according to any one of claims 1 to 7, comprising the following steps: S1. Get input image; S2, the cross-modal image generation model generates a virtual lateral skull radiograph based on the input image; S3, the landmark extraction model extracts bony landmarks based on the virtual lateral skull radiograph and generates diagnostic suggestions based on the bony landmarks; S4. Upload the virtual skull lateral film and bony landmarks, and update the cross-modal image generation model and landmark extraction model.
9. The non-radiation primary screening method for malocclusion according to claim 8, characterized in that: The cross-modal image generation model generates a virtual cephalogram according to the input image in step S2, comprising the following steps: S21, the adaptive illumination module eliminates the ambient light difference of the input image; S22, the attention module calculates the attention weight; S23, the encoder extracts the features of the input image according to the attention weight and obtains a feature map; S24. The decoder restores the feature map to the size of the input image and generates a virtual lateral skull film.
10. The non-radiation primary screening method for malocclusion according to claim 8, characterized in that: The S3, extracting the landmark point model, extracting the bony landmark points based on the virtual lateral skull film, and generating a diagnosis suggestion based on the bony landmark points includes the following steps: S31, convolutional neural network extracts bony landmarks based on virtual lateral skull radiographs; S32, graph convolutional network models the spatial relationship of bony landmarks and calculates clinical parameters; S33. The diagnosis suggestion module generates diagnosis suggestions based on clinical parameters.