Automatic Identification Method and Device for Key Landmark Points in Lateral Cephalogram
Through the integration of the U-shaped structure deep learning model and the dependency relationship between key landmark points, the problem of time-consuming and laborious identification of head lateral tablet mark points and high resource demand is solved, and the rapid and accurate multi-class landmark point recognition is achieved to meet the needs of common measurement methods.
Patent Information
- Application Number
- CN202311530112.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-16
- Publication Date
- 2025-07-08
- Estimated Expiration
- 2043-11-16
AI Technical Summary
The existing method for identifying key marker points of skull lateral films relies on manual annotation, which is time-consuming and error-prone. The existing deep learning models are slow to calculate and have high hardware resource requirements, so it is impossible to fully identify key marker points of cephalogram measurement, bone age and airway.
The U-shaped structure deep learning model is adopted to obtain the skull lateral film, generate a heat map and solve the key marker points, and modify and integrate the dependencies of the key marker points to identify the key marker points.
It realizes fast and accurate identification of key marking points, reduces the hardware resource requirements, and can identify key marking points of cephalogram measurement, bone age and airway to meet the needs of common measurement methods.
Smart Images

Figure CN117558030B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and particularly to an automatic recognition method for key landmark points of a lateral cephalogram and an automatic recognition device for key landmark points of a lateral cephalogram. Background Art
[0002] The shooting of a lateral cephalogram by X-ray is a routine examination before orthodontics and is also an important step for orthodontic diagnosis and maxillofacial surgical treatment planning. However, the determination of key landmark points of a lateral cephalogram highly depends on manual annotation by doctors. The whole process is time-consuming and laborious and is extremely prone to deviation. Therefore, it has great practical significance to study a stable and reliable automatic recognition method for key landmark points of a lateral cephalogram.
[0003] Nowadays, with the booming development of deep learning technology, it has in-depth applications in all walks of life. It digs out the commonalities between things through mathematical methods such as probability theory and mathematical statistics and realizes the automation of operations based on this. At present, there are mainly two problems in the method of identifying landmark points of a lateral cephalogram through deep learning: one is that there are few key cephalometric landmark points, which cannot meet the needs of some common measurement methods. The other is that most of the deep learning models adopted directly identify key landmark points, and the convolutional attention mechanism in the inference process, Taylor expansion in decoding coordinates, etc. are too heavy, making the method computationally slow and requiring high hardware resources. Summary of the Invention
[0004] In order to solve the above technical problems, the present invention provides an automatic recognition method and device for key landmark points of a lateral cephalogram, which can avoid directly identifying key landmark points in a lateral cephalogram by using a complex deep learning model, has a faster calculation speed and high robustness, requires lower hardware resources, and can accurately and comprehensively identify key landmark points in a lateral cephalogram to meet the needs of most common measurement methods.
[0005] The technical solution adopted by the present invention is as follows:
[0006] An automatic recognition method for key landmark points of a lateral cephalogram includes the following steps: S1, obtaining a lateral cephalogram to be recognized; S2, inputting the lateral cephalogram to be recognized into a key landmark point recognition model to obtain a heat map of each first key landmark point; S3, resolving the heat map to obtain the first key landmark points of the lateral cephalogram; S4, calculating to obtain second key landmark points according to the dependency relationship of the key landmark points, and correcting some of the first key landmark points obtained by resolution to obtain third key landmark points, and integrating to obtain final key landmark points, wherein the set of the uncorrected first key landmark points, the second key landmark points, and the third key landmark points is the final key landmark points.
[0007] In addition, the automatic recognition method for key landmark points on a lateral cephalogram proposed above according to the present invention may also have other additional technical features:
[0008] According to an embodiment of the present invention, the key landmark points on the lateral cephalogram include: cephalometric key landmark points, skeletal age key landmark points, and airway key landmark points.
[0009] According to an embodiment of the present invention, before step S2, it further includes: obtaining a sample lateral cephalogram and a corresponding standard result file; preprocessing the sample lateral cephalogram, determining the positions of the first key landmark points on the sample lateral cephalogram according to the positions of the key landmark points in the standard result file to obtain a sample data set; dividing the sample data set into a training set, a validation set, and a test set; performing image enhancement on the sample lateral cephalograms in the training set; converting the position data of the first key landmark points on the sample lateral cephalograms in the sample data set into heatmaps; inputting the heatmaps into a deep learning model network for training to obtain the key landmark point recognition model.
[0010] Specifically, pixel-level transformation or spatial-level transformation is used to perform image enhancement on the sample lateral cephalograms in the training set.
[0011] Specifically, the deep learning model network is a U-shaped structure deep learning model network, and the loss function of the U-shaped structure deep learning model is binary cross-entropy, and its calculation formula is:
[0012]
[0013] Among them, represents binary cross-entropy, represents the output heatmap calculated by the model, o represents the heatmap converted based on the first key landmark points, N represents the number of samples, H is the length of the output heatmap, W is the width of the output heatmap, represents the pixel value of the i-th row and j-th column of the heatmap converted from the k-th first key landmark point.
[0014] Specifically, the second key landmark points are calculated according to the positional relationship of the key landmark points or the line segment ratio relationship between the line segments formed by the key landmark points, and the calculated partial first key landmark points are corrected according to the positional relationship of the key landmark points to obtain the third key landmark points.
[0015] In addition, to achieve the above object, the present invention also proposes an automatic recognition device for key landmark points on a lateral cephalogram.
[0016] An automatic recognition device for key landmark points of a lateral cephalogram, comprising: a first acquisition module for acquiring a lateral cephalogram to be recognized; an output module for inputting the lateral cephalogram to be recognized into a key landmark point recognition model to obtain a heat map of each key landmark point; a calculation module for calculating the heat map to obtain the initial key landmark points of the lateral cephalogram; an integration module for calculating to obtain second key landmark points according to the dependency relationship of the key landmark points, and correcting some of the first key landmark points obtained by calculation to obtain third key landmark points, and finally obtaining the final key landmark points through integration, wherein the set of the uncorrected first key landmark points, the second key landmark points and the third key landmark points is the final key landmark points.
[0017] In addition, the automatic recognition device for key landmark points of a lateral cephalogram proposed above according to the present invention may also have other additional technical features:
[0018] According to an embodiment of the present invention, the key landmark points of the lateral cephalogram include: cephalometric key landmark points, skeletal age key landmark points and airway key landmark points.
[0019] According to an embodiment of the present invention, the device further comprises: a second acquisition module for acquiring a sample lateral cephalogram and a corresponding standard result file; a preprocessing module for preprocessing the sample lateral cephalogram and determining the positions of the key landmark points of the sample lateral cephalogram according to the positions of the key landmark points in the standard result file to obtain a sample data set; a data division module for dividing the sample data set into a training set, a validation set and a test set; an image enhancement module for performing image enhancement on the sample lateral cephalograms in the training set; a conversion module for converting the position data of the key landmark points of the sample lateral cephalograms in the sample data set into heat maps; a training module for inputting the heat maps into a deep learning model network for training to obtain the key landmark point recognition model.
[0020] Specifically, the image enhancement module is used to perform image enhancement on the sample lateral cephalograms in the training set by using pixel-level transformation or space-level transformation.
[0021] Specifically, the deep learning model network is a U-shaped structure deep learning model network, and the loss function of the U-shaped structure deep learning model is binary cross entropy, and its calculation formula is:
[0022]
[0023] wherein, denotes binary cross-entropy, represents the output heatmap calculated by the model, o represents the heatmap converted based on the first key landmark points, N represents the number of samples, H is the length of the output heatmap, and W is the width of the output heatmap, represents the pixel value at the i-th row and j-th column of the heatmap converted from the k-th first key landmark point.
[0024] Specifically, the integration module is used to calculate the second key landmark points according to the positional relationship of the key landmark points or the line segment ratio relationship between the line segments formed by the key landmark points, and correct some of the first key landmark points obtained by solution according to the positional relationship of the key landmark points to obtain the third key landmark points.
[0025] Advantages of the present invention:
[0026] The method for automatically identifying key landmark points of a lateral cephalogram according to an embodiment of the present invention inputs the lateral cephalogram to be identified into a key landmark point recognition model to obtain the heatmap of each first key landmark point, and calculates the first key landmark points from the heatmap, avoiding directly identifying the key landmark points in the lateral cephalogram using a complex deep learning model, with faster calculation speed, high robustness, and lower requirements for hardware resources; calculating and correcting according to the dependency relationship of the key landmark points, and integrating to obtain the final key landmark points, which can effectively utilize the mutual dependence and constraints between the key landmark points, accurately and comprehensively identify the key landmark points in the lateral cephalogram, and meet the needs of most common measurement methods. Brief Description of the Drawings
[0027] Figure 1 is a flowchart of the method for automatically identifying key landmark points of a lateral cephalogram according to an embodiment of the present invention;
[0028] Figure 2 is a distribution schematic diagram of the key landmark points of a lateral cephalogram according to an embodiment of the present invention;
[0029] Figure 3 is a dependency schematic diagram of key landmark points with a dependency relationship according to a specific embodiment of the present invention;
[0030] Figure 4 is a block schematic diagram of the device for automatically identifying key landmark points of a lateral cephalogram according to an embodiment of the present invention. Detailed Embodiments
[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0032] As Figure 1 shown, the method for automatically identifying key landmark points of a lateral cephalogram in an embodiment of the present invention includes the following steps:
[0033] S1. Obtain the lateral cephalogram to be identified.
[0034] S2. Input the lateral cephalogram to be identified into the key landmark point recognition model to obtain the heat map of each first key landmark point.
[0035] The key landmark points of the airway can be used to judge the development status of the airway and diagnose diseases such as airway stenosis; the key landmark points of bone age can not only determine the biological age of children, but also understand the growth and development potential and the trend of sexual maturity of children through bone age at an early stage. Bone age can also be used to predict the adult height of children, and the determination of bone age is also very helpful for the diagnosis of some pediatric endocrine diseases and has great guiding significance for the treatment of some patients with short stature.
[0036] However, most of the current methods for identifying landmark points of lateral cephalograms through deep learning only include cephalometric key landmark points, without bone age key landmark points and airway key landmark points, and cannot diagnose the degree of airway stenosis and determine bone age.
[0037] As Figure 2 shown, in an embodiment of the present invention, the key landmark points of the lateral cephalogram may include: cephalometric key landmark points, bone age key landmark points, and airway key landmark points. In the figure, the circular points are cephalometric key landmark points, the pentagram points are bone age key landmark points, and the triangular points are airway key landmark points. Identifying these key landmark points is beneficial to the diagnosis of the degree of airway stenosis and the determination of bone age, meeting the needs of common measurement methods; taking the two points with a fixed distance on the positioning ruler as cephalometric key landmark points is beneficial to determining the scale of the lateral cephalogram and improving the accuracy of recognition.
[0038] In an embodiment of the present invention, before step S2, it may further include:
[0039] Sa. Obtain the sample lateral cephalogram and the corresponding standard result file.
[0040] In a specific embodiment of the present invention, a sufficient amount of lateral cephalometric X-ray samples are collected, including some low-quality lateral cephalometric samples, such as unclear soft tissue imaging; non-overlapping of the mechanical ear points and mandibular angles on both sides; the head not being perpendicular to the ground; anterior open bite; directly taking a "paper lateral cephalometric film" with a mobile phone, resulting in paper warping; pictures intercepted using a screenshot tool, etc.; and obtaining a standard result file corresponding to the lateral cephalometric sample by having 4 experts perform detailed annotation on a total of 78 key landmark points required for analysis methods such as cephalometric analysis, bone age analysis, and airway analysis. Each landmark point is as Figure 2 shown.
[0041] Sb, preprocess the lateral cephalometric sample, and determine the position of the first key landmark point of the lateral cephalometric sample according to the position of the key landmark points in the standard result file to obtain a sample dataset.
[0042] In an embodiment of the present invention, step Sb may specifically include:
[0043] Delete the lateral cephalometric samples with extremely poor image quality. Since too much information is lost in these lateral cephalometric samples, deleting these samples is beneficial to improving the effectiveness of modeling.
[0044] Adopt the 3σ principle to delete the lateral cephalometric samples with excessive divergence in the annotation results of multiple experts.
[0045] To reduce the influence of the subjectivity of each expert, the average value of the annotation results of multiple experts is determined as the actual position of the key landmark points, that is, the position of the first key landmark point of the lateral cephalometric sample.
[0046] Sc, divide the sample dataset into a training set, a validation set, and a test set.
[0047] In an embodiment of the present invention, the training set, the validation set, and the test set are divided in a ratio of 0.8:0.15:0.05. At the same time, in order to maintain the proportion of lateral cephalometric samples in the mixed dentition period among all lateral cephalometric samples, stratified non-replacement sampling is adopted.
[0048] Sd, perform image enhancement on the lateral cephalometric samples in the training set.
[0049] It should be noted that in this embodiment, "performing image enhancement on the lateral cephalometric samples in the training set" is to simulate various low-quality images to cope with the situation that the actually obtained lateral cephalometric films to be recognized may be low-quality images, and improve the effectiveness of modeling.
[0050] In an embodiment of the present invention, pixel-level transformation or spatial-level transformation can be used to perform image enhancement on the lateral cephalometric samples in the training set.
[0051] In another embodiment of the present invention, pixel-level transformation and spatial-level transformation can also be simultaneously used to enhance the training set sample lateral cephalometric radiographs, so as to more comprehensively simulate various low-quality images.
[0052] Specifically, the pixel-level transformation includes: Gaussian blur, where the size of the convolution kernel of the Gaussian blur can be 7x7; random brightness transformation, and the range of the random brightness transformation can be [-0.2, 0.2]; gamma transformation, and the set of gamma values can be {0.2, 0.4, 0.67, 1, 1.5, 2.5, 5.0}, where the probability of selecting a gamma value of 1 is 50%, and the gamma transformation can better simulate the effect of picture degradation; histogram equalization. When performing pixel-level transformation, the probabilities of selecting these several pixel-level image enhancement methods are 0.2, 0.1, 0.6, and 0.1 respectively.
[0053] The spatial-level transformation includes random edge cropping, and the cropping ratio is [0, 0.2]; random edge filling, and the filling ratio is [0, 0.2], and the filling value is 255 (u8); random left-right flipping; affine transformation, where the rotation angle can be [-30°, 30°], and the scaling ratio can be [-0.5, 2]; grid transformation, and the size of the grid can be 5x5, and the distortion ratio is [-0.03, 0.03]; elastic transformation, where the α value of the elastic transformation is 200 and the β value is 20. The probabilities of selecting these several spatial-level image enhancement methods are 0.2, 0.2, 0.2, 0.2, and 0.4 respectively.
[0054] Se, convert the position data of the first key landmark point of the sample lateral cephalometric radiograph in the sample data set into a heat map.
[0055] Specifically, assume that the coordinate of a certain first key landmark point after image enhancement is P A =(x A , y A ), and then generate a heat map centered on P A through the following two-dimensional Gaussian distribution probability density function:
[0056]
[0057] where (x, y) represents the position of the pixel in the heat map, and σ represents the standard deviation of this distribution. In this embodiment, the value of σ is 5.
[0058] Therefore, each first key landmark point of each sample lateral cephalometric radiograph has a corresponding heat map.
[0059] Sf, input the heat map into the deep learning model network for training to obtain a key landmark point recognition model.
[0060] In an embodiment of the present invention, the deep learning model network can be a U-shaped deep learning model network. The overall model is a U-shaped structure. The left half of the U is the feature extractor of the model, and the right half is the feature interpreter. The feature extractor and the feature interpreter not only transfer feature information through the lowest layer, but also make long connections between the outputs of each layer of the feature extractor and each layer of the feature interpreter. This structural design not only avoids the problem of gradient disappearance caused by overly deep networks, but also fuses feature maps of different scales together, making up for the information loss during the downsampling process of the feature extractor.
[0061] Specifically, the feature extractor has a total of 5 layers. Each layer receives the output of the previous layer as the input of this layer. The input of the first layer is the lateral radiograph image. From top to bottom, the shallow layer processes the high-resolution feature map of the image, which contains more image detail texture information, while the deep layer processes the low-resolution feature map of the image, which contains more image semantic information. Each layer consists of a downsampling module and a convolution module. The downsampling is responsible for reducing the size of the image feature map, transitioning the feature map from high resolution to low resolution. Its internal structure is a stack of Conv+BN+LeakyReLU sequences, where the convolution kernel size of Conv is 3, the stride is 2, and the size of the output feature map is half of the size of its input feature. The convolution module is responsible for extracting the feature information of this feature map. Its internal structure is similar to the sequence stacking structure of the downsampling layer, but it is thickened by one layer, that is, (Conv+BN+LeakyReLU)x2, where the convolution kernel size of the Conv layer is 3, the stride is 1, and the padding is 1, ensuring the same size of the input and output. In addition, the number of channels in each layer increases sequentially to store more refined picture feature information. From top to bottom, the number of channels in each layer is 64, 128, 256, 512, and 1024.
[0062] Correspondingly, the feature interpreter also has five layers (the bottom layer is shared by both). Different layers process feature maps of different sizes of the image. Specifically, the input of these layers is composed of the output of the previous layer and the output of the corresponding depth layer of the feature extractor. Each layer consists of an upsampling module and a convolution module. Among them, the convolution module is similar to the convolution module in the feature extractor, and the upsampling module is implemented by nearest neighbor interpolation with a scaling factor of 2. Therefore, the size of the output feature map is twice the size of the input feature map. Similarly, the number of channels in each layer of the feature interpreter also changes. From bottom to top, they are 1024, 512, 256, 128, and 64, representing the decoding of feature information. Finally, the last layer is subjected to convolution calculation with a convolution kernel of 1 to generate the heat map of the first key landmark.
[0063] The loss function of the U-shaped deep learning model is binary cross-entropy, and its calculation formula is:
[0064]
[0065] Among them, represents binary cross-entropy, represents the output heatmap calculated by the model, o represents the heatmap converted based on the first key landmark, N represents the number of samples, H is the length of the output heatmap, and W is the width of the output heatmap. represents the pixel value at the i-th row and j-th column of the heatmap converted from the k-th first key landmark.
[0066] Meanwhile, during the model training process, it is also necessary to make corresponding adjustments to the learning rate. The learning rate is an important hyperparameter when training a neural network, which determines the magnitude of the weight update in each iteration. The impact of the learning rate on the performance of the neural network is very significant. Too high or too low a learning rate may lead to unstable or inefficient training. At the initial stage of model training, the model has a large optimization space, and a larger learning rate can accelerate the convergence of the model; however, in the later stage of training, the model optimization space becomes smaller, and a large learning rate may cause the loss function to oscillate around the global minimum and fail to converge.
[0067] In an embodiment of the present invention, a custom method is adopted such that the learning rate automatically decreases as the number of training epochs increases, and its formula is as follows:
[0068]
[0069] where lr epoch represents the learning rate of the current training epoch, lr0 represents the initial learning rate, and epoch represents the current training. In the present invention, lr0 = 0.01.
[0070] In an embodiment of the present invention, a validation set can be selected to verify whether the model training converges, and a test set can also be selected to input the converged model and calculate the accuracy of the first key landmark by solving the output heatmap.
[0071] S3. Solve the heatmap to obtain the first key landmark of the lateral cephalogram.
[0072] It should be noted that by inputting the lateral cephalogram to be recognized into the key landmark recognition model to obtain the heatmap of each first key landmark and then solving the heatmap to obtain the first key landmark of the lateral cephalogram, compared with directly recognizing the key landmarks using a deep learning model, outputting the heatmap of the key landmarks consumes less resources, has a faster calculation speed and high robustness, and has lower requirements for hardware resources.
[0073] S4. Calculate the second key landmark based on the dependency relationship of the key landmarks, correct some of the first key landmarks obtained by the solution to obtain the third key landmark, and integrate to obtain the final key landmark. Among them, the set of the uncorrected first key landmarks, the second key landmark, and the third key landmark is the final key landmark.
[0074] In an embodiment of the present invention, the second key landmark can be calculated according to the positional relationship of the key landmarks or the line segment ratio relationship between the line segments formed by the key landmarks, and some of the first key landmarks obtained by the solution are corrected according to the positional relationship of the key landmarks to obtain the third key landmark.
[0075] As Figure 3 shown, in a specific embodiment of the present invention, the second key landmark BL1 (the starting point of the tangent line of the outer surface of the clivus) can be calculated according to the polar coordinate positional relationship between the first key landmark S (sella turcica) and the first key landmark Ba (basion). The specific calculation formula is as follows:
[0076]
[0077] Among them, is the coordinate of the second key landmark to be obtained, represents the rectangular coordinates of the origin of the polar coordinates, ratio is the scaling ratio, θ is the rotation angle, and R is the distance between two first key landmarks that have a dependency relationship with the second key landmark to be obtained. Among them, the origin of the polar coordinates can be any one of the two first key landmarks, and the scaling ratio and the rotation angle are preset values, which are determined according to the polar coordinate positional relationship between the second key landmark to be obtained and the two first key landmarks. Among them, x1 and x2 are the abscissas of the two first key landmarks, and y1 and y2 are the ordinates of the two first key landmarks.
[0078] The second key landmark BL2 can also be calculated according to the line segment ratio relationship between every two points connected by the three points of the second key landmark BL1, the first key landmark D` (the connection point of the root of the pterygoid plate and the outer surface of the clivus), and the second key landmark BL2 (the end point of the tangent line of the outer surface of the clivus). The specific calculation formula is as follows:
[0079]
[0080] Among them, represents the coordinate of the point to be obtained, represents the coordinate of the dependent point, and λ represents the line segment ratio, and the value here is 2.
[0081] According to the positional relationship of the key landmark points, the intersection point of the straight line connecting the first key landmark point Go (gonion) and the first key landmark point B (subspinale) and the straight line connecting the first key landmark point MPW (mid-pharyngeal wall point) and the first key landmark point LPW (lower pharyngeal wall point) is denoted as the second key landmark point TPPW (the intersection of the Go-B line and the posterior pharyngeal wall), and its calculation formula is as follows:
[0082]
[0083] Among them, a1, b1, and c1 are the parameters of the straight line equation a1x + b1y + c1 = 0 of the straight line Go-B, and a2, b2, and c2 are the parameters of the straight line equation a2x + b2y + c2 = 0 of the straight line LPW-MPW. Its calculation formula is as follows:
[0084]
[0085] Among them, (x1, y1), (x2, y2) represent the coordinates of two first key landmark points used to solve the straight line equation.
[0086] Some of the first key landmark points obtained by calculation can also be corrected according to the positional relationship of the key landmark points to obtain the third key landmark point.
[0087] For example: The foot of the perpendicular of the first key landmark point UPW (upper pharyngeal wall point) on the straight line formed by the first key landmark point Ba and the first key landmark point PNS (posterior nasal spine point) is denoted as the third key point UPW, so that the three points UPW, Ba, and PNS are collinear, realizing the correction of the first key landmark point UPW. The foot of the perpendicular of the first key landmark point TB (the intersection of the Go-B line and the root of the tongue) on the straight line formed by the first key landmark point B and the first key landmark point Go is denoted as the third key point TB, so that the three points B, Go, and TB are collinear, realizing the correction of the first key landmark point TB. Among them, the calculation formula of the foot of the perpendicular is as follows:
[0088]
[0089] Among them, is the coordinate of the corrected point, (x, y) is the coordinate of the point before correction, and the calculation methods of a, b, and c are the same as those of a i 、b i 、c i in the previous text.
[0090] The steps for correcting the first key landmark point AD2 (the intersection point of the perpendicular line from PNS to the Ba-S line and the posterior pharyngeal wall) are as follows: First, calculate the foot of the perpendicular of the first key landmark point PNS on the line formed by the first key landmark point Ba and the first key landmark point S, denoted as Cache. Then, using the first key landmark point AD2, the first key landmark point AD (the most convex point of the adenoid), and the first key landmark point UPW predicted by the model, form a parabola, and calculate the intersection point of this parabola and the line connecting Cache and the first key landmark point PNS. This intersection point is the final third key landmark point AD2. Among them, the coordinate calculation formula for the third key landmark point AD2 is as follows:
[0091]
[0092] Among them, a, b, and c are the coefficients of the parabola equation ax 2 +bx + c = 0, and m, n are the coefficients of the line equation y = mx + n. Among them, the specific calculation formulas for a, b, c, m, and n are as follows:
[0093]
[0094]
[0095]
[0096] Among them, (x1, y1), (x2, y2), and (x3, y3) represent the coordinates of three points participating in the calculation of the coefficients in the parabola equation, and (x4, y4) and (x5, y5) represent the coordinates of two points participating in the calculation of the coefficients in the line equation.
[0097] Since the calculation and correction are carried out according to the dependency relationship of the key landmark points, and the final key landmark points are obtained through integration, it can effectively utilize the mutual dependence and constraints among the key landmark points, improve the recognition accuracy, obtain more comprehensive key landmark points from the lateral cephalogram, and meet the needs of most common measurement methods.
[0098] In a specific embodiment of the present invention, on a common desktop i7 CPU, all the key landmark point positions as shown Figure 2 can be calculated within one second.
[0099] According to the method for automatically identifying key landmark points on a lateral cephalogram according to an embodiment of the present invention, by inputting the lateral cephalogram to be identified into a key landmark point recognition model, a heat map of each first key landmark point is obtained, and the heat map is resolved to obtain the first key landmark point, avoiding directly identifying the key landmark points in the lateral cephalogram using a complex deep learning model, with faster calculation speed, higher robustness, and lower requirements for hardware resources; calculating and correcting according to the dependency relationship of the key landmark points, and integrating to obtain the final key landmark points, which can effectively utilize the mutual dependence and constraints between the key landmark points, accurately and comprehensively identify the key landmark points in the lateral cephalogram, and meet the needs of most common measurement methods.
[0100] Corresponding to the method for automatically identifying key landmark points on a lateral cephalogram in the above embodiment, the present invention also proposes an apparatus for automatically identifying key landmark points on a lateral cephalogram.
[0101] As Figure 4 shown, the apparatus for automatically identifying key landmark points on a lateral cephalogram according to an embodiment of the present invention includes: a first acquisition module 10, an output module 20, a resolution module 30, and an integration module 40. Among them, the first acquisition module 10 is used to acquire the lateral cephalogram to be identified; the output module 20 is used to input the lateral cephalogram to be identified into a key landmark point recognition model to obtain a heat map of each key landmark point; the resolution module 30 is used to resolve the heat map to obtain the initial key landmark points of the lateral cephalogram; the integration module 40 is used to calculate to obtain the second key landmark points according to the dependency relationship of the key landmark points, and correct some of the first key landmark points obtained by resolution to obtain the third key landmark points, and integrate to obtain the final key landmark points, where the set of the uncorrected first key landmark points, the second key landmark points, and the third key landmark points is the final key landmark points.
[0102] It should be noted that inputting the lateral cephalogram to be identified into a key landmark point recognition model, obtaining a heat map of each first key landmark point, and resolving the heat map to obtain the first key landmark points of the lateral cephalogram consume less resources than directly identifying the key landmark points using a deep learning model, with faster calculation speed, higher robustness, and lower requirements for hardware resources. Since the calculation and correction are performed according to the dependency relationship of the key landmark points, and the final key landmark points are obtained through integration, the mutual dependence and constraints between the key landmark points can be effectively utilized, the recognition accuracy can be improved, more comprehensive key landmark points can be obtained from the lateral cephalogram, and the needs of most common measurement methods can be met.
[0103] Key landmark points of the airway can be used to judge the development status of the airway and diagnose diseases such as airway stenosis; key landmark points of bone age can not only determine the biological age of children, but also help to understand the growth and development potential and the trend of sexual maturity of children at an early stage through bone age. Bone age can also be used to predict the adult height of children. The determination of bone age is also of great help to the diagnosis of some pediatric endocrine diseases and has great guiding significance for the treatment of some patients with short stature.
[0104] However, most of the current methods for identifying landmark points on lateral cephalograms through deep learning only include cephalometric key landmark points, without bone age key landmark points and airway key landmark points, and cannot diagnose airway stenosis and determine bone age.
[0105] As Figure 2 shown, in an embodiment of the present invention, the key landmark points of the lateral cephalogram may include: cephalometric key landmark points, bone age key landmark points, and airway key landmark points. In the figure, the circular points are cephalometric key landmark points, the pentagram points are bone age key landmark points, and the triangular points are airway key landmark points. Identifying these key landmark points is beneficial to the diagnosis of airway stenosis and the determination of bone age, meeting the needs of common measurement methods; taking two points with a fixed distance on the positioning ruler as cephalometric key landmark points is beneficial to determining the scale of the lateral cephalogram and improving the accuracy of identification.
[0106] In an embodiment of the present invention, the automatic recognition device for key landmark points of the lateral cephalogram further includes: a second acquisition module, a preprocessing module, a data division module, an image enhancement module, a conversion module, and a training module. Among them, the second acquisition module is used to acquire sample lateral cephalograms and corresponding standard result files; the preprocessing module is used to preprocess the sample lateral cephalograms, determine the positions of the key landmark points of the sample lateral cephalograms according to the positions of the key landmark points in the standard result files, and obtain a sample data set; the data division module is used to divide the sample data set into a training set, a validation set, and a test set; the image enhancement model is used to perform image enhancement on the sample lateral cephalograms in the training set; the conversion module is used to convert the position data of the key landmark points of the sample lateral cephalograms in the sample data set into a heat map; the training module is used to input the heat map into the deep learning model network for training to obtain a key landmark point recognition model.
[0107] In a specific embodiment of the present invention, the second acquisition module is used to collect a sufficient amount of lateral cephalometric X-ray films of the head, including some low-quality lateral cephalometric X-ray films of the head, such as unclear soft tissue imaging; non-overlapping of the mechanical ear points and mandibular angles on both sides; the head not being perpendicular to the ground; anterior open bite; directly taking a "paper lateral cephalometric X-ray film of the head" with a mobile phone, resulting in warping of the paper; pictures intercepted using a screenshot tool, etc.; and obtaining a standard result file corresponding to the lateral cephalometric X-ray film of the head by four experts making detailed annotations on 78 key landmark points required for methods such as cephalometric analysis, bone age analysis, and airway analysis. Each landmark point is as Figure 2 shown.
[0108] In an embodiment of the present invention, the preprocessing module is specifically used for: deleting the lateral cephalometric X-ray films of the head with too poor image quality; deleting the lateral cephalometric X-ray films of the head with too large differences in the annotation results of multiple experts using the 3σ principle; in order to reduce the influence of the subjectivity of each expert, determining the average value of the annotation results of multiple experts as the actual position of the key landmark points, that is, the position of the first key landmark points of the lateral cephalometric X-ray films of the head.
[0109] In an embodiment of the present invention, the data division module is used to divide the training set, validation set, and test set according to a ratio of 0.8∶0.15∶0.05. At the same time, in order to maintain the proportion of the lateral cephalometric X-ray films of the head in the mixed dentition stage in all the lateral cephalometric X-ray films of the head, stratified sampling without replacement is adopted.
[0110] In an embodiment of the present invention, the image enhancement module can be used to perform image enhancement on the lateral cephalometric X-ray films of the head in the training set.
[0111] It should be noted that in this embodiment, "performing image enhancement on the lateral cephalometric X-ray films of the head in the training set" is to simulate various low-quality images to cope with the situation that the actually obtained lateral cephalometric X-ray films of the head to be recognized may be low-quality images, and improve the effectiveness of modeling.
[0112] In an embodiment of the present invention, the image enhancement module can perform image enhancement on the lateral cephalometric X-ray films of the head in the training set by using pixel-level transformation or spatial-level transformation.
[0113] In another embodiment of the present invention, the image enhancement module can also perform image enhancement on the lateral cephalometric X-ray films of the head in the training set by using pixel-level transformation and spatial-level transformation at the same time, so as to more comprehensively simulate various low-quality images.
[0114] Specifically, the pixel-level transformation includes: Gaussian blur, where the size of the convolution kernel of Gaussian blur can be 7x7; random brightness transformation, and the range of the random brightness transformation can be [-0.2, 0.2]; gamma transformation, and the set of gamma values can be {0.2, 0.4, 0.67, 1, 1.5, 2.5, 5.0}, where the probability of selecting a gamma value of 1 is 50%, and the gamma transformation can better simulate the effect of image degradation; histogram equalization. When performing pixel-level transformation, the probabilities of selecting these several pixel-level image enhancement methods are 0.2, 0.1, 0.6, and 0.1 respectively.
[0115] The spatial-level transformation includes random edge cropping, and the cropping ratio is [0, 0.2]; random edge filling, the filling ratio is [0, 0.2], and the filling value is 255 (u8); random left-right flipping; affine transformation, where the rotation angle can be [-30°, 30°], and the scaling ratio can be [-0.5, 2]; grid transformation, the size of the grid can be 5x5, and the distortion ratio is [-0.03, 0.03]; elastic transformation, where the α value of the elastic transformation is 200 and the β value is 20. The probabilities of selecting these several spatial-level image enhancement methods are 0.2, 0.2, 0.2, 0.2, and 0.4 respectively.
[0116] In an embodiment of the present invention, the conversion module is used to set the coordinate of a certain first key landmark after image enhancement to be P A =(x A , y A ), and then a heat map centered on P A is generated through the following two-dimensional Gaussian distribution probability density function:
[0117]
[0118] where (x, y) represents the position of the pixel in the heat map, and σ represents the standard deviation of the distribution. In this embodiment, the value of σ is 5.
[0119] Therefore, each first key landmark of each sample lateral cephalogram has a corresponding heat map.
[0120] In an embodiment of the present invention, the deep learning model network is a U-shaped deep learning model network. The overall model is a U-shaped structure. The left half of the U shape is the feature extractor of the model, and the right half is the feature interpreter. The feature extractor and the feature interpreter not only transfer feature information through the lowest layer, but also perform long connections between the outputs of each layer of the feature extractor and each layer of the feature interpreter. This structural design not only avoids the problem of gradient disappearance caused by an overly deep network, but also fuses feature maps of different scales together, making up for the information loss in the downsampling process of the feature extractor.
[0121] Specifically, the feature extractor has a total of 5 layers. Each layer takes the output of the previous layer as its input, and the input of the first layer is the lateral radiograph image. From top to bottom, the shallow layers process the high-resolution feature maps of the image, which contain more detailed texture information of the image, while the deep layers process the low-resolution feature maps of the image, which contain more semantic information of the image. Each layer consists of two modules: downsampling and convolution. The downsampling is responsible for reducing the size of the image feature map, transitioning the feature map from high resolution to low resolution. Its internal structure is a stack of Conv+BN+LeakyReLU sequences, where the Conv has a kernel size of 3, a stride of 2, and the output feature map size is half of the size of its input feature. The convolution module is responsible for extracting the feature information of the feature map. Its internal structure is similar to the sequence stacking structure of the downsampling layer, but it is one layer thicker, i.e., (Conv+BN+LeakyReLU)x2. Among them, the Conv layer has a kernel size of 3, a stride of 1, and a padding of 1, ensuring the same size of the input and output. In addition, the number of channels in each layer increases sequentially to store more refined image feature information. From top to bottom, the number of channels in each layer is 64, 128, 256, 512, and 1024 respectively.
[0122] Correspondingly, the feature interpreter also has five layers (the bottom layer is shared by both). Different layers process image feature maps of different sizes. Specifically, the input of these layers is composed of the concatenation of the output of the previous layer and the output of the corresponding depth layer of the feature extractor. Each layer consists of an upsampling module and a convolution module. Among them, the convolution module is similar to the convolution module in the feature extractor, and the upsampling module is implemented using nearest neighbor interpolation with a scaling factor of 2. Therefore, the size of its output feature map is twice the size of the input feature map. Similarly, each layer of the feature interpreter also has a change in the number of channels, which are 1024, 512, 256, 128, and 64 from bottom to top, representing the decoding of feature information. Finally, the last layer is subjected to convolution calculation with a convolution kernel of 1 to generate the heat map of the first key landmark.
[0123] The loss function of the U-shaped structure deep learning model is binary cross-entropy, and its calculation formula is:
[0124]
[0125] where, represents binary cross-entropy, represents the output heat map calculated by the model, o represents the heat map converted based on the first key landmark, N represents the number of samples, H is the length of the output heat map, W is the width of the output heat map, represents the pixel value at the i-th row and j-th column of the heat map converted from the k-th first key landmark.
[0126] Meanwhile, during the model training process, it is also necessary to make corresponding adjustments to the learning rate. The learning rate is an important hyperparameter when training a neural network, which determines the magnitude of weight updates in each iteration. The impact of the learning rate on the performance of the neural network is very significant. Too high or too low a learning rate may lead to unstable or inefficient training. In the initial stage of model training, the model has a large optimization space, and a relatively large learning rate can accelerate the convergence of the model; however, in the later stage of training, the model optimization space becomes smaller, and a large learning rate may cause the loss function to oscillate around the global minimum and fail to converge.
[0127] In one embodiment of the present invention, a custom method is adopted such that the learning rate automatically decreases as the number of training epochs increases, and the formula is as follows:
[0128]
[0129] where lr epoch represents the learning rate of the current training epoch, lr0 represents the initial learning rate, and epoch represents the current training. In the present invention, lr0 = 0.01.
[0130] In one embodiment of the present invention, a validation set can be selected to verify whether the model training converges, and a test set can also be selected to input the converged model, and the accuracy of the first key landmark can be calculated by solving the heat map obtained from the test output.
[0131] In one embodiment of the present invention, the integration module 40 is used to calculate the second key landmark according to the positional relationship or line segment ratio relationship of the key landmarks, and correct some of the first key landmarks obtained by calculation according to the positional relationship of the key landmarks to obtain the third key landmark.
[0132] In one embodiment of the present invention, the integration module 40 can calculate the second key landmark according to the positional relationship of the key landmarks or the line segment ratio relationship between the line segments formed by the key landmarks, and correct some of the first key landmarks obtained by calculation according to the positional relationship of the key landmarks to obtain the third key landmark.
[0133] As Figure 3 shown, in a specific embodiment of the present invention, the integration module 40 can calculate the second key landmark BL1 according to the polar coordinate positional relationship between the first key landmark S (sella turcica point) and the first key landmark Ba (basion) and the second key landmark BL1 (starting point of the outer tangent of the clivus cranii), and the specific calculation formula is as follows:
[0134]
[0135] where is the coordinate of the second key landmark to be obtained, Indicates the rectangular coordinates of the origin of the polar coordinates, ratio is the scaling ratio, θ is the rotation angle, and R is the distance between two first key landmarks that have a dependency relationship with the second key landmark to be found. Among them, the origin of the polar coordinates can be any one of the two first key landmarks, and the scaling ratio and rotation angle are preset values, which are determined according to the polar coordinate position relationship between the second key landmark to be found and the two first key landmarks. Among them, x1 and x2 are the abscissas of two first key landmarks, and y1 and y2 are the ordinates of two first key landmarks.
[0136] The integration module 40 can also calculate the second key landmark BL2 according to the line segment ratio relationship between every two points connected by the three points of the second key landmark BL1, the first key landmark D` (the connection point of the root of the wing plate and the outer surface of the clivus), and the second key landmark BL2 (the end point of the tangent line of the outer surface of the clivus). The specific calculation formula is as follows:
[0137]
[0138] Among them Indicates the coordinates of the point to be found, Indicates the coordinates of the dependent point, and λ represents the line segment ratio, which is 2 according to the value here.
[0139] The integration module 40 can also, according to the positional relationship of the key landmarks, record the intersection point of the straight line connecting the first key landmark Go (gonion point) and the first key landmark B (subspinale point) and the straight line connecting the first key landmark MPW (mid-pharyngeal wall point) and the first key landmark LPW (lower pharyngeal wall point) as the second key landmark TPPW (the intersection point of the Go-B connection line and the posterior pharyngeal wall). The calculation formula is as follows:
[0140]
[0141] Among them, a1, b1, and c1 are the parameters of the straight line equation a1x + b1y + c1 = 0 of the straight line Go-B, and a2, b2, and c2 are the parameters of the straight line equation a2x + b2y + c2 = 0 of the straight line LPW-MPW. The calculation formula is as follows:
[0142]
[0143] Among them, (x1, y1) and (x2, y2) represent the coordinates of two first key landmarks used to solve the straight line equation.
[0144] The integration module 40 can also correct some of the first key landmarks obtained by calculation according to the positional relationship of the key landmarks to obtain the third key landmark.
[0145] For example: Denote the foot of the perpendicular of the first key landmark point UPW (upper pharyngeal wall point) on the line formed by the first key landmark point Ba and the first key landmark point PNS (posterior nasal spine point) as the third key point UPW, such that the three points UPW, Ba, and PNS are collinear, thus realizing the correction of the first key landmark point UPW. Denote the foot of the perpendicular of the first key landmark point TB (the intersection point of the Go-B connection line and the root of the tongue) on the line formed by the first key landmark point B and the first key landmark point Go as the third key point TB, such that the three points B, Go, and TB are collinear, thus realizing the correction of the first key landmark point TB. Among them, the calculation formula for the foot of the perpendicular is as follows:
[0146]
[0147] Among them, is the coordinate of the corrected point, (x, y) is the coordinate of the point before correction, and the calculation methods of a, b, and c are the same as those of a i , b i , c i in the previous text.
[0148] The correction steps of the integration module 40 for the first key landmark point AD2 (the intersection point of the perpendicular line from the PNS to the Ba-S connection line and the posterior pharyngeal wall) are as follows: First, calculate the foot of the perpendicular of the first key landmark point PNS on the line formed by the first key landmark point Ba and the first key landmark point S, and denote it as Cache. Then, use the first key landmark point AD2, the first key landmark point AD (the most prominent point of the adenoid), and the first key landmark point UPW predicted by the model to form a parabola, and calculate the intersection point of this parabola and the line connecting Cache and the first key landmark point PNS. This intersection point is the final third key landmark point AD2. Among them, the coordinate calculation formula for the third key landmark point AD2 is as follows:
[0149]
[0150] Among them, a, b, and c are the coefficients of the parabola equation ax 2 +bx + c = 0, and m, n are the coefficients of the line equation y = mx + n. Among them, the specific calculation formulas for a, b, c, m, and n are as follows:
[0151]
[0152]
[0153]
[0154] Among them, (x1, y1), (x2, y2), and (x3, y3) represent the coordinates of three points participating in the calculation of the coefficients in the parabola equation, and (x4, y4) and (x5, y5) represent the coordinates of two points participating in the calculation of the coefficients in the straight line equation.
[0155] In a specific embodiment of the present invention, on an ordinary desktop i7 CPU, all the key landmark positions as Figure 2 shown can be calculated within one second.
[0156] According to the automatic recognition device for key landmarks in a lateral cephalogram according to the embodiment of the present invention, by inputting the to-be-recognized lateral cephalogram into the key landmark recognition model, a heat map of each first key landmark is obtained, and the first key landmark is obtained by resolving the heat map, avoiding directly recognizing the key landmarks in the lateral cephalogram using a complex deep learning model, with faster calculation speed, high robustness, and lower requirements for hardware resources; calculating and correcting according to the dependency relationship of the key landmarks, and integrating to obtain the final key landmarks, which can effectively utilize the mutual dependence and constraint between the key landmarks, accurately and comprehensively recognize the key landmarks in the lateral cephalogram, and meet the needs of most common measurement methods.
[0157] In the description of the present invention, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The meaning of "plural" is two or more unless otherwise specifically defined.
[0158] In the present invention, unless otherwise clearly specified and limited, the terms "installed", "connected", "connected", "fixed", etc. shall be understood in a broad sense. For example, it may be a fixed connection, a detachable connection, or integrated; it may be a mechanical connection or an electrical connection; it may be directly connected or indirectly connected through an intermediate medium, and it may be the internal connection of two components or the interaction relationship between two components. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0159] In the present invention, unless otherwise clearly defined or limited, the first feature being "on" or "under" the second feature may mean that the first and second features are in direct contact, or the first and second features are in indirect contact through an intermediate medium. Moreover, the first feature being "above", "over" and "on top of" the second feature may mean that the first feature is directly above or obliquely above the second feature, or simply indicates that the horizontal height of the first feature is higher than that of the second feature. The first feature being "under", "below" and "beneath" the second feature may mean that the first feature is directly below or obliquely below the second feature, or simply indicates that the horizontal height of the first feature is lower than that of the second feature.
[0160] In the description of this specification, the description of reference terms such as "an embodiment", "some embodiments", "examples", "specific examples", or "some examples", etc. means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0161] Any process or method description shown in a flowchart or described in other ways herein can be understood to represent a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process, and the scope of the preferred embodiments of the present invention includes additional implementations, where the functions can be executed in a manner that is not in the order shown or discussed, including in a substantially simultaneous manner according to the functions involved or in the reverse order, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0162] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definitional sequence list of executable instructions for implementing logical functions, which can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in conjunction with these instruction execution systems, apparatus, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion with one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other suitable processing as necessary, and then storing it in a computer memory.
[0163] It should be understood that various parts of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), and the like.
[0164] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0165] In addition, each functional unit in various embodiments of the present invention may be integrated into one processing module, may exist separately as individual units physically, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0166] Although the embodiments of the present invention have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. An automatic recognition method for key landmark points on a lateral cephalogram, characterized in that Including the following steps: S1. Obtain a lateral cephalogram to be recognized; S2. Input the lateral cephalogram to be recognized into a key landmark recognition model to obtain a heat map of each first key landmark; S3. Solve the heat map to obtain the first key landmarks of the lateral cephalogram; S4. Calculate the second key landmarks according to the dependency relationship of the key landmarks, and correct some of the first key landmarks obtained by the solution to obtain the third key landmarks, and obtain the final key landmarks through integration, where the set of the uncorrected first key landmarks, the second key landmarks and the third key landmarks is the final key landmarks; wherein, the second key landmarks are calculated according to the positional relationship of the key landmarks or the line segment ratio relationship between the line segments formed by the key landmarks, and some of the first key landmarks obtained by the solution are corrected according to the positional relationship of the key landmarks to obtain the third key landmarks.
2. The automatic recognition method of key landmark points on a lateral cephalogram according to claim 1, wherein The key landmarks of the lateral cephalogram include: cephalometric key landmarks, skeletal age key landmarks and airway key landmarks.
3. The automatic recognition method for key landmark points on a lateral cephalogram according to claim 1 or 2, characterized in that, Before step S2, it further includes: Obtain sample lateral cephalograms and corresponding standard result files; Preprocess the sample lateral cephalograms, and determine the positions of the first key landmarks of the sample lateral cephalograms according to the positions of the key landmarks in the standard result files to obtain a sample data set; Divide the sample data set into a training set, a validation set and a test set; Perform image enhancement on the sample lateral cephalograms in the training set; Convert the position data of the first key landmarks of the sample lateral cephalograms in the sample data set into heat maps; Input the heat maps into a deep learning model network for training to obtain the key landmark recognition model.
4. The method for automatically identifying key landmark points on a lateral cephalogram according to claim 3, characterized in that, Perform image enhancement on the sample lateral cephalograms in the training set by using pixel-level transformation or spatial-level transformation.
5. The method for automatically identifying key landmark points on a lateral cephalogram according to claim 3, wherein The deep learning model network is a U-shaped structure deep learning model network, and the loss function of the U-shaped structure deep learning model is binary cross-entropy, and its calculation formula is: Among them, represents binary cross-entropy, represents the output heat map calculated by the model, o represents the heat map converted based on the first key landmark, N represents the number of samples, H is the length of the output heat map, and W is the width of the output heat map. represents the pixel value at the i-th row and j-th column of the heat map converted from the k-th first key landmark.
6. An automatic recognition device for key landmark points on a lateral cephalogram, characterized in that, Including: A first acquisition module, which is used to obtain a lateral cephalogram to be recognized; An output module, which is used to input the lateral cephalogram to be recognized into a key landmark recognition model to obtain a heat map of each key landmark; A solution module, which is used to solve the heat map to obtain the first key landmarks of the lateral cephalogram; An integration module, which is used to calculate the second key landmarks according to the dependency relationship of the key landmarks, and correct some of the first key landmarks obtained by the solution to obtain the third key landmarks, and obtain the final key landmarks through integration, where the set of the uncorrected first key landmarks, the second key landmarks and the third key landmarks is the final key landmarks; Among them, the integration module is used to calculate the second key landmark point according to the positional relationship or line segment ratio relationship of the key landmark points, and correct some of the first key landmark points obtained by solution according to the positional relationship of the key landmark points to obtain the third key landmark point.
7. The automatic recognition device for key landmark points on a lateral cephalogram according to claim 6, characterized in that, The key landmark points of the lateral cephalogram include: cephalometric key landmark points, skeletal age key landmark points, and airway key landmark points.
8. The automatic recognition device for key landmark points on a lateral cephalogram according to claim 6 or 7, characterized in that, It also includes: A second acquisition module, which is used to acquire a sample lateral cephalogram and a corresponding standard result file; A preprocessing module, which is used to preprocess the sample lateral cephalogram, and determine the positions of the key landmark points of the sample lateral cephalogram according to the positions of the key landmark points in the standard result file to obtain a sample data set; A data division module, which is used to divide the sample data set into a training set, a validation set, and a test set; An image enhancement module, which is used to perform image enhancement on the sample lateral cephalograms in the training set; A conversion module, which is used to convert the position data of the key landmark points of the sample lateral cephalograms in the sample data set into a heat map; A training module, which is used to input the heat map into a deep learning model network for training to obtain the key landmark point recognition model.
9. The key landmark automatic recognition device for lateral cephalogram according to claim 8, characterized in that, The image enhancement module is used to perform image enhancement on the sample lateral cephalograms in the training set by using pixel-level transformation or spatial-level transformation.
10. The automatic recognition device for key landmark points on a lateral cephalogram according to claim 8, characterized in that, The deep learning model network is a U-shaped structure deep learning model network, and the loss function of the U-shaped structure deep learning model is binary cross-entropy, and its calculation formula is: Among them, represents binary cross-entropy, represents the output heat map calculated by the model, o represents the heat map converted based on the first key landmark, N represents the number of samples, H is the length of the output heat map, and W is the width of the output heat map. represents the pixel value at the i-th row and j-th column of the heat map converted from the k-th first key landmark.
Citation Information
Patent Citations
Face key point correction method and device and computer equipment
CN111444775A
Automatic identification method and equipment for head shadow survey mark point of X-ray head normal position film
CN113948190A