A deep learning-based auxiliary diagnosis system for developmental dysplasia of the hip in children

Through the assisted diagnosis system for children with developmental hip dysplasia based on deep learning, using the example segmentation model and feature point positioning model, the complex and inaccurate diagnosis of children with developmental hip dysplasia is solved, and the rapid and accurate image quality control and hierarchical diagnosis is achieved, which improves the consistency and reliability of the diagnosis.

CN119444726BActive Publication Date: 2025-06-06ZHEJIANG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411581810.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-07
Publication Date
2025-06-06
Estimated Expiration
2044-11-07

AI Technical Summary

Technical Problem

The diagnosis of developmental hip dysplasia in children is complex and has problems with inaccurate diagnosis. Especially due to position errors and insufficient medical resources, it is difficult for the existing technology to achieve fully automated image quality control and diagnosis.

Method used

A deep learning-based auxiliary diagnosis system for children's developmental hip dysplasia, including an instance segmentation model and feature point positioning model, is adopted. Through multi-scale feature extraction and local autosimilarity calculation, the rapid and accurate quality control and hierarchical diagnosis of hip images are achieved.

Benefits of technology

It realizes rapid and accurate image quality control and hierarchical diagnosis of children's developmental hip dysplasia, reduces the demand for human resources, avoids misdiagnosis and misdiagnosis caused by human factors, and improves the consistency and reliability of diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119444726B_ABST
    Figure CN119444726B_ABST
Patent Text Reader

Abstract

The invention discloses a child developmental hip dysplasia auxiliary diagnosis system based on deep learning, including a trained instance segmentation model and a feature point positioning model; the instance segmentation model is used to segment the bilateral obturator foramen and ilium from the child pelvic image; the feature point positioning model is used to accurately locate 8 hip joint feature points from the child pelvic image; the diagnosis steps are as follows: the child pelvic image to be diagnosed is input into the instance segmentation model, the bilateral obturator foramen and ilium are segmented, the maximum transverse diameter and area of ​​the obturator foramen and ilium, and the rotation index and area ratio of the obturator foramen and ilium are obtained to judge the symmetry of the pelvic image; if the pelvic image is determined to be symmetrical, it meets the quality control standard; after the quality control process is completed, the image that meets the quality control standard or all images are selected to be sent to the feature point positioning model for feature point detection and IHDI classification. The invention can realize rapid and accurate image quality control and graded diagnosis of child developmental hip dysplasia.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of medical image segmentation and feature point positioning, and in particular to a children's developmental hip dysplasia auxiliary diagnosis system based on deep learning. Background Art

[0002] Developmental dysplasia of the hip (DDH) is a relatively common disease in newborns, with an incidence of about 1 to 5 cases per 1,000 newborns. Long-term abnormal stress on the hip joint of children may cause changes in its anatomical structure and accelerate the degradation of articular cartilage. In the case of partial femoral dislocation, the femoral head only contacts the outer upper edge of the acetabulum, so that the cartilage and labrum have to bear the impact of the force that should be distributed throughout the hip joint, making the outer upper edge of the acetabulum blunt or defective, aggravating the misalignment of the femoral head and acetabulum, and ultimately leading to early-onset osteoarthritis, bone deformity and lower limb dysfunction. In children with complete femoral dislocation, the acetabulum lacks appropriate mechanical stimulation from the femoral head, which leads to limited acetabular bone development. Hip dislocation may also cause systemic abnormalities in posture and gait, induce movement disorders, and have a lasting negative impact on the patient's quality of life. Therefore, abnormal acetabulum and femoral morphology make the diagnosis of developmental dysplasia of the hip more complicated.

[0003] Deep learning is rapidly advancing towards accurate diagnosis and personalized treatment. In recent years, deep learning has not only been used in segmented application scenarios such as medical imaging report generation, breast cancer screening and diagnosis, but has also gradually expanded to clinical decision support, drug development, personalized treatment, and telemedicine. In particular, in algorithms based on deep learning, the system can process and analyze massive amounts of data and discover potential pathological features that are difficult to discover with traditional methods. At the same time, it can also solve data limitations caused by scarce and weak labels through active learning or interactive segmentation. Deep learning is also driving innovation in medical research. By integrating genomic data, imaging data, and electronic medical records, it can achieve cross-domain, multi-dimensional data analysis, thereby improving diagnosis and treatment results.

[0004] Deep learning has been widely used in medical imaging and orthopedic auxiliary diagnosis. For example, Gaillochet M proposed an improved MRI image reconstruction method. By combining the unsupervised learning reconstruction algorithm with the N4 bias field estimation method, the influence of the bias field on the reconstruction results was effectively reduced, thereby improving the reconstruction quality, especially in terms of visual effects and root mean square error (RMSE). Larson et al. used confounding variables in nonlinear deep learning models to estimate the bone age of children's hand X-rays, achieving expert-level results. Zech et al. used the target detection framework to accurately identify X-rays containing pediatric wrist fractures. Zheng et al. used deep learning to segment the femur and tibia to calculate leg length and determine the lower limb length difference (LLD). The algorithm measured the leg length of each patient in about 1 second, and there was no significant difference with the evaluation results of radiologists.

[0005] With the help of deep learning models, the diagnosis of developmental dysplasia of the hip in children is gradually moving towards semi-automation or even full automation. Deep learning technology is gradually optimizing and reshaping the diagnostic process of developmental dysplasia of the hip, providing solid technical support for more efficient medical services. However, there is no fully automated system for image quality control and diagnosis of developmental dysplasia of the hip in children that has been proposed and conducts continuous and detailed evaluations of children to minimize the potential risk of "late" dislocation.

[0006] As people pay more and more attention to children's health, the demand for disease screening of newborns is also increasing, and the current situation of shortage of medical resources is becoming more and more serious. Therefore, it is urgent to introduce a fully automated image quality control and diagnosis solution to alleviate the inaccurate diagnosis caused by body position errors and insufficient medical resources, and to improve the consistency and reliability of diagnosis. Summary of the invention

[0007] The present invention provides a deep learning-based auxiliary diagnosis system for developmental dysplasia of the hip in children, which can achieve rapid and accurate image quality control and graded diagnosis of developmental dysplasia of the hip in children.

[0008] A deep learning-based auxiliary diagnosis system for developmental hip dysplasia in children, comprising a computer memory, a computer processor, and a computer program stored in the computer memory and executable on the computer processor, wherein the computer memory stores a trained instance segmentation model and a feature point positioning model;

[0009] The instance segmentation model is used to segment the bilateral obturator foramen and ilium from a child's pelvic image;

[0010] The feature point positioning model is used to accurately locate 8 hip joint feature points from the child's pelvic image through multi-scale feature extraction, including the inner edge of the acetabulum on both sides, the upper outer edge of the acetabulum, the midpoint of the femoral metaphysis, and the femoral head ossification center;

[0011] When the computer processor executes the computer program, the following steps are implemented:

[0012] The pelvic images of children to be diagnosed are input into the trained instance segmentation model to segment the bilateral obturator foramen and ilium; based on the segmentation results, the maximum transverse diameter and area of ​​the obturator foramen and ilium are obtained, and the rotation index and area ratio of the obturator foramen and ilium are further calculated to determine the symmetry of the pelvic images; if the pelvic images are determined to be symmetrical, they meet the quality control standards; after the quality control process is completed, the images that meet the quality control standards or all images are selected and sent to the feature point positioning model for feature point detection and IHDI grading.

[0013] Furthermore, the structure of the instance segmentation model is as follows:

[0014] The feature pyramid network (FPN) is used to perform multi-level feature extraction on children's pelvic images of different scales to generate a feature map with spatial information. After the feature map is generated, the region proposal network (RPN) is used to generate candidate region boxes to identify potential target areas in the image. The candidate boxes and feature maps are input into the RoIAlign layer for spatial transformation and pooling operations, and the output results include pixel-level segmentation masks, category predictions, and bounding box predictions, which are used to accurately distinguish the bilateral obturators and ilium in the images.

[0015] When training the instance segmentation model, a segmentation quality control dataset for children with developmental hip dysplasia was constructed. An end-to-end approach was adopted, and the instance segmentation model was fine-tuned with the ImageNet pre-trained model parameters fixed at the bottom two layers;

[0016] In the process of network optimization, stochastic gradient descent is used to optimize the multi-task loss function L. The formula is as follows:

[0017] L=L cls +L box +L mask

[0018] Among them, L cls is the classification loss, L box is the border loss, L mask It is the mask loss, and a two-stage learning rate adjustment strategy is introduced in network optimization.

[0019] Furthermore, the feature point localization model adopts a parallel stacking architecture of multi-resolution sub-networks. The input child pelvic image is first processed by two convolution layers with a convolution kernel of 3×3 and a stride of 2. Batch normalization and ReLU activation are performed after each convolution operation. After two convolutions, the spatial resolution of the feature map is downsampled to one-fourth of the original image.

[0020] The feature maps processed by the two convolutional layers then enter the bottleneck module and transition layer structure in turn. The transition layer structure includes three transition layers, and a new scale branch is added to each transition layer to capture features of different scales.

[0021] Each scale branch first extracts the image features at the current scale layer by layer through 2 Basic Blocks and 2 Basic-SABlocks; then, the branches integrate cross-scale information through feature interaction; finally, the output of the third transition layer is scale-fused, and the features of all different scales are added together and activated through ReLU to generate a fused multi-scale output.

[0022] A new scale branch is added in each transition layer to capture features of different scales, including:

[0023] The first transition layer uses the output of the bottleneck module to generate feature maps that are downsampled by 4 times and 8 times respectively through two parallel 3×3 convolution operations; in the second transition layer, based on the feature map that has been downsampled to 8 times, a convolution operation with a convolution kernel size of 3x3 and a stride of 2 is used to add a feature map that is downsampled by 16 times; in the third transition layer, based on the feature map that has been downsampled to 16 times, a convolution operation with a convolution kernel size of 3x3 and a stride of 2 is used to further add a feature map that is downsampled by 32 times.

[0024] The Basic-SA Block includes two convolutional layers and a SimAM module, and the specific structure is as follows:

[0025] The feature map is first processed by the first convolutional layer, which uses a kernel size of 3×3, a stride of 1, and a padding setting of 1. After the first convolution operation is completed, the output feature map is batch normalized. Next, the normalized feature map is activated by the ReLU function. Then, it is processed by the second convolutional layer, which also uses a 3×3 kernel, a stride of 1, and keeps the padding setting of 1. The convolutional feature map is batch normalized again.

[0026] The SimAM module calculates the local self-similarity of the feature map by minimizing the energy function, ensuring that the more important features have the lowest similarity with their surrounding pixels; and uses the Sigmoid function to enhance these important features;

[0027] Finally, the SimAM module performs weighted processing on the pelvic feature map output by the convolutional layer; the weighted feature map will be fused with the feature map before input into the SimAM module by element-by-element addition, forming an effective residual connection mechanism.

[0028] The formula for minimizing the energy function is:

[0029]

[0030] In the formula, represents the energy value of neuron t; t represents the value of the target neuron, which represents the activation value of the neuron selected in the feature map; represents the mean of all neurons in the channel except the target neuron t; λ is the regularization parameter; finally, the lower the energy value, the more obvious the difference between the target neuron t and other neurons, and the higher its importance;

[0031] pass To measure the importance of key features, the formula of the Sigmoid function is:

[0032]

[0033] Where X is the input feature map; ⊙ represents the element-by-element multiplication operation.

[0034] When training the feature point positioning model, a feature point positioning dataset for children with developmental hip dysplasia was constructed, and an end-to-end approach was adopted. The feature point positioning model was fine-tuned using the CoCo pre-trained model parameters;

[0035] The mean square error MSE is used as the loss function, and the formula is as follows:

[0036]

[0037] In the formula, y i is the true value of the feature point of the i-th sample, is the prediction result of the feature point positioning model for the i-th sample, and n is the number of samples.

[0038] Compared with the prior art, the present invention has the following beneficial effects:

[0039] The method of the present invention uses the acquired diagnostic data set of developmental hip dysplasia in children to train a deep learning instance segmentation model to achieve instance segmentation of the ilium and obturator, and then obtains the obturator rotation index, obturator area ratio, ilium rotation index, and ilium area ratio to measure the symmetry of the image and achieve image quality assessment. In addition, a feature point positioning model for auxiliary diagnosis is designed, which captures the fine features of the hip bones at different scales through a high-resolution network framework, and by measuring the linear separability of neurons, calculates the local self-similarity of the feature map and amplifies the key features, and can automatically identify and locate the key anatomical feature points of the hip joint, including 8 feature points such as the inner edge of the bilateral acetabulum, the upper outer edge of the acetabulum, the midpoint of the femoral epiphysis, and the femoral head ossification center. Finally, a graded diagnosis is performed according to the diagnostic criteria of the International Hip Dysplasia Research Institute. The instance segmentation quality control and feature point positioning method of the present invention only needs to establish a labeled sample library for model training in the early stage, so as to achieve rapid and accurate image quality control and graded diagnosis of developmental hip dysplasia in children. Compared with traditional assessment methods, it reduces manpower and material resources, and effectively avoids missed diagnosis and misdiagnosis caused by human factors in the assessment process, alleviates inaccurate diagnosis caused by posture errors and insufficient medical resources, and improves the consistency and reliability of diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] Figure 1 This is an implementation flow chart of a deep learning-based auxiliary diagnosis system for developmental dysplasia of the hip in children according to the present invention.

[0041] Figure 2 Schematic diagram of the instance segmentation model constructed for the present invention.

[0042] Figure 3 This is a schematic diagram of the result of segmenting the bilateral iliac bones and obturator foramen in the present invention. ;

[0043] Figure 4 Schematic diagram of Basic-SABlock constructed in the present invention.

[0044] Figure 5 Schematic diagram of the feature point positioning model constructed by the present invention.

[0045] Figure 6 Schematic diagram of evaluation of positioning effect of each hip joint feature point in an embodiment of the present invention.

[0046] Figure 7 It is a bar diagram for evaluating the positioning effect of each hip joint feature point in an embodiment of the present invention.

[0047] Figure 8 Schematic diagram of the IHDI grading standard for developmental hip dysplasia.

[0048] Fig. 9Schematic diagram of the grading effect of developmental dysplasia of the hip. DETAILED DESCRIPTION

[0049] The present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be pointed out that the embodiments described below are intended to facilitate the understanding of the present invention and do not have any limiting effect on the present invention.

[0050] like Figure 1 As shown in the figure, a deep learning-based auxiliary diagnosis system for children's developmental hip dysplasia is implemented as follows:

[0051] S01. Obtain pelvic AP X-ray images for the diagnosis of developmental dysplasia of the hip, and screen and construct the corresponding data set for auxiliary diagnosis.

[0052] The children's pelvic anteroposterior X-ray data (DICOM format) obtained from the database were uniformly converted into JPEG format. Using the image data of children aged six months to three years old, and excluding those with other hip joint diseases or postoperative images, a dataset of 1,250 pelvic anteroposterior X-rays was successfully established.

[0053] S02. Perform instance segmentation and annotation on the original images in the constructed children's developmental dysplasia of the hip diagnostic dataset, and construct a children's developmental dysplasia of the hip quality control dataset for training the instance segmentation model based on deep learning.

[0054] The labelme annotation tool was used to segment and annotate the bilateral ilium and obturator areas in the pediatric pelvic images for quality control of pelvic anteroposterior X-ray images. Specifically, on the screened images, the annotation process of the segmented area is achieved by point-by-point marking. First, the hip joint bone area is accurately manually annotated based on the ilium edge and obturator foramen, and these points define the boundaries of the target object. Then, the connection lines of the points will form a complete segmentation contour, so that the final segmented area can truly reflect the bone morphology. The generated training dataset will be used to train the instance segmentation model, so that it has the ability to automatically identify and locate the ilium and obturator bone structures, thereby providing support for subsequent quality control.

[0055] S03. Perform feature point positioning and annotation on the original images in the constructed developmental dysplasia of the hip diagnostic dataset, and construct a feature point positioning dataset for children's developmental dysplasia of the hip for training a feature point positioning model based on deep learning (i.e., HR-SANet).

[0056] The CoCo Annotator annotation tool was used to annotate the upper lateral edge of the bilateral acetabulum, the inner edge of the acetabulum, the ossification center of the femoral head, and the midpoint of the upper edge of the femoral metaphysis, a total of 8 feature points, from right to left, for feature point positioning of HR-SANet (i.e., feature point positioning model) and subsequent diagnosis.

[0057] S04. Use the labeled instance segmentation dataset to train the model and use it to accurately distinguish the bilateral obturator foramen and ilium in the image. Based on these segmentation results, obtain the maximum transverse diameter and area of ​​the obturator foramen and ilium and calculate the rotation index and area ratio to measure the symmetry of the pelvic image.

[0058] like Figure 2 As shown in the figure, the instance segmentation model uses the Feature Pyramid Network (FPN) to extract multi-level features from anteroposterior pelvic X-ray images of different scales and generate feature maps with spatial information. While retaining the image detail information, high-level semantic features are extracted in the layer-by-layer convolution operation. After the feature map is generated, the region proposal network (RPN) is used to generate candidate region boxes to identify potential target areas in the image. Then, the candidate boxes and feature maps are input to the RoIAlign layer for spatial transformation and pooling operations, and the output results include pixel-level segmentation masks, category predictions, and bounding box predictions, which are used to accurately distinguish the bilateral obturators and ilium in the image. Based on these segmentation results, the maximum transverse diameters and areas of the obturators and ilium are obtained, and the rotation index and area ratio are calculated, so as to quantitatively evaluate the symmetry of anteroposterior pelvic X-ray images in non-standard postures (obturator transverse diameters and area ratios are too different).

[0059] The model training is based on the constructed pediatric developmental hip dysplasia dataset in an end-to-end manner. The backbone feature extraction network uses the ImageNet pre-trained model parameters to fix the bottom two layers for fine-tuning. In the network optimization process, stochastic gradient descent is used to optimize the multi-task loss function L.

[0060] L=L cls +L box +L mask

[0061] Among them, L cls is the classification loss, L box is the border loss, L mask It is the mask loss, and a two-stage learning rate adjustment strategy is introduced in network optimization.

[0062] Instance segmentation model pair Figure 3The bilateral ilium (green mark) and obturator foramen (blue mark) shown in the figure were segmented, and then the maximum transverse diameter and area of ​​the obturator foramen and ilium were extracted as the basis for image symmetry analysis. and Represents the maximum transverse diameter of the left and right ilium, respectively, and defines and Respectively represent the maximum transverse diameter of the left and right obturator foramen; meanwhile, define and The left and right iliac bone areas, and =areas of the left and right obturator foramina. If the transverse diameter and area of ​​the bilateral obturator foramina or ilium are significantly different, it means that the pelvic anteroposterior X-ray image has been rotated during the shooting process, and it is necessary to consider retaking the pelvic anteroposterior X-ray to obtain an accurate diagnosis.

[0063] Calculate the rotation index:

[0064] (1) Calculate the maximum transverse diameter of the ilium: First, determine the pixel coordinates on both sides of the minimum circumscribed rectangle of the ilium. Right As shown in Formula 1 and Formula 2, the left horizontal coordinates of the minimum circumscribed rectangle of the bilateral iliac bones are Subtract the right horizontal coordinate Obtain the maximum transverse diameter of the ilium

[0065]

[0066]

[0067] (2) Calculation of the maximum transverse diameter of the obturator: First, determine the pixel coordinates on both sides of the minimum circumscribed rectangle of the pelvic ilium. Right As shown in Formula 3 and Formula 4, the left horizontal coordinates of the minimum circumscribed rectangle of the bilateral iliac bones are Subtract the right horizontal coordinate Obtain the maximum transverse diameter of the closed foramen

[0068]

[0069]

[0070] The results of the ilium rotation index and obturator rotation index are shown in formula 5 and formula 6, which directly compare the maximum transverse diameter of the bilateral ilium. Maximum transverse diameter of closed foramen Method obtained.

[0071]

[0072]

[0073] Calculate the ilium area ratio (Ilium Area Ratio) and obturator area ratio (Obturator Area Ratio):

[0074] (3) By segmenting the pixel size of the ilium and obturator area (left ilium Right ilium Left obturator foramen ) directly reflects the area size of each bone. As shown in Formula 7 and Formula 8, the ilium rotation index is obtained by directly comparing the pixel sizes of the bilateral ilium, which simplifies the calculation while ensuring the accuracy of the measurement unit.

[0075]

[0076]

[0077] These indices can quantitatively evaluate the degree of pelvic rotation and provide an objective basis for determining the symmetry of the hip joint area.

[0078] S05, the Basic-SABlock module is proposed to calculate the local self-similarity of the feature map by minimizing the energy function. The more important the feature, the lower the similarity with its surrounding pixels. Measure the importance of key features and enhance them.

[0079] like Figure 4As shown in the figure, Basic-SABlock consists of a standard convolutional layer and a SimAM module. For the input feature map of the anteroposterior pelvic X-ray film with a width of w, a height of h, and a number of feature channels of c, it is first processed by the first convolutional layer in the middle of the module. The convolution kernel size of this convolutional layer is 3×3, the stride is 1, and the padding is set to 1 to ensure that the output size is consistent with the input. After the convolution operation is completed, the output feature map is standardized by batch normalization to ensure that the mean and variance of each small batch of data remain stable. Next, the normalized feature map is activated by the ReLU function to introduce nonlinear transformation to enrich the feature expression ability. Subsequently, it is processed by the second convolutional layer, which also uses a 3×3 convolution kernel, a stride of 1, and keeps the padding as 1 to further extract deeper pelvic features. The convolutional feature map is batch normalized again to ensure the uniform scale of the data and enhance the convergence and stability of the model.

[0080] The SimAM module plays a role in amplifying key features by measuring the linear separability of neurons. In particular, in pelvic images, since adjacent pixels tend to have high similarity, while the similarity between distant pixels is significantly reduced, this phenomenon is particularly evident in the edge features of the image. When dealing with such complex edge details, the SimAM module calculates the local self-similarity of the feature map by minimizing the energy function (Formula 9), ensuring that the more important features have lower similarity with their surrounding pixels.

[0081] This module is implemented through The importance of key features is measured, and the Sigmoid function is used to enhance these important features (as shown in Formula 10).

[0082]

[0083] Represents the energy value of neuron t, which is used to measure the importance of the neuron. The lower the energy value, the more different or significant the neuron is in visual processing. T is the value of the target neuron, which represents the activation value of the neuron selected in the feature map. It is hoped that the value of this neuron is significantly different from the surrounding neurons so that it can be given a higher weight in the attention mechanism. is the mean of all neurons in the channel except the target neuron t. It represents the average activation value of other neurons in the channel and is used to calculate the difference between the target neuron and its surrounding neurons. is the variance of all neurons in the channel except the target neuron t, indicating the degree of discreteness of the activation values ​​of these neurons. A larger variance means that the activation values ​​of the neurons in the channel are more different, and vice versa, it means that the activation values ​​are more consistent. λ is a regularization parameter, which is used to control the stability of the formula and avoid overfitting. This parameter can prevent oversensitivity to certain extreme values ​​in the calculation.

[0084] The numerator of the formula is used for normalization to make the energy calculation relatively balanced. The denominator represents the difference between the target neuron and other neurons plus the regularization term. Ultimately, the lower the energy value, the more obvious the difference between the target neuron t and other neurons, and the higher its importance.

[0085]

[0086] The Sigmoid function maps the input to a fixed range (0,1), with the left end approaching 0 and the right end approaching 1, which is convenient for interpreting probability. X is the input feature map. ⊙ represents the element-by-element multiplication operation, which applies the weights calculated by Sigmoid to the feature map X element by element, thereby weighting the eigenvalues ​​of each channel. The effect of this formula is to weight the feature map according to the energy value of the neuron, so that important neurons contribute more to the feature map, thereby increasing the model's attention to key information.

[0087] Finally, the SimAM module will perform weighted processing on the pelvic feature map output by the convolutional layer. Specifically, the module adaptively adjusts the weights of pixels to highlight key features and suppress redundant information. The weighted feature map will be fused with the initial input feature map by element-wise addition to form an effective residual connection mechanism. The original information of the input feature map is retained, and enhanced features are introduced to further enrich the model expression. In addition, this residual connection can effectively alleviate the problem of gradient vanishing in deep networks, ensure that gradient information can be smoothly forwarded during training, and promote more efficient training and optimization of the model. The model's perception of key anatomical structures is effectively improved, thereby more accurately capturing the acetabulum and femoral margins. The performance of Basic-SABlock in processing complex bone images such as blurred boundaries and subtle structures has been greatly improved.

[0088] Before constructing the feature point positioning model, the present invention implements data preprocessing and enhancement to optimize the training effect of the model. First, a fixed aspect ratio (height: width = 4:3) is set for pelvic detection, and a standard detection frame of 256×196 pixels is generated according to the ratio, and the original image is cropped to ensure the uniformity of the data input format. Subsequently, a variety of data enhancement strategies are used to improve the model's adaptability to different deformations, scales, and orientation changes, including random scaling (ranging from 0.7 to 1.3 times), random rotation (angle range of -30° to 30°), 50% probability of random horizontal flipping, and normalization.

[0089] S06. The labeled feature point positioning data set was enhanced and model trained. The HR-SANet deep learning model (i.e., feature point positioning model) was designed based on the proposed Basic-SABlock to achieve hip joint feature point positioning and auxiliary diagnosis of developmental hip dysplasia. The feature point positioning model captures the subtle features of the hip bones at different scales through multi-scale feature extraction, ensuring that feature points can be effectively positioned at different resolutions. It is used to accurately locate 8 feature points including the bilateral acetabulum inner edge, the acetabulum upper outer edge, the midpoint of the femoral metaphysis, and the femoral head ossification center.

[0090] like Figure 5 As shown in the figure, in the feature point positioning model, a multi-resolution sub-network parallel stacking architecture is adopted. While maintaining the high-resolution output of the image, it can extract feature information at different levels, aiming to achieve the fusion of multi-scale features and accurately identify the edge features of the hip bones. The input pelvic anteroposterior X-ray image is first processed by two convolution layers with a convolution kernel of 3×3 and a stride of 2. Batch normalization and ReLU activation are performed after each convolution operation to ensure the standardized processing and nonlinear modeling of the features. After two convolutions, the spatial resolution of the feature map is downsampled to one-fourth of the original image.

[0091] The transition layer structure further refines the processing of feature maps. In each transition layer, the model adds a new branch to capture features of different scales. Specifically, the first transition layer uses the output of Layer1 to generate feature maps that are downsampled 4 times and 8 times respectively through two parallel 3×3 convolution operations. In the second transition layer, based on the feature map that has been downsampled to 8 times, the model uses a convolution operation with a convolution kernel size of 3x3 and a stride of 2 to further generate a feature map that is downsampled 16 times. At this time, the output feature size is 16×12×128. In the third transition layer, based on the feature map that has been downsampled to 16 times, a convolution operation with a convolution kernel size of 3x3 and a stride of 2 is still used to generate a feature scale that is downsampled 32 times, and the output feature size is reduced to 8×6×256.

[0092] Each scale branch of the scale fusion structure first extracts the image features at the current scale layer by layer through 2 Basic Blocks and 2 Basic-SABlocks. Then, the integration of cross-scale information is achieved through feature interaction between branches. For example, for the branch that is downsampled 4 times, its output directly maintains the current scale without additional sampling processing; while the branch from the downsampled 8 times is upsampled 2 times through the convolution layer to match the scale of the downsampled 4 times. Similarly, the branch from the downsampled 16 times is upsampled 4 times to achieve scale alignment with the downsampled 4 times input. Finally, all features of different scales are added and activated by ReLU to generate a fused multi-scale output. This cross-scale feature fusion mechanism can effectively transmit information between scales, realize multi-level expression of features, and enhance the model's ability to recognize bone edge feature points.

[0093] The model training is done using a pediatric developmental hip dysplasia dataset constructed from the end-to-end approach, and the backbone feature extraction network is fine-tuned using the CoCo pre-trained model parameters;

[0094] In the feature point localization task of the feature point localization model, the serious imbalance between the number of positive samples and negative samples will affect the convergence speed and accuracy of the network. In order to alleviate this problem, the mean square error (MSE) is used as the loss function, as shown in formula (11), y i is the true value of the feature point of the i-th sample, is the prediction result of the feature point positioning model for the i-th sample.

[0095] Specifically, a two-dimensional Gaussian distribution is applied to the actual coordinate position of each feature point to generate a heat map. These heat maps not only reflect the precise position of the feature points, but also assign certain weights to the areas around the feature points, thereby alleviating the impact of the imbalance of positive and negative samples. In this way, the loss function can measure the average error between the generated predicted heat map and the actual heat map, allowing the network to more effectively learn the spatial distribution of hip joint feature points and accelerate the convergence of the model.

[0096]

[0097] S07. According to the diagnostic criteria of the International Hip Dysplasia Institute, establish the bilateral hip coordinate axis and perform graded diagnosis (based on the principle of independent diagnosis of bilateral hip joints, one image produces two unilateral diagnosis results, such as left hip IHDI grade I and right hip IHDI grade III). Integrate the IHDI grading standard code for diagnosing developmental dysplasia of the hip into the prediction process of the feature point positioning model.

[0098] After the model locates the characteristic points of the hip joint, the grading results of developmental hip dysplasia are generated according to the IHDI standard. At this time, the midpoints of the femoral metaphysis (characteristic points 4 and 8) are both located in the medial and inferior 90° quadrants, and the grading results are bilateral IHDI grade I. Figure 8 shown.

[0099] The standard draws the horizontal Hilgenreiner line by connecting the inner edge of the acetabulum (characteristic points 2 and 6), and draws the vertical and horizontal Perkin line through the outer edge of the acetabulum (characteristic points 1 and 5). The intersection of the two lines constitutes the origin of the coordinate system, forming a coordinate axis. The outer edge of the acetabulum (characteristic points 1 and 5) and the inner edge of the acetabulum (characteristic points 2 and 6) are connected and extended, and the angle between the obtained orange line and the Hilgenreiner line is defined as the acetabular index. The relative position of the midpoint of the femoral metaphysis (characteristic points 4 and 8) is used to evaluate the degree of displacement of the femoral head relative to the acetabulum. The detailed grading standards are shown in Table 1 below. The quadrant of the midpoint of the femoral metaphysis (characteristic points 4 and 8) in the coordinate axis determines the final grading result of IHDI.

[0100] Table 1

[0101]

[0102] Figure 6In the experimental results, the positioning performance of multiple feature points of bilateral hip joints (lateral superior rim of acetabulum, medial rim of acetabulum, center of femoral head, midpoint of epiphysis) was analyzed in detail. The experiment recorded the performance of these feature points under PCK, OKS and AP indicators respectively. From the results, the performance of the medial rim of acetabulum (R) is relatively weak, with a PCK value of 0.861, an OKS value of 0.877, and an AP of 0.815. In contrast, the performance of the midpoint of epiphysis (R) is better, with a PCK of 0.933, an OKS of 0.922, and an AP of 0.904. The model has a strong accuracy in the positioning of this feature point. In the left hip joint, the center of femoral head (L) performed outstandingly, with a PCK value of up to 0.957, and OKS and AP of 0.928 and 0.907 respectively. Overall, the performance of the medial rim of acetabulum (L) and the lateral superior rim of acetabulum (L) is also relatively consistent, with PCK and AP both above 0.9.

[0103] In summary, although there are some differences in the detection performance of different feature points, the model's performance in the three indicators of PCK, OKS, and AP is relatively balanced, especially in the positioning of the upper lateral edge of the acetabulum and the midpoint of the epiphysis, showing higher robustness and accuracy. Figure 7 A visual display was performed, where blue, orange, and green represent the performance of the three indicators PCK, OKS, and AP, respectively.

[0104] like Fig. 9 As shown in the figure, based on the principle of independent diagnosis of bilateral hip joints, a total of 420 diagnostic grading results were obtained from the 210 test set images without quality control, and the diagnostic accuracy of IHDI grade I was 92.63%; after quality control, 176 pelvic anteroposterior X-ray films identified by the quality control model as meeting the standard position produced a total of 352 grading results, and the diagnostic accuracy of IHDI grade I increased to 93.42%. For the diagnosis of IHDI grade II, III and IV images, the accuracy after quality control also showed a similar upward trend.

[0105] The embodiments described above provide a detailed description of the technical solutions and beneficial effects of the present invention. It should be understood that the above are only specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, supplements and equivalent substitutions made within the scope of the principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A deep learning-based auxiliary diagnosis system for developmental hip dysplasia in children, comprising a computer memory, a computer processor, and a computer program stored in the computer memory and executable on the computer processor, characterized in that: The computer memory stores the trained instance segmentation model and feature point positioning model; The instance segmentation model is used to segment the bilateral obturator foramen and ilium from a child's pelvic image; The feature point positioning model is used to accurately locate 8 hip joint feature points from the child's pelvic image through multi-scale feature extraction, including the inner edge of the acetabulum on both sides, the upper outer edge of the acetabulum, the midpoint of the femoral metaphysis, and the femoral head ossification center; The feature point localization model adopts a parallel stacking architecture of multi-resolution sub-networks. The input child pelvic image is first processed by two convolution layers with a convolution kernel of 3×3 and a stride of 2. Batch normalization and ReLU activation are performed after each convolution operation. After two convolutions, the spatial resolution of the feature map is downsampled to one-fourth of the original image. The feature maps processed by the two convolutional layers then enter the bottleneck module and transition layer structure in turn. The transition layer structure includes three transition layers, and a new scale branch is added to each transition layer to capture features of different scales. Each scale branch first extracts the image features at the current scale layer by layer through two Basic Blocks and two Basic-SA Blocks. Then, the branches integrate cross-scale information through feature interaction. Finally, the output of the third transition layer is scale-fused, and all features of different scales are added together and activated through ReLU to generate a fused multi-scale output. When the computer processor executes the computer program, the following steps are implemented: The pelvic image of the child to be diagnosed is input into the trained instance segmentation model to segment the bilateral obturator foramen and ilium. Based on the segmentation results, the maximum transverse diameter and area of ​​the obturator foramen and ilium are obtained, and the rotation index and area ratio of the obturator foramen and ilium are further calculated to determine the symmetry of the pelvic image. If the pelvic image is judged to be symmetrical, it meets the quality control standard; after the quality control process is completed, the images that meet the quality control standard or all images are selected to be sent to the feature point positioning model for feature point detection and IHDI grading.

2. The deep learning-based auxiliary diagnosis system for developmental dysplasia of the hip in children according to claim 1 is characterized in that: The structure of the instance segmentation model is as follows: The feature pyramid network (FPN) is used to perform multi-level feature extraction on children's pelvic images of different scales to generate a feature map with spatial information. After the feature map is generated, the region proposal network (RPN) is used to generate candidate region boxes to identify potential target areas in the image. The candidate boxes and feature maps are input into the RoIAlign layer for spatial transformation and pooling operations, and the output results include pixel-level segmentation masks, category predictions, and bounding box predictions, which are used to accurately distinguish the bilateral obturators and ilium in the images.

3. The deep learning-based auxiliary diagnosis system for developmental dysplasia of the hip in children according to claim 2 is characterized in that: When training the instance segmentation model, a segmentation quality control dataset for children with developmental hip dysplasia was constructed. An end-to-end approach was adopted, and the instance segmentation model was fine-tuned with the ImageNet pre-trained model parameters fixed at the bottom two layers; In the process of network optimization, stochastic gradient descent is used to optimize the multi-task loss function L. The formula is as follows: L=L cls +L box +L mask Among them, L cls is the classification loss, L box is the border loss, L mask It is the mask loss, and a two-stage learning rate adjustment strategy is introduced in network optimization.

4. The deep learning-based auxiliary diagnosis system for developmental hip dysplasia in children according to claim 1 is characterized in that: A new scale branch is added in each transition layer to capture features of different scales, including: The first transition layer uses the output of the bottleneck module to generate feature maps that are downsampled by 4 times and 8 times respectively through two parallel 3×3 convolution operations; in the second transition layer, based on the feature map that has been downsampled to 8 times, a convolution operation with a convolution kernel size of 3x3 and a stride of 2 is used to add a feature map that is downsampled by 16 times; in the third transition layer, based on the feature map that has been downsampled to 16 times, a convolution operation with a convolution kernel size of 3x3 and a stride of 2 is used to further add a feature map that is downsampled by 32 times.

5. The deep learning-based auxiliary diagnosis system for developmental dysplasia of the hip in children according to claim 1 is characterized in that: The Basic-SA Block includes two convolutional layers and a SimAM module, and the specific structure is as follows: The feature map is first processed by the first convolutional layer, which uses a kernel size of 3×3, a stride of 1, and a padding setting of 1. After the first convolution operation is completed, the output feature map is batch normalized. Next, the normalized feature map is activated by the ReLU function. Then, it is processed by the second convolutional layer, which also uses a 3×3 kernel, a stride of 1, and keeps the padding setting of 1. The convolutional feature map is batch normalized again. The SimAM module calculates the local self-similarity of the feature map by minimizing the energy function, ensuring that the more important features have the lowest similarity with their surrounding pixels; and uses the Sigmoid function to enhance these important features; Finally, the SimAM module performs weighted processing on the pelvic feature map output by the convolutional layer; the weighted feature map will be fused with the feature map before input into the SimAM module by element-by-element addition, forming an effective residual connection mechanism.

6. The deep learning-based auxiliary diagnosis system for developmental dysplasia of the hip in children according to claim 5 is characterized in that: The formula for minimizing the energy function is: In the formula, represents the energy value of neuron t; t represents the value of the target neuron, which represents the activation value of the neuron selected in the feature map; represents the mean of all neurons in the channel except the target neuron t; λ is the regularization parameter; finally, the lower the energy value, the more obvious the difference between the target neuron t and other neurons, and the higher its importance; pass To measure the importance of key features, the formula of the Sigmoid function is: Where X is the input feature map; ⊙ represents the element-by-element multiplication operation.

7. The deep learning-based auxiliary diagnosis system for developmental dysplasia of the hip in children according to claim 1 is characterized in that: When training the feature point positioning model, a feature point positioning dataset for children with developmental hip dysplasia was constructed, and an end-to-end approach was adopted. The feature point positioning model was fine-tuned using the CoCo pre-trained model parameters; The mean square error MSE is used as the loss function, and the formula is as follows: In the formula, y i is the true value of the feature point of the i-th sample, is the prediction result of the feature point positioning model for the i-th sample, and n is the number of samples.