Children skeletal development image recognition system
By designing a children's bone development image recognition system, using separable convolutional neural network and recognition network model, the problems of subjectivity and inefficiency in traditional evaluation methods are solved, and more accurate and efficient bone development assessment is achieved.
Patent Information
- Application Number
- CN202510184063.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-19
- Publication Date
- 2025-05-27
AI Technical Summary
Traditional children's skeletal development assessment relies on physician experience, and there are problems of subjectivity and inefficiency, resulting in a lack of consistency and accuracy of the evaluation results.
A children's bone development image recognition system is designed, and the image development results are automatically generated through the image reception module, image information reading module, image processing module and image recognition module.
It improves the accuracy and efficiency of bone development assessment, reduces the subjectivity and workload of manual assessment, and provides an objective and quantitative reference basis.
Smart Images

Figure CN120047427A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of image recognition, and particularly relates to a child bone development image recognition system. Background Art
[0002] With the development of medical imaging technology, image recognition technology for medical images has emerged. This technology can quickly and accurately process and analyze image data, extract and recognize features in the image through algorithms, and obtain objective conclusions.
[0003] In traditional technology, the assessment of children's bone development mainly relies on doctors' experience and visual observation of X-ray films. Doctors obtain bone age features by observing the bone morphology, size, density, etc. in the X-ray films, and then judge whether the children's bone development is normal based on the knowledge and experience they have accumulated.
[0004] However, in the current traditional assessment system, on the one hand, it is greatly affected by doctors' personal experience and subjective factors. Different doctors may have different judgments on the same X-ray film, resulting in a lack of consistency and accuracy in the assessment results. On the other hand, when making an artificial assessment, doctors need to compare the bone morphology of the X-ray film with the standard bone age atlas to obtain the bone development situation, resulting in low assessment efficiency. Summary of the Invention
[0005] Based on this, it is necessary to provide a child bone development image recognition system that can automatically generate bone development results based on children's bone images to accurately assess children's growth status and provide reference diagnostic opinions for medical staff in view of the above technical problems.
[0006] In a first aspect, the present application provides a child bone development image recognition system, including:
[0007] An image receiving module for receiving an input image; the input image includes a hand X-ray image;
[0008] An image information reading module for parsing the input image to obtain corresponding image information; the image information includes width, height, number of channels, and pixel set;
[0009] An image processing module for processing the input image into a standardized image according to the image information by using an image registration technology based on a separable convolutional neural network;
[0010] An image recognition module for recognizing the standardized image by using a recognition network model to obtain a bone development result corresponding to the input image.
[0011] In one of the embodiments, the image recognition module includes an entrance module, an intermediate module, and an exit module;
[0012] The input module is used to perform a convolution operation on the standardized image to obtain a feature map representing the features of the standardized image; the convolution operation includes two layers of standard convolution and three layers of separable convolution;
[0013] The middle module is used to perform feature extraction and feature enhancement on the feature map through a residual convolution network to obtain an enhanced feature map; the residual convolution network includes eight repeated three-layer residual convolutions;
[0014] The output module is used to convert the enhanced feature map into a set of feature vectors by using a global average pooling operator, and transfer the set of feature vectors to the fully connected layer of the output module to obtain the skeletal development result.
[0015] In one embodiment, performing a convolution operation on the standardized image to obtain a feature map representing the features of the standardized image includes:
[0016] Using the following formula to perform a convolution operation of two layers of standard convolution to obtain a preliminary feature representation feature map:
[0017]
[0018] V' = V(V(m,n,o out ))
[0019] where V is the standard convolution; K is the convolution kernel; X is the input image; m is the horizontal coordinate of the position of the standardized image; n is the vertical coordinate of the position of the standardized image; o in is the number of channels of the standardized image; o out is the number of channels of the output feature map; j is the offset of the convolution kernel in the horizontal direction relative to m; k is the offset of the convolution kernel in the vertical direction relative to n; V’ is the two layers of standard convolution;
[0020] Using the following formula to perform a convolution operation of three layers of separable convolution on the preliminary feature representation feature map to obtain a feature map representing the features of the standardized image:
[0021]
[0022]
[0023] G(m,n,o out ) = P(m,n,o in )Q(m,n,o out )
[0024] G” = G(G(G(m,n,o out )))
[0025] where P is the depth convolution; Q is the pointwise convolution; o in is the number of channels of the standardized image; o outis the number of channels of the output feature map; K is the convolution kernel; X is the input image; m is the horizontal coordinate of the position of the normalized image; n is the vertical coordinate of the position of the normalized image; j is the offset of the convolution kernel in the horizontal direction relative to m; k is the offset of the convolution kernel in the vertical direction relative to n; G is the depthwise separable convolution; G” is the three-layer depthwise separable convolution.
[0026] In one embodiment, the residual convolutional network includes eight repeated three-layer residual convolutions, and each three-layer residual convolution is used for:
[0027] Performing the convolution operation of the three-layer residual convolution by using the following formula:
[0028] U(m,n,o out ) = Y(m,n,o in ) + X(m,n,o in )
[0029] U” = Y(Y(Y(m,n,o in ))) + X(m,n,o in )
[0030] where U is the single-layer residual convolution; Y(m,no in ) is the feature map generated after convolution by the entrance module; U” is the three-layer residual convolution; X is the image information of the normalized image.
[0031] In one embodiment, the image processing module includes a segmentation module, a rotation module, and an alignment module;
[0032] The segmentation module is used to extract the hand region from the input image according to the image information to obtain the hand region image;
[0033] The rotation module is used for:
[0034] Detecting key points of the hand region image to obtain the position coordinates of the key points; wherein, the position coordinates of the key points include the position coordinates of the thumb tip, the position coordinates of the middle finger tip, the position coordinates of the little finger tip, and the position coordinates of the midpoint at the bottom of the carpal region;
[0035] Obtaining the rotation angle according to the difference between the position coordinates of the middle finger tip and the position coordinates of the midpoint at the bottom of the carpal region, and rotating the hand region image according to the rotation angle to obtain the upright hand image;
[0036] The alignment module is used to align the upright hand image so that the hand is located in the middle of the image to obtain the normalized image.
[0037] In one embodiment, the segmentation module includes a feature extraction module, an ASPP module, a mask generation module, and an image vector operation module;
[0038] The feature extraction module is used to perform multi-level preliminary feature extraction on the input image to obtain a basic feature vector;
[0039] The ASPP module is used to perform feature sampling on the input image from different scales using parallel downsampling branches with different dilation rates to obtain a differential feature vector;
[0040] The mask generation module is used to generate a mask label with the same size as the input image according to the basic feature vector and the differential feature vector;
[0041] The image vector operation module is used to obtain the hand region image according to the mask label and the image information.
[0042] In one embodiment, obtaining the hand region image according to the mask label and the image information includes:
[0043] Using the following calculation formula, the hand region image is obtained according to the mask label and the image information:
[0044] X 1 = X mask ⊙ X
[0045] where X 1 is the hand region image; X mssk is the mask label; X is the image information.
[0046] In one embodiment, obtaining the rotation angle according to the difference between the position coordinates of the middle finger fingertip and the position coordinates of the midpoint at the bottom of the standard carpal region includes:
[0047] Using the following calculation formula, the rotation angle is obtained according to the position coordinates of the middle finger fingertip and the position coordinates of the midpoint at the bottom of the carpal region:
[0048]
[0049] where KP 2 is the position coordinates of the middle finger fingertip; KP 4 is the position coordinates of the midpoint at the bottom of the carpal region; n is the ordinate of the position coordinates; m is the abscissa of the position coordinates, and θ is the rotation angle.
[0050] In one embodiment, the alignment module includes a horizontal alignment module, a vertical alignment module, and a zero-padding processing module;
[0051] The horizontal alignment module is used to coincide the midpoint coordinates of the position coordinates of the thumb fingertip and the position coordinates of the little finger fingertip with the horizontal midpoint coordinates of the upright hand image through a rigid translation transformation;
[0052] The vertical alignment module is used to align the midpoint coordinates of the position coordinates of the middle finger fingertip and the midpoint of the bottom of the carpal region with the horizontal midpoint coordinates of the upright hand image through rigid translation transformation;
[0053] The zero-padding processing module is used to fill pixel points in the blank edges of the upright hand image to ensure that the standardized image is output according to the set size.
[0054] In one embodiment, the bone development result includes bone age;
[0055] Pass the set of feature vectors to the fully connected layer of the exit module to obtain the bone development result, including:
[0056] Use the following formula to generate the bone age in the bone development result:
[0057]
[0058] where R BA is the bone age; ω is the max pooling operator; V i is the output result after two-layer standard convolution of the entrance module; V j is the output result after three-layer separable convolution of the entrance module; τ is the global average pooling operator; N is the fully connected layer; X φ2 is the enhanced feature map.
[0059] In the above-mentioned child bone development image recognition system, the image receiving module receives hand X-ray images, etc. as input images. The bone images that are easily obtained through medical channels are used as key data for child bone development assessment, which improves the convenience of system popularization and application. The image information reading module analyzes the detailed information of the input image, providing a basis for subsequent image processing and analysis, ensuring that the system comprehensively grasps the image features, and providing data support for accurate image processing and bone development assessment. The separable convolutional neural network adopted by the image processing module has high computational efficiency, can effectively reduce the amount of calculation, and improve the processing speed. The image registration technology converts the image into a standard image, unifies the position, angle, etc. of the hand in the image, removes interference factors caused by shooting postures and angle differences, enables the subsequent recognition network model to focus more on bone development-related features, and improves the accuracy of the results. The image recognition module recognizes the standardized image. The trained recognition network model can efficiently extract and analyze features, accurately identify the bone features in the image, and obtain the bone development result. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] To more clearly illustrate the technical solutions in the embodiments of the present application or in the related art, the following will briefly introduce the drawings required for use in the description of the embodiments or the related art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0061] Figure 1 Schematic diagram of the composition of a child's skeletal development image recognition system;
[0062] Figure 2 Schematic diagram of the composition of the image recognition module of a child's skeletal development image recognition system;
[0063] Figure 3 Schematic diagram of the composition of the image processing module of a child's skeletal development image recognition system;
[0064] Figure 4 Schematic diagram of the composition of the segmentation module of a child's skeletal development image recognition system;
[0065] Figure 5 Schematic diagram of the composition of the alignment module of a child's skeletal development image recognition system;
[0066] Figure 6 Method flowchart adapted to a child's skeletal development image recognition system. Detailed implementation manners
[0067] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the following further details the present application in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0068] In one embodiment, as Figure 1 shown, a child's skeletal development image recognition system is provided. In this embodiment, it is exemplified that the system is applied to a terminal. It can be understood that the system can also be applied to a server, and can also be applied to a system including a terminal and a server, and is implemented through the interaction between the terminal and the server. In this embodiment, the system includes:
[0069] An image receiving module, configured to receive an input image. The input image includes a hand X-ray image.
[0070] Generally, the assessment of children's skeletal development is usually carried out by examining the bones of the hand and wrist through X-ray films. This area contains multiple growth plates. The number of bones in the hand is large and the changes are obvious, which can better reflect the overall skeletal development status of the body. Hand X-ray images include the morphology, size, density of hand bones, and the condition of bone growth plates, etc., and are used to judge skeletal development. Specifically, the bone growth plate will gradually become thinner during children's growth. The hand bones of children of different ages show different characteristics in X-ray images. In infancy, the bones are small and the ossification centers have not yet fully developed. While in adolescence, the bones grow and the ossification centers increase and the degree of fusion changes. Particularly, in clinical practice, the X-ray image of the left hand, which is the less commonly used hand, is used.
[0071] An image information reading module, which is used to parse the input image to obtain the corresponding image information. The image information includes width, height, number of channels, and pixel set.
[0072] Exemplarily, the hand X-ray image is saved in the PNG (Portable Network Graphic) format. The reading operation of the image information reading module will parse the image file header information to obtain the width, height, and channel information of the image, and analyze the pixel distribution data of the input image to obtain the pixel set. These information can describe the characteristics such as the size and color mode of the input image, and provide guidance for subsequent operations. Schematically, when the image processing module performs image processing, the channel information will affect the way the separable convolutional neural network processes the input image data. Different channels are processed according to different convolutional operations, and then different features are extracted. Generally, the X-ray image is a grayscale image with 1 channel. If it is a color image, the number of channels is generally 3.
[0073] An image processing module, which is used to process the input image into a standardized image according to the image information by using the image registration technology based on the separable convolutional neural network.
[0074] There are various differences in the original input hand X-ray images, such as the shooting posture, angle, and differences in image resolution and contrast. These differences will interfere with the accurate extraction and analysis of skeletal development characteristics in the subsequent process. Standardizing the image through image registration technology can unify the position and angle of the hand in the image, reduce the interference of skeletal morphological changes caused by different shooting postures, and make the subsequent bone age assessment more focused on the development characteristics of the bones themselves, improving the accuracy of the assessment.
[0075] An image recognition module, which is used to recognize the standardized image by using the recognition network model to obtain the skeletal development result corresponding to the input image.
[0076] The recognition network model is an Xception-41 regressor built based on deep learning, which contains multiple layers and different types of convolutional operations. It can automatically extract multi-level and multi-dimensional features from images, capture features closely related to bone development such as the morphology of bones, the structure of joints, and the growth plates of bones. After a large amount of training and optimization, it can accurately establish a mapping relationship between image features and bone development results.
[0077] The obtained bone development results (such as bone age) provide an objective and quantitative reference basis for medical staff. In clinical diagnosis, doctors can combine the actual age of the patient and the predicted bone age to judge whether their growth and development are normal. If there is a large difference between the bone age and the actual age, it may indicate problems such as growth disorders in the patient, which helps doctors detect potential health risks in a timely manner and formulate corresponding treatment plans. In addition, the automated processing method of this module also improves the diagnostic efficiency and reduces the subjectivity and workload of manual evaluation.
[0078] In one embodiment, as Figure 2 shown, the image recognition module includes an entrance module, an intermediate module, and an exit module.
[0079] The entrance module is used to perform convolutional operations on the standardized image to obtain a feature map of the feature representation of the standardized image. The convolutional operations include two layers of standard convolution and three layers of separable convolution.
[0080] The entrance module undertakes the key task of initial feature extraction from the standardized image. Through two layers of standard convolution and three layers of separable convolution operations, it transforms the standardized image into a feature map with rich feature representation. The feature map of the feature representation of the standardized image after the convolution operations of two layers of standard convolution and three layers of separable convolution not only contains rich spatial features but also has more advanced semantic information, such as features closely related to bone development like the maturity and growth trend of bones.
[0081] The intermediate module is used to extract and strengthen features from the feature map through a residual convolutional network to obtain a strengthened feature map. The residual convolutional network includes eight repeated three-layer residual convolutions.
[0082] The intermediate module is mainly composed of eight repeated three-layer residual convolutional networks. Each network uses separable convolution operations, that is, decomposing the standard convolution into depthwise convolution and pointwise convolution, which can efficiently extract the spatial and channel information of the image while reducing the computational amount and the number of parameters. As the data is sequentially passed through the eight repeated modules, the network can gradually mine deeper and more representative features in the image, thereby obtaining an enhanced feature map. Schematically, in each residual convolutional network, the input feature is added to the feature after the convolution operation. This fusion method can not only retain the important information in the feature map of the entrance module, but also introduce new features through the convolution operation, further strengthening the feature representation. After being processed by the eight repeated modules, the enhanced feature map output by the intermediate module contains rich image features, providing support for the subsequent exit module to predict the skeletal development result.
[0083] The exit module is used to convert the enhanced feature map into a set of feature vectors by using the global average pooling operator, and pass the set of feature vectors to the fully connected layer of the exit module to obtain the skeletal development result.
[0084] The exit module uses the global average pooling operator to reduce the size of the enhanced feature map to 1×1×2048 and processes each channel of the enhanced feature map separately. For each channel in the enhanced feature map, the global average pooling calculates the average value of all elements in the channel, aggregating the feature information of the entire channel into a scalar value. Exemplarily, if the enhanced feature map has A channels, after global average pooling, a feature vector of length A will be obtained. This set of feature vectors contains the global feature information of each channel in the enhanced feature map. The fully connected layer receives the set of feature vectors obtained by global average pooling as input, combines and maps the input features through a series of linear transformations, activation functions, etc., and finally identifies the skeletal development result.
[0085] The multi-layer convolution operations of the entrance and intermediate modules enable the model to extract image features at different scales and levels, enhancing the adaptability to various hand X-ray images. The residual connection enables the network to better capture subtle differences and improves the model's ability to identify differences in skeletal features of different individuals. The combination of the global average pooling and the fully connected layer in the exit module converts complex features into stable prediction results, enabling the model to accurately predict the skeletal development result even when facing new and unseen images, improving the generalization ability of the model and making it more valuable for practical applications.
[0086] In one embodiment, performing a convolution operation on the normalized image to obtain a feature map of the feature representation of the normalized image includes:
[0087] Using the following formula to perform a convolution operation of two layers of standard convolution to obtain a preliminary feature representation feature map:
[0088]
[0089] V' = V(V(m, n, o out ))
[0090] where V is the standard convolution; K is the convolution kernel; X is the normalized image; m is the horizontal coordinate of the position of the normalized image; n is the vertical coordinate of the position of the normalized image; o in is the number of channels of the normalized image; o out is the number of channels of the output feature map; j is the offset of the convolution kernel in the horizontal direction relative to m; k is the offset of the convolution kernel in the vertical direction relative to n; V’ is the two-layer standard convolution;
[0091] This formula means that on the normalized image, through the convolution operation of the convolution kernel at different positions and channels, the feature map on the output channel is obtained. Specifically, first use the first standard convolution kernel K 1 to perform a convolution operation on the normalized image X. During the convolution process, the convolution kernel K 1 slides on the image X according to the set stride, and performs a convolution operation on the input channel o in at each position (m, n) to obtain the feature map Y 1 output by the first standard convolution layer. Then, taking Y 1 as the input of the second standard convolution layer, use the second standard convolution kernel K 2 to perform a similar convolution operation to obtain the preliminary feature representation feature map Y 2 after two layers of standard convolution. These two layers of standard convolution initially extract the basic features of the image, such as low-level features like edges and textures.
[0092] Among them, the convolution operation is essentially a process of the convolution kernel sliding on the input image and multiplying and accumulating element by element. In a two-dimensional plane, the convolution kernel needs to move in both the horizontal and vertical directions to traverse the entire input image. Here, m and n determine the current position of the convolution kernel on the input image, and j and k respectively represent the offsets of the convolution kernel in the horizontal and vertical directions relative to the current position. Specifically, when calculating a certain position (m, n, o out) When calculating the value, the convolutional kernel will take (m, n) as the center position and operate within a certain range around it. The value ranges of j and k are determined by the size of the convolutional kernel. By changing the values of j and k, the convolutional kernel can cover pixel regions at different positions in the input image. Multiply the pixel values in these regions by the values at the corresponding positions of the convolutional kernel and then sum them up to finally obtain the value at the corresponding position in the output feature map. Exemplarily, assuming the size of the convolutional kernel is 3×3, then the value ranges of j and k are usually {-1, 0, 1}. When j = -1 and k = -1, it means that the element in the upper left corner of the convolutional kernel corresponds to the position that is one pixel offset to the upper left of the current position (m, n) in the input image; when j = 0 and k = 0, the central element of the convolutional kernel corresponds to the current position (m, n) of the input image. By traversing all combinations of j and k within their value ranges, the convolution operation of the convolutional kernel at the current position can be completed, thereby calculating the value of the output feature map at this position.
[0093] Use the following formula to perform a convolution operation of three separable convolutions on the preliminary feature representation feature map to obtain the feature map of the feature representation of the normalized image:
[0094]
[0095] G(m,n,o out ) = P(m,n,o in )Q(m,n,o out )
[0096] G” = G(G(G(m,n,o out )))
[0097] Where P is the depthwise convolution; Q is the pointwise convolution; o in is the number of channels of the normalized image; o out is the number of channels of the output feature map; K is the convolutional kernel; X is the input image; m is the horizontal coordinate of the position of the normalized image; n is the vertical coordinate of the position of the normalized image; j is the offset of the convolutional kernel in the horizontal direction relative to m; k is the offset of the convolutional kernel in the vertical direction relative to n; G is the separable convolution; G” is the three-layer separable convolution.
[0098] The depthwise convolution performs a convolution operation on each input channel separately, only considering the convolution in the spatial dimension and not involving the fusion between channels. While the pointwise convolution, based on the result of the depthwise convolution, performs a convolution operation on the result of the depthwise convolution in the channel dimension to achieve information fusion between channels, improving the computational efficiency while realizing the feature extraction function of the standard convolution.
[0099] Exemplarily, take the preliminary feature representation feature map Y 2 after two layers of standard convolutions as the input. First, perform the first layer of depthwise convolution on Y 2Each input channel is convolved using a depth convolution kernel, performing convolution operations only in the spatial dimension without involving cross-channel fusion, to obtain the intermediate result Z of depth convolution. 1 Then, the first layer of pointwise convolution is performed on Z 1 by convolving in the channel dimension to achieve cross-channel information fusion, obtaining the output Z of the first layer of separable convolution. 2 In the same way, the second and third layers of separable convolution are sequentially performed, finally obtaining the feature map Z after two standard convolutions and three layers of separable convolutions. 4 Z 4 is added to the standardized image X through a residual convolution connection for operations such as size adjustment, finally obtaining the feature map of the feature representation of the standardized image.
[0100] Through two layers of standard convolution, the entry module extracts basic features from the standardized image, and the obtained feature map changes in terms of the number of channels, size, and semantic level. In terms of the number of channels, according to the setting of the number of output channels in the standard convolution formula, the number of channels of the feature map will be adjusted according to the model design, and the increased number of channels can accommodate more different types of features. In terms of size, it usually changes according to parameters such as the convolution kernel size and stride, generally reducing the size of the feature map, which helps to reduce the subsequent computational amount while retaining key features. In terms of the semantic level, the semantics of the feature map become more abstract, no longer just simple pixel information, but contain higher-level image features, such as the shape of finger bones, the position of joints, etc., and features related to bone development start to appear, providing more valuable data for the subsequent three layers of separable convolution and the operation of the entire bone development image recognition system. Based on the preliminary feature map obtained through two layers of standard convolution, the three layers of separable convolution further extract and refine these features. It is no longer just simple basic features such as edges and textures, but contains more representative features related to bone development, such as the maturity and growth trend of bones. For example, it may capture subtle changes in bone texture and subtle morphological differences at joints, and these features are very important for accurately judging bone development.
[0101] In one embodiment, the residual convolution network includes eight repeated three-layer residual convolutions, and each three-layer residual convolution is used for:
[0102] Performing the convolution operation of the three-layer residual convolution using the following formula:
[0103] U(m,n,o out ) = Y(m,n,o in ) + X(m,n,o in )
[0104] U” = Y(Y(Y(m,n,oin )))+X(m,n,o in )
[0105] where U is a single-layer residual convolution; Y(m,no in ) is the feature map generated after the convolution of the input module; U” is a three-layer residual convolution; X is the image information of the normalized image.
[0106] The residual convolution network of the intermediate module uses the same number of filters in three separable convolutions. This structure enables the residual or skip connection to maintain the same size as the input and be combined with the main branch through the addition operator, effectively avoiding the problem of gradient disappearance, enhancing the network's learning ability for complex features, and further extracting and strengthening features in the intermediate module by the residual convolution.
[0107] In one embodiment, as Figure 3 shown, the image processing module includes a segmentation module, a rotation module, and an alignment module;
[0108] The segmentation module is used to extract the hand region from the input image according to the image information to obtain a hand region image;
[0109] Specifically, the segmentation module uses Xception-65 as the backbone network and is equipped with the DeepLab V3plus network. It is constructed in the encoder-decoder format. The decoder performs two upsamplings through the composite operation of the separable convolutional layer and bilinear interpolation, and uses this network to extract the hand region image from the background of the input image, removing the influence of background objects on the judgment of the bone development result.
[0110] The rotation module is used for:
[0111] performing key point detection on the hand region image to obtain the key point position coordinates; among them, the key point position coordinates include the position coordinates of the thumb fingertip, the position coordinates of the middle finger fingertip, the position coordinates of the little finger fingertip, and the position coordinates of the midpoint at the bottom of the carpal region;
[0112] The key points can be automatically detected by a separable convolutional neural network regressor based on the MobileNet V1 network architecture. Among them, the last softmax activation function of the MobileNet V1 network should be replaced with a linear activation function, and the network will output the key point position coordinates in the form of an eight-element vector. Schematically, the key point regression loss function is used during training to improve the accuracy of key point detection and the precision of obtaining the key point position coordinates, where is the predicted key point position coordinates, is the true key point position coordinates:
[0113]
[0114] Based on the difference between the position coordinates of the middle finger fingertip and the position coordinates of the midpoint at the bottom of the carpal region, the rotation angle is obtained, and the hand region image is rotated according to the rotation angle to obtain an upright hand image;
[0115] When the hand X-ray image is taken, due to the different postures of the subject's hand, the angle of the hand in the image will vary. This variation will affect the accuracy of subsequent analysis and evaluation of bone features. Rotating the hand region image according to the calculated rotation angle to make it an upright hand image can unify the angle of the hand in the image and eliminate the interference factors caused by the shooting angle. When performing bone development evaluation subsequently, the model can focus more on the features of the hand bones themselves, avoiding misjudgment caused by different hand angles, thereby improving the accuracy and reliability of bone age assessment.
[0116] The alignment module is used to align the upright hand image so that the hand is located in the middle of the image to obtain a standardized image.
[0117] Perform position alignment so that the hand is in the middle of the image and keep the image size unchanged. Exemplarily, it is stipulated that the size of the standardized image is 299×299 pixels.
[0118] The segmentation module accurately extracts the hand region image, removes interference from irrelevant backgrounds such as medical equipment and labels, so that subsequent analysis only focuses on the hand bones, avoiding background information from misleading bone age assessment. The rotation module, based on key point detection and rotation angle calculation, adjusts the hand region image to an upright state, unifying the angle of the hand in the image and reducing the impact of bone morphological changes caused by shooting angle differences. The alignment module further places the hand in the middle of the image, ensuring the consistency of the hand position in different images. These operations together ensure that the images used for bone age assessment are highly consistent in the hand region, angle, and position, enabling the subsequent image recognition module to more accurately identify bone features, thereby improving the accuracy of bone development assessment.
[0119] In one embodiment, as Figure 4 shown, the segmentation module includes a feature extraction module, an ASPP module, a mask generation module, and an image vector operation module;
[0120] The feature extraction module is used to perform multi-level preliminary feature extraction on the input image to obtain a basic feature vector;
[0121] The feature extraction module uses Xception-65 as the basic architecture of the encoder to support image feature extraction. The feature extraction module adopts depthwise separable convolution, which decomposes the standard convolution into two operations: depthwise convolution and pointwise convolution. This can not only reduce the number of parameters and computational complexity, but also effectively extract the spatial and channel information of the image. Through the stacking of multiple convolutional layers, Xception-65 can gradually extract features at different levels and scales in the image, from low-level edge features and texture features to high-level semantic features, providing rich feature representations, namely basic feature vectors, for subsequent segmentation tasks. These basic feature vectors are the basis for accurately segmenting the hand region image and help the segmentation module better identify the differences between the hand and the background in the input image.
[0122] The ASPP (Atrous Spatial Pyramid Pooling) module is used to obtain differential feature vectors by sampling the features of the input image at different scales from different parallel downsampling branches with different dilation rates.
[0123] The ASPP module consists of three parallel downsampling branches with different dilation rates and a residual network. Exemplarily, the dilation rates D = 6, 12, 18. The branches with different dilation rates can sample the image under different receptive fields to capture feature information at different scales. Specifically, the branch with a large dilation rate can obtain the context information of a larger region in the image, while the branch with a small dilation rate focuses on detailed features, thus obtaining differential feature vectors. This enables the network to better handle different-sized structures and complex background situations in the hand region during segmentation. In addition, the residual network is used to solve the problem of gradient vanishing, ensuring the stability of the network during the training process and after deployment, enabling information to be effectively transmitted in the network, and further improving the effect of features and the accuracy of segmentation.
[0124] The mask generation module is used to generate a mask label with the same size as the input image based on the basic feature vector and the differential feature vector.
[0125] The basic feature vector and the differential feature vector will be fed into the subsequent network structure composed of Xception-65 and DeepLab V3plus. The network will perform a series of operations such as convolution, pooling, and activation functions based on these rich feature information of the basic feature vector and the differential feature vector to generate a mask label with the same size as the input image. In the mask label, different values are assigned to the hand region and the background region. Exemplarily, the hand region is set as the foreground and assigned a value of 1, while the background region is assigned a value of 0, thus distinguishing the hand region and the background region in the image in the form of a label. In addition, the mask generation module uses the binary cross-entropy loss function to optimize the mask label, where is a mask label generated by the network, is the real mask label:
[0126]
[0127] The accuracy of the mask generation module for automatically predicting and generating mask labels is adjusted using a function.
[0128] The image vector operation module is used to obtain the hand region image based on the mask label and the image information.
[0129] The generated mask label is element-wise multiplied with the image information, i.e., the input image vector, to separate the hand region from the background.
[0130] The feature extraction module performs multi-level preliminary feature extraction, which can obtain rich basic features from the input image, covering information from simple edges, textures to more complex shapes, providing a basis for subsequent precise segmentation. The ASPP module samples features from different scales using parallel downsampling branches with different dilation rates, and can capture the features of different-sized objects in the image. The large dilation rate branch can focus on the features of larger regions in the image, while the small dilation rate branch focuses on the details. The combination of the two enables the model to effectively capture the hand features at different scales. The mask generation module synthesizes the basic feature vector and the difference feature vector to generate the mask label, making full use of multi-faceted feature information and improving the accuracy of the mask. Finally, the image vector operation module obtains the hand region image based on the accurate mask label, and the entire process greatly improves the segmentation accuracy of the hand region and reduces mis-segmentation.
[0131] In one embodiment, obtaining the hand region image based on the mask label and the image information includes:
[0132] Using the following calculation formula, the hand region image is obtained based on the mask label and the image information:
[0133] X 1 = X mask ⊙ X
[0134] where X 1 is the hand region image; X mssk is the mask label; X is the image information.
[0135] ⊙ represents element-wise multiplication. The image information of X is mainly a pixel set. The pixels of the input image corresponding to the background area with a value of 0 in the mask label are 0, while the pixel values of the image corresponding to the hand region with a value of 1 in the mask label are retained, thereby accurately extracting the hand region image in the input image and separating the hand region from the background.
[0136] In one embodiment, the rotation angle is obtained according to the difference between the position coordinates of the middle finger fingertip and the position coordinates of the midpoint at the bottom of the standard carpal region, including:
[0137] Using the following calculation formula, the rotation angle is obtained according to the position coordinates of the middle finger fingertip and the position coordinates of the midpoint at the bottom of the carpal region:
[0138]
[0139] where KP 2 is the position coordinate of the middle finger fingertip; KP 4 is the position coordinate of the midpoint at the bottom of the carpal region; n is the ordinate of the position coordinate; m is the abscissa of the position coordinate, and θ is the rotation angle.
[0140] By calculating the difference in the ordinates of the two key points, i.e., the position coordinate KP 2 of the middle finger fingertip and the position coordinate KP 4 of the midpoint at the bottom of the carpal region, KP 2 (n) - KP 4 (n), and combining with the difference in abscissas to calculate the modulus length of the vector Using the definition of the trigonometric function relationship cosθ, the rotation angle θ is calculated. This rotation angle θ is used to rotate the hand region to the standard upright position, so that the segmented hand image has a unified orientation in subsequent processing.
[0141] In one embodiment, as Figure 5 shown, the alignment module includes a horizontal alignment module, a vertical alignment module, and a zero-padding processing module;
[0142] The horizontal alignment module is used to make the midpoint coordinate of the position coordinates of the thumb fingertip and the little finger fingertip coincide with the horizontal midpoint coordinate of the upright hand image through a rigid translation transformation;
[0143] Based on the two key points of the thumb fingertip and the little finger fingertip, since the position coordinates of the key points can be directly obtained during key point detection, the midpoint coordinate of the position coordinates of the thumb fingertip and the little finger fingertip is calculated, and the midpoint coordinate is compared with the horizontal midpoint coordinate of the upright hand image to obtain the coordinate gap. Through a rigid translation transformation, according to the coordinate gap, the midpoint coordinate is made to coincide with the horizontal midpoint coordinate of the image. Specifically, the image will be moved and adjusted along the horizontal direction to ensure that the hand is located at the center position of the image in the horizontal direction, avoiding the hand from shifting to the left or right side of the image and ensuring the symmetry of the hand in the horizontal direction.
[0144] The vertical alignment module is used to make the midpoint coordinate of the position coordinates of the middle finger fingertip and the position coordinates of the midpoint at the bottom of the carpal region coincide with the horizontal midpoint coordinate of the upright hand image through a rigid translation transformation;
[0145] Similarly, taking the fingertips of the middle finger and the midpoint at the bottom of the carpal region as two key points, the position coordinates of the key points during key point detection are obtained, the midpoint coordinates of the position coordinates of the fingertips of the middle finger and the midpoint at the bottom of the carpal region are calculated, and the midpoint coordinates are matched with the vertical midpoint of the upright hand image. Using rigid translation transformation, the midpoint coordinates are overlapped with the vertical midpoint of the upright hand image, so that the hand can also be accurately located at the center of the upright hand image in the vertical direction, preventing the hand from shifting in the vertical direction and maintaining symmetry in the vertical direction.
[0146] The zero-padding processing module is used to ensure that the standardized image is output in a set size by filling pixel points at the blank edges of the upright hand image.
[0147] After completing the horizontal and vertical alignment, since there may be blank edge regions during the image translation process, in order to make the size of the aligned image consistent with the input image, a zero-padding operation is performed on the image. Specifically, pixel points with a pixel value of 0 are filled at the blank edges of the image to ensure that the size of the image remains unchanged and to guarantee the consistency and stability of the subsequent processing flow.
[0148] The horizontal and vertical alignment modules operate respectively based on the position coordinates of the fingertips of the thumb, little finger, and middle finger and the midpoint at the bottom of the carpal region, and accurately place the hand at the center position of the image. This operation ensures the consistency of the spatial position of the hand in different images, enabling the subsequent image recognition module to more stably extract the hand bone features.
[0149] In one embodiment, the bone development result includes bone age;
[0150] The set of feature vectors is passed to the fully connected layer of the exit module to obtain the bone development result, including:
[0151] Using the following formula to generate the bone age in the bone development result:
[0152]
[0153] where R BA is the bone age; ω is the max pooling operator; V i is the output result after two-layer standard convolution of the entrance module; V j is the output result after three-layer separable convolution of the entrance module; τ is the global average pooling operator; N is the fully connected layer; X φ2 is the enhanced feature map.
[0154] X φ2The feature map generated for the intermediate module contains rich skeletal development information, including feature information related to bone age, which can be used to predict bone age. ω is the result of performing max pooling on the output of the three-layer separable convolution of the entrance module by the max pooling operator. Max pooling selects the maximum value from the local area, which can highlight the key information in the features, reduce the feature dimension while retaining important features, enhance the robustness of the model to features, and make subsequent predictions focus more on significant features. It represents the operation of multiplying the output results of two-layer standard convolution on the entrance module to construct a more representative feature representation, integrating information from different aspects, and providing more comprehensive data support for bone age prediction. Global average pooling averages the feature map in the spatial dimension, reducing the size of the feature map to 1×1×2048, obtaining a fixed-length feature vector, which is input into the fully connected layer for the final bone age prediction calculation to obtain R BA 。
[0155] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are shown in sequence according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0156] Schematically, as Figure 6 shown, a method for recognizing children's skeletal development images is applicable to the children's skeletal development image recognition system of the present application:
[0157] S601. Obtain a hand X-ray image as the input image.
[0158] S602. Analyze the input image to obtain the corresponding image information.
[0159] S603. Process the input image into a standardized image according to the image information.
[0160] S604. Input the standardized image into the recognition network model to recognize the standardized image and obtain the skeletal development result corresponding to the input image.
[0161] The above-described embodiments merely represent several implementation manners of the embodiments of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the embodiments of the application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the embodiments of the present application, several modifications and improvements can still be made, and these all fall within the protection scope of the embodiments of the present application.
Claims
1. A children's skeletal development image recognition system, characterized in that: include: An image receiving module, used for receiving an input image; the input image includes a hand X-ray image; An image information reading module, used for parsing the input image to obtain corresponding image information; the image information includes width, height, number of channels and pixel set; An image processing module, used for processing the input image into a standardized image according to the image information by using an image registration technology based on a separable convolutional neural network; The image recognition module is used to recognize the standardized image using a recognition network model to obtain the bone development result corresponding to the input image.
2. The system according to claim 1, characterized in that: The image recognition module includes an entry module, an intermediate module and an exit module; The entry module is used to perform a convolution operation on the standardized image to obtain a feature map of the feature representation of the standardized image; the convolution operation includes two layers of standard convolution and three layers of separable convolution; The intermediate module is used to extract features and enhance features of the feature map through a residual convolution network to obtain an enhanced feature map; the residual convolution network includes eight repeated three-layer residual convolutions; The export module is used to convert the enhanced feature map into a feature vector set using a global average pooling operator, and pass the feature vector set to the fully connected layer of the export module to obtain the bone development result.
3. The system according to claim 2, characterized in that The step of performing a convolution operation on the standardized image to obtain a feature map representing the features of the standardized image includes: Using the following formula, we perform two layers of standard convolution to obtain a preliminary feature representation feature map: V'=V(V(m,n,o out )) Where V is the standard convolution; K is the convolution kernel; X is the input image; m is the horizontal coordinate of the standardized image; n is the vertical coordinate of the standardized image; o in is the number of channels of the standardized image; o out is the number of channels of the output feature map; j is the horizontal offset of the convolution kernel relative to m; k is the vertical offset of the convolution kernel relative to n; V' is a two-layer standard convolution; Using the following formula, a three-layer separable convolution operation is performed on the preliminary feature representation feature map to obtain a feature map of the feature representation of the standardized image: G(m,n,o out )=P(m,n,o in )Q(m,n,o out G”=G(G(G(m,n,o out ))) Where P is the depth convolution; Q is the point-by-point convolution; o in is the number of channels of the standardized image; o out is the number of channels of the output feature map; K is the convolution kernel; X is the input image; m is the horizontal coordinate of the normalized image; n is the vertical coordinate of the normalized image; j is the horizontal offset of the convolution kernel relative to m; k is the vertical offset of the convolution kernel relative to n; G is separable convolution; G” is a three-layer separable convolution.
4. The system according to claim 2, characterized in that The residual convolution network includes eight repeated three-layer residual convolutions, each of which is used to: Use the following formula to perform the convolution operation of three layers of residual convolution: U(m,n,o out )=Y(m,n,o in )+X(m,n,o in ) U”=Y(Y(Y(m,n,o in )))+X(m,n,o in ) Where U is a single layer of residual convolution; Y(m, no in ) is the feature map generated after the convolution of the entry module; U" is the three-layer residual convolution; X is the image information of the standardized image.
5. The system according to claim 1, characterized in that: The image processing module includes a segmentation module, a rotation module and an alignment module; The segmentation module is used to extract the hand region of the input image according to the image information to obtain a hand region image; The rotation module is used for: Perform key point detection on the hand area image to obtain key point position coordinates; wherein the key point position coordinates include the position coordinates of the thumb tip, the middle finger tip, the little finger tip and the position coordinates of the midpoint of the bottom of the carpal bone area; Obtaining a rotation angle according to the difference between the position coordinates of the tip of the middle finger and the position coordinates of the midpoint of the bottom of the carpal bone region, and rotating the hand region image according to the rotation angle to obtain an upright hand image; The alignment module is used to align the upright hand image so that the hand is located in the middle of the image to obtain a standardized image.
6. The system according to claim 5, characterized in that: The segmentation module includes a feature extraction module, an ASPP module, a mask generation module and an image vector operation module; The feature extraction module is used to perform multi-level preliminary feature extraction on the input image to obtain a basic feature vector; The ASPP module is used to perform feature sampling on the input image from different scales using parallel downsampling branches with different expansion rates to obtain a difference feature vector; The mask generation module is used to generate a mask label having the same size as the input image according to the basic feature vector and the difference feature vector; The image vector operation module is used to obtain the hand area image according to the mask label and the image information.
7. The system according to claim 6, characterized in that The obtaining the hand region image according to the mask label and the image information includes: The hand region image is obtained according to the mask label and the image information using the following calculation formula: X1=X mask ⊙X Where X1 is the hand area image; X mssk is the mask label; X is the image information.
8. The system according to claim 5, characterized in that The step of obtaining the rotation angle according to the difference between the position coordinates of the tip of the middle finger and the position coordinates of the midpoint of the bottom of the standard carpal bone region includes: The rotation angle is obtained according to the position coordinates of the tip of the middle finger and the position coordinates of the midpoint of the bottom of the carpal bone region using the following calculation formula: Wherein KP2 is the position coordinate of the tip of the middle finger; KP4 is the position coordinate of the midpoint of the bottom of the carpal bone area; n is the ordinate of the position coordinate; m is the abscissa of the position coordinate, and θ is the rotation angle.
9. The system according to claim 5, characterized in that: The alignment module includes a horizontal alignment module, a vertical alignment module and a zero padding processing module; The horizontal alignment module is used to make the midpoint coordinates of the position coordinates of the thumb tip and the position coordinates of the little finger tip coincide with the horizontal midpoint coordinates of the upright hand image through rigid translation change; The vertical alignment module is used to make the midpoint coordinates of the position coordinates of the tip of the middle finger and the midpoint of the bottom of the carpal bone area coincide with the horizontal midpoint coordinates of the upright hand image through rigid translation change; The zero-filling processing module is used to ensure that the standardized image is output in a set size by filling pixel points at the edge blanks of the upright hand image.
10. The system according to claim 2, characterized in that: The skeletal development results include bone age; The feature vector set is transferred to the fully connected layer of the export module to obtain the bone development result, including: The bone age in the bone development result is generated using the following formula: Where R BA is the bone age; ω is the maximum pooling operator; V i V is the output result of the entry module after two layers of standard convolution; j is the output result of the entry module after three layers of separable convolution; τ is the global average pooling operator; N is the fully connected layer; X φ2 is the enhanced feature map.