Multi-modal feature fusion lightweight system for metacarpal and phalangeal epiphysis classification

By building a lightweight system with multimodal feature fusion, the problems of low efficiency and high misidentification rate in bone age assessment are solved, and fast and accurate bone age assessment and epiphyseal grade classification are achieved, which is suitable for bone age-assisted diagnosis in primary medical scenarios.

CN120766081APending Publication Date: 2025-10-10TONGBAN YOUKANG (HEBEI) TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510907384.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

The efficiency of bone age assessment in existing technologies is low and the recognition error rate is high. The high complexity of deep learning models leads to slow reasoning speed, which makes it difficult to meet the needs of primary medical scenarios. The misrecognition rate is high, especially in the classification of metacarpal and phalangeal epiphyseal grades.

Method used

A lightweight multimodal feature fusion system for metacarpophalangeal epiphysis is constructed. A feature fusion network architecture with multimodal group convolution and axial attention mechanism is adopted. Dynamic image enhancement and multimodal data fusion strategies are combined to design an improved loss function to train a lightweight classification model, including dynamic weighted cross entropy, anatomical center loss and feature orthogonality constraint.

Benefits of technology

It achieves rapid and accurate bone age assessment in primary medical scenarios, reduces the amount of calculation, and improves the accuracy of epiphyseal grade classification, making it suitable for clinical bone age auxiliary diagnosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120766081A_ABST
    Figure CN120766081A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and medical image interdiscipline, in particular to a metacarpal and phalangeal epiphysis classification-oriented multi-modal feature fusion lightweight system. The system comprises a construction module used for constructing a preset classification model which is a lightweight classification model; and the training module is used for taking the acquired metacarpal and phalanx images as input, taking the preset classification categories corresponding to the metacarpal and phalanx images as output, and training a preset classification model based on the improved loss function to obtain a target classification model. The lightweight classification model for classifying the metacarpal and phalanx images is small in parameter quantity and high in calculation speed. In addition, the lightweight classification model is trained by adopting an improved loss function to obtain a target classification model, so that the classification is more accurate, and the problems of low efficiency and high recognition error rate during bone age evaluation in the prior art can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of the intersection of artificial intelligence and medical imaging, and in particular to a lightweight system for multimodal feature fusion for metacarpal and phalangeal epiphysis classification. Background Art

[0002] Bone age assessment plays a central role in pediatric growth and development monitoring and the diagnosis and treatment of endocrine diseases. It provides a key basis for assessing children's growth and development and diagnosing related diseases. With advancements in medical technology and growing medical needs, the development of bone age assessment technology has become increasingly important.

[0003] Existing technologies, such as artificial intelligence (AI), for automated bone age assessment utilize deep learning and other AI technologies to attempt to rapidly diagnose bone age through computer algorithms. However, deep learning models have complex architectures and a large number of parameters, resulting in slow inference speeds and a difficulty meeting the needs of primary care. Furthermore, they have a high false positive rate, making it easy to misjudge critical levels. Summary of the Invention

[0004] The embodiment of the present invention provides a lightweight system for multimodal feature fusion for metacarpal and phalangeal epiphysis classification to solve the problems of low efficiency and high recognition error rate in bone age assessment in the prior art.

[0005] In a first aspect, an embodiment of the present invention provides a lightweight system for multimodal feature fusion for metacarpal and phalangeal epiphysis classification, comprising:

[0006] Building modules and training modules;

[0007] The construction module is used to construct a preset classification model;

[0008] The training module is used to take the acquired metacarpal and phalangeal images as input, the preset classification categories corresponding to the metacarpal and phalangeal images as output, and train the preset classification model based on the improved loss function to obtain a target classification model.

[0009] In a possible implementation, the preset classification model includes:

[0010] The input layer expands the single input channel into M input channels, so that each channel corresponds to a metacarpal epiphyseal region, and M is a positive integer greater than or equal to 2;

[0011] The convolution layer is a grouped convolution, and an axial attention layer is embedded in the MBConv module to extract the features of the long-range dependency relationship between the longitudinal and transverse directions of the epiphysis;

[0012] The feature splicing layer splices all the extracted features into a feature matrix through channels;

[0013] The fully connected layer is a double-layer fully connected classification layer.

[0014] In one possible implementation, M is set to 14 based on clinical bone age assessment requirements.

[0015] In one possible implementation, the improved loss function is

[0016]

[0017] Among them, L WCE represents the weighted cross entropy loss function, N represents the total number of samples, represents the weight coefficient of the Cth category in the tth round of training, y i,c Indicates that the i-th sample belongs to the true label of the C-th category, p i,c It represents the probability that the target classification model predicts that the i-th sample belongs to the C-th category, C represents the total number of classification categories, n c represents the number of samples in the Cth category, T max Indicates the maximum number of training rounds.

[0018] In one possible implementation, the improved loss function is

[0019]

[0020] Among them, α represents the hyperparameter, L ACL represents the anatomical center loss function, x i represents the feature vector of the i-th sample, represents the center vector of category y to which the i-th sample belongs, λ ortho represents the orthogonality constraint weight, c j represents the center vector of category j, c p represents the center vector of category p, Represents the dot product between the center vectors of different categories, used for orthogonality constraints.

[0021] In one possible implementation, the improved loss function is

[0022]

[0023] Among them, L total Represents the total loss function, β(t) represents the balance factor that decays exponentially with the training process, β0 represents the initial balance factor, and k represents the decay rate.

[0024] In a possible implementation, the method further includes: a classification module;

[0025] The classification module is used to input the metacarpal and phalangeal bone images to be classified into the target classification model and output the metacarpal and phalangeal bone classification category.

[0026] An embodiment of the present invention provides a lightweight system for multimodal feature fusion for metacarpophalangeal epiphysis classification, wherein a preset classification model is constructed through a construction module; a training module takes the acquired metacarpophalangeal images as input and the preset classification categories corresponding to the metacarpophalangeal images as output, and trains the preset classification model based on an improved loss function to obtain a target classification model. In this embodiment, a lightweight classification model is constructed for classifying metacarpophalangeal images, which has a small number of parameters and a fast calculation speed. In addition, the lightweight classification model is trained using an improved loss function to obtain a target classification model, so that the classification of metacarpophalangeal epiphysis is more accurate, and further a more accurate bone age prediction is obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0028] Figure 1 is a schematic diagram of a lightweight system for multimodal feature fusion for metacarpal and phalangeal epiphysis classification provided by an embodiment of the present invention;

[0029] Figure 2 Schematic diagram of a lightweight system for classifying metacarpal and phalangeal epiphyses using multimodal feature fusion for metacarpal and phalangeal epiphyseal classification provided by an embodiment of the present invention;

[0030] Figure 3 3 is a schematic diagram of a lightweight system for multimodal feature fusion for metacarpal and phalangeal epiphysis classification provided by another embodiment of the present invention. DETAILED DESCRIPTION

[0031] In the following description, specific details such as particular system structures and techniques are provided for purposes of illustration, not limitation, to facilitate a thorough understanding of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the present invention may be practiced in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted so as not to obscure the description of the present invention with unnecessary detail.

[0032] In order to make the purpose, technical solutions and advantages of the present invention more clear, specific embodiments will be described below with reference to the accompanying drawings.

[0033] Existing manual bone age assessments are time-consuming, labor-intensive, and highly dependent on the physician's personal experience. This is subject to significant subjectivity, limiting their accuracy and efficiency in clinical applications. While the use of artificial intelligence for automated bone age assessments improves bone age diagnosis efficiency, existing solutions still face systemic bottlenecks in clinical translation.

[0034] The high complexity of the deep learning model makes it difficult for the inference speed to meet the needs of primary medical scenarios, and the design paradigm of directly migrating natural image processing models ignores the characteristics of anatomical structure classification, resulting in poor classification results and high false positive rates. In addition, the long-tail distribution characteristics of the developmental levels of various parts of the metacarpophalangeal bones, such as the low proportion of high-grade samples in each epiphyseal bone, make the traditional cross-entropy loss function highly sensitive to class imbalance. Even with the use of weighted strategies, it is still impossible to effectively distinguish morphologically similar levels, resulting in false positives in key levels. For example, the difference between the epiphyseal lines of ulna level 7(2) and 7(3) is not obvious, which may lead to false positives in key levels. At the data level, fixed parameter image enhancement methods are also difficult to adapt to the dynamic characteristics of multi-source heterogeneous and non-uniform quality X-rays, such as the coexistence of low-contrast images from computed radiography (CR) equipment and overexposed images from direct digital radiography (DR). Such preprocessing biases can increase the mean absolute error of the model, resulting in poor generalization ability. All of these are hindering the automatic bone age recognition algorithm based on artificial intelligence solutions from becoming a core tool for diagnosis and treatment.

[0035] In order to solve the above problems, the embodiment of the present invention provides a lightweight system for multimodal feature fusion for metacarpal and phalangeal epiphysis classification. Figure 1 The following is a schematic diagram of a lightweight system for multimodal feature fusion for metacarpal and phalangeal epiphysis classification.

[0036] A lightweight system for multimodal feature fusion for metacarpal and phalangeal epiphysis classification, comprising: a construction module 11 and a training module 12;

[0037] A construction module 11 is used to construct a preset classification model, which is a lightweight classification model;

[0038] The training module 12 is used to take the acquired metacarpal and phalangeal images as input, and the preset classification categories corresponding to the metacarpal and phalangeal images as output, and train the preset classification model based on the improved loss function to obtain a target classification model.

[0039] In this embodiment, to meet the processing speed and load requirements of primary care scenarios, we set a lightweight classification model as the default model. This model is based on an improvement of EfficientNetB3 using multimodal group convolution and axial attention fusion. Combining dynamic image enhancement with multimodal data fusion strategies, it achieves the developmental grade classification of the 13 metacarpal epiphyses of pediatric patients, which is suitable for clinical bone age-assisted diagnosis scenarios. Therefore, this embodiment adopts a feature fusion network architecture of multimodal group convolution and axial attention mechanism to target the overall correlation characteristics of epiphyseal development.

[0040] In one embodiment, the preset classification model may include:

[0041] The input layer expands the single input channel into M input channels, so that each channel corresponds to a metacarpal epiphyseal region;

[0042] The convolution layer is a grouped convolution, and an axial attention layer is embedded in the MBConv module to extract the features of the long-distance longitudinal and transverse dependencies of the epiphysis. This effectively retains and captures the long-distance longitudinal and transverse dependencies of the epiphysis, improving the ability to extract epiphyseal morphological features.

[0043] The feature splicing layer splices all the extracted features into a feature matrix through channels;

[0044] The fully connected layer is a double-layer fully connected classification layer.

[0045] Optionally, in this embodiment, based on the anatomically grouped input design, we expand the input channels of the input layer so that each channel corresponds to a specific epiphyseal region. For example, the original single-channel input is expanded to an M-channel anatomically grouped input, where M can be the number of channels. It should be noted that the value of M is not limited to 14; it can be a positive integer greater than or equal to 2.

[0046] In this embodiment, based on the clinical needs of bone age assessment, M can be set to 14. Generally, according to the Chinese National Standard for Bone Age Assessment TY / T 3001-2006 and the synergy of biological growth and development, 14 metacarpal bones are required for bone age assessment. Among them, 13 metacarpal and phalangeal bone images in the epiphyseal region, such as the radius, ulna, metacarpal phalanges I / III / V, are used for bone age recognition. Then, one full metacarpal bone is added, for a total of 14 feature images. Each channel corresponds to a specific epiphyseal region, such as the specific epiphyseal region is the radius, ulna, metacarpal phalanges I / III / V, etc. The size of the input metacarpal and phalangeal bone images is uniformly 320*320 pixels. See Figure 2 The schematic diagram shown is a schematic diagram of using a lightweight system for multimodal feature fusion for metacarpal phalangeal epiphysis classification to classify metacarpal phalanges. 14 metacarpal phalangeal images are input into the target classification model, and each metacarpal phalangeal image corresponds to one channel.

[0047] In the grouped convolution reconstruction, the first convolution layer is modified to a grouped convolution with 14 groups and 1 channel per group. The initial convolution kernel size can be 3*3 with a stride of 2. During the convolution operation, each group independently extracts spatial features. Parameters within each group are shared, but parameters between groups are independent. This improvement to the convolution layer reduces the computational complexity to 1 / 14 of the original single-channel input. The theoretical computational complexity was originally 3*3*3*32=864, but is now 14*(3*3*3*32)=1296.

[0048] The axial spatial attention enhancement module decomposes the global attention of a 2D image into long-range dependent attention along the height and width axes, enabling better localization of epiphyseal boundaries. Embedding the axial attention layer in the MBConv module of EfficientNetB3 extracts long-range texture features along the long axis of the epiphysis (alternating between the X and Y axes), improving the model's feature extraction capabilities.

[0049] The main design of the non-fusion feature splicing and classification layer optimization part is to maintain the independence of 14 groups of features at the end of the backbone network and generate a 1000*14-dimensional feature matrix through channel splicing, rather than traditional multimodal fusion methods such as addition or averaging.

[0050] Then build a two-layer fully connected classification network, the classification structure is as follows:

[0051] Classifier = FC(1000×14→512)→Dropout(0.3)→BN→RELU→FC(512→C). Dropout represents random masking of neurons and is set to 0.3 in this example. BN stands for BatchNorm, which represents normalization and accelerates model training convergence. RELU represents the activation function, which is connected after the convolutional layer to introduce nonlinearity to the network, effectively solving the vanishing gradient problem.

[0052] In this embodiment, a lightweight classification network for holistic modeling of epiphyseal development is constructed. It not only retains the independent feature expression of each epiphysis, but also expands the single channel to 14 channels, so that each channel corresponds to a specific area of ​​the metacarpal epiphysis. Each channel extracts the corresponding features, and each group of model parameters is shared only within this group. Taking into account the synergy of growth and development, the 14 groups of feature results are simultaneously taken into consideration in the fully connected layer as the basis for classification. An axial attention layer is also embedded, which allows the extraction of long-range texture features along the long axis of the metacarpal epiphysis (alternating between the X and Y axes); the experimental results are significantly higher than the original EfficientNetB3 architecture.

[0053] In this embodiment, after constructing a preset classification network, the parameters within the network are trained to classify metacarpal and phalangeal epiphyses. During training of the preset classification network, a loss function is employed to modify the network. Therefore, a ternary hybrid loss function is designed that integrates dynamic weighted cross entropy, anatomical center loss, and feature orthogonality constraints. This hybrid loss function embeds anatomical prior knowledge into the feature learning process, mathematically achieving physical consistency modeling of epiphyseal developmental patterns. Its mathematical form and optimization mechanism are as follows: 1. Dynamic category-aware weighted cross entropy loss; 2. Anatomical center loss; 3. Adaptive weighted hybrid mechanism.

[0054] Optionally, the improved loss function can be a dynamic category-aware weighted cross entropy loss function, whose mathematical expression is

[0055]

[0056] Among them, L WCE represents the weighted cross entropy loss function, N represents the total number of samples, represents the weight coefficient of the Cth category in the tth round of training, which is dynamically adjusted with the rounds of training to alleviate the long-tail distribution problem. i,c Indicates that the i-th sample belongs to the true label of the C-th category, p i,c It represents the probability that the target classification model predicts that the i-th sample belongs to the C-th category, and C represents the total number of classification categories.

[0057] in, n c represents the number of samples in the Cth category, T max represents the maximum training round, that is, the maximum number of iterations; α represents a hyperparameter, and its value range is 0-1. In this embodiment, it can be set to 0.01.

[0058] In the mathematical expression of the above dynamic category-aware weighted cross entropy loss function, compared with the cross entropy loss function in the prior art, exist In the expression of c The square root of 1 is taken during the calculation process. The reason for taking the reciprocal is to make the categories with fewer samples in the metacarpal and phalangeal image dataset have a larger weight. The reason for taking the square root is to prevent the weight of the categories with a large number of samples in the metacarpal and phalangeal image dataset from being too small. A term that increases with the number of training rounds is also added at the end. In this way, at the end of model training, the extremely small categories in the epiphyseal classification will receive higher weights, thereby improving the classification accuracy of the target classification model on small categories.

[0059] The above design of weight coefficients can make rare categories such as level 0 of radius obtain higher weights in the later stage of training.

[0060] In one embodiment, the improved loss function may be an anatomical center loss function, whose calculation expression is:

[0061]

[0062] Among them, L ACL represents the anatomical center loss function, x i represents the feature vector of the i-th sample, represents the center vector of category y to which the i-th sample belongs, λ ortho represents the orthogonality constraint weight, fixed at 0.01, c j represents the center vector of category j, c p represents the center vector of category p, Represents the dot product between the center vectors of different categories, used for orthogonality constraints.

[0063] Each category center is initialized with the standard anatomical features of the corresponding epiphysis and iteratively updated during training.

[0064] The difference between the anatomical center loss function in this embodiment and the center loss function in the prior art is that the original 1 / 2 is modified to 1 / 2N, where N is the total number of samples. The added 1 / N is used to average the sum of the squared distances of all samples. Then a term for orthogonality constraint is added at the end. λ ortho It is a regularization parameter used to control the strength of the orthogonality constraint. The orthogonality of the following two vectors is to reduce the overlap between different categories, making the feature representation of each category more independent, reducing the feature similarity of adjacent levels, and thus improving classification accuracy.

[0065] In one embodiment, the loss function is set to

[0066] Among them, L WCE , L ACL The definitions in the above embodiments may be adopted, for example and We can also use the corresponding parameter definition in the prior art, and we will L WCE and L ACL Combined into the total loss function.

[0067] Among them, L totalWe add β(t) to the total loss function. β(t) represents a balancing factor that decays exponentially with training, balancing classification loss with feature constraints. This design results in the weight of the anatomical center loss function gradually decreasing with each training round, allowing the model to make more informed use of the weighted cross-entropy loss function in the later stages of classification. β0 represents the initial balancing factor, fixed at 0.5, and k represents the decay rate, fixed at 0.005.

[0068] In one embodiment, see Figure 3 As shown, after the preset classification model is trained to obtain the target classification model, the lightweight system for multimodal feature fusion for metacarpal epiphysis classification may further include: a classification module 13;

[0069] The classification module 13 is used to input the metacarpal and phalangeal bone images to be classified into the target classification model and output the metacarpal and phalangeal bone classification category.

[0070] Based on the trained target classification model, the metacarpophalangeal epiphyseal grade classification is performed on the metacarpophalangeal target images to be classified. The grade of each epiphyseal can be obtained. Combined with the corresponding scoring method, the grade score is obtained and then the bone age prediction value is obtained to assist clinical bone age assessment.

[0071] In one embodiment, see Figure 2 As shown, in order to further improve the classification accuracy of the metacarpal phalangeal epiphysis, the metacarpal phalangeal images to be classified and the metacarpal phalangeal images used in training the preset classification model can be adaptively enhanced to improve the image clarity and highlight the epiphyseal features.

[0072] The embodiment of the present invention constructs a preset classification model through a construction module, and the preset classification model is a lightweight classification model; then a training module is used to take the acquired metacarpophalangeal images as input, and the preset classification categories corresponding to the metacarpophalangeal images as output, and the preset classification model is trained based on an improved loss function to obtain a target classification model. In this embodiment, a lightweight classification model is constructed for classifying metacarpophalangeal images, which has a small number of parameters and a fast calculation speed. In addition, the lightweight classification model is trained using an improved loss function to obtain a target classification model, which makes the classification of the metacarpophalangeal epiphysis more accurate and further accurately predicts bone age.

[0073] In this embodiment, a preset classification model is designed. Aiming at the overall correlation characteristics of epiphyseal development, a feature fusion network architecture based on multimodal group convolution and axial attention mechanism is proposed. The 13 metacarpal phalangeal epiphyses and the entire metacarpal phalangeal bone are input into the network at the same time. When performing epiphyseal grade classification, the basic principle of coordination of human growth and development is fully considered. The axial spatial attention mechanism and the anatomical-based mixed loss function are added to optimize the model. This allows the trained target classification model to improve the accuracy of epiphyseal grading when performing epiphyseal grading, thereby assisting in clinical bone age prediction applications.

[0074] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0075] The embodiments described above are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A lightweight system for multimodal feature fusion for metacarpal and phalangeal epiphysis classification, characterized by: include: Building blocks and training blocks; The construction module is used to construct a preset classification model, and the preset classification model is a lightweight classification model; The training module is used to take the acquired metacarpal and phalangeal images as input, the preset classification categories corresponding to the metacarpal and phalangeal images as output, and train the preset classification model based on the improved loss function to obtain a target classification model.

2. The lightweight system for multimodal feature fusion for metacarpal and phalangeal epiphysis classification according to claim 1 is characterized in that: The preset classification model includes: The input layer expands the single input channel into M input channels, so that each channel corresponds to a metacarpal epiphyseal region, and M is a positive integer greater than or equal to 2; The convolution layer is a grouped convolution, and an axial attention layer is embedded in the MBConv module to extract the features of the long-range dependency relationship between the longitudinal and transverse directions of the epiphysis; The feature splicing layer splices all the extracted features into a feature matrix through channels; The fully connected layer is a double-layer fully connected classification layer.

3. The lightweight system for multimodal feature fusion for metacarpal and phalangeal epiphysis classification according to claim 2 is characterized in that: The value of M is 14 based on the clinical bone age assessment requirements.

4. The lightweight system for multimodal feature fusion for metacarpal and phalangeal epiphysis classification according to any one of claims 1 to 3, characterized in that: The improved loss function is Among them, L WCE represents the weighted cross entropy loss function, N represents the total number of samples, represents the weight coefficient of the Cth category in the tth round of training, y i,c Indicates that the i-th sample belongs to the true label of the C-th category, p i,c It represents the probability that the target classification model predicts that the i-th sample belongs to the C-th category, C represents the total number of classification categories, n c represents the number of samples in the Cth category, T max Indicates the maximum number of training rounds.

5. The lightweight system for multimodal feature fusion for metacarpal and phalangeal epiphysis classification according to claim 4 is characterized in that: The improved loss function is Among them, α represents the hyperparameter, L ACL represents the anatomical center loss function, x i represents the feature vector of the i-th sample, represents the center vector of category y to which the i-th sample belongs, λ ortho represents the orthogonality constraint weight, c j represents the center vector of category j, c p represents the center vector of category p, Represents the dot product between the center vectors of different categories, used for orthogonality constraints.

6. The lightweight system for multimodal feature fusion for metacarpal and phalangeal epiphysis classification according to claim 5, characterized in that: The improved loss function is Among them, L total Represents the total loss function, β(t) represents the balance factor that decays exponentially with the training process, β0 represents the initial balance factor, and k represents the decay rate.

7. The lightweight system for multimodal feature fusion for metacarpal and phalangeal epiphysis classification according to claim 1, characterized in that: Also includes: Classification module; The classification module is used to input the metacarpal and phalangeal bone images to be classified into the target classification model and output the metacarpal and phalangeal bone classification category.