Intelligent image typing method and device for children's femoral neck fracture based on image-text fusion

By constructing a deep learning model that integrates image and text features, the problem of relying on doctors' experience and insufficient data in the classification of femoral neck fractures in children has been solved. This has enabled intelligent and accurate classification of femoral neck fractures in children, improving diagnostic accuracy and robustness, and reducing the risk of complications.

CN122434862APending Publication Date: 2026-07-21SHENZHEN TRADITIONAL CHINESE MEDICINE HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHENZHEN TRADITIONAL CHINESE MEDICINE HOSPITAL
Filing Date
2026-04-23
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing technologies rely on doctors' experience in classifying femoral neck fractures in children, making it difficult for inexperienced doctors to make accurate assessments. Furthermore, traditional deep learning models suffer from insufficient accuracy due to limited data sample size and single feature capture dimension.

Method used

A deep learning model for image-text fusion was constructed. By combining image and text features, a convolutional neural network was used to extract image features and a category semantic prototype library was applied to extract text features. Feature alignment and fusion were performed, and the deep learning model was trained to achieve intelligent and accurate classification of femoral neck fractures in children.

Benefits of technology

It improves the accuracy and robustness of imaging classification of femoral neck fractures in children, breaks through the dependence on doctors' experience, helps inexperienced doctors to quickly standardize diagnosis and treatment, reduces the risk of complications, and provides a new approach to building intelligent diagnosis and treatment models for rare diseases.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122434862A_ABST
    Figure CN122434862A_ABST
Patent Text Reader

Abstract

The application provides a kind of image intelligent typing method and device for children's femoral neck fracture based on image-text fusion, relating to medical image detection technical field.The invention content includes image-text data set construction and preprocessing, image-text information fusion, multi-modal model training, cascade diagnosis and result output, model test and evaluation steps;The application extracts data set image features through VGG16 model;Apply CLIP encoder to extract data set text features;Map the output image feature vector and text feature vector to the same semantic space to complete image-text fusion;Apply training module to train the model and generate structured diagnostic report;Calculate the model accuracy, recall rate and F1 score based on the test set to evaluate the model performance;The application captures rich data features from the perspective of vision and text, makes up for the lack of single-modal model accuracy caused by limited data features, and further realizes rapid and accurate assessment of children's femoral neck fracture typing, with significant clinical significance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of medical imaging technology, and more specifically, to a method and device for intelligent classification of femoral neck fracture images in children based on image-text fusion. Background Technology

[0002] Femoral neck fracture in children is a serious orthopedic trauma with a high incidence of various postoperative complications, including avascular necrosis of the femoral head (approximately 20%-25%) and nonunion (approximately 10%-24%), resulting in extremely poor prognosis. Because the risk of complications and subsequent treatment plans vary depending on the fracture type, accurate fracture classification is crucial for guiding standardized treatment plans and reducing the risk of complications.

[0003] Currently, the classification proposed by Wang et al. is widely used in clinical practice to classify femoral neck fractures in children based on imaging. Specifically, Wang et al.'s classification divides femoral neck fractures in children into Type I (no displacement), Type II (anterior displacement), Type III (posterior displacement), and Type IV (comminuted fracture of the medial and posterior column of the neck) based on the direction of displacement of the distal fragment relative to the proximal fragment in the sagittal view and the stability of the medial and posterior column of the neck. However, this method is highly dependent on the experience of physicians in clinical application. Due to the uneven distribution of high-quality medical resources, experienced physicians often cannot promptly guide less experienced physicians to accurately assess the classification of femoral neck fractures in children, which is detrimental to the rapid and standardized diagnosis and treatment of this type of fracture.

[0004] With the development of deep learning technology, image recognition technologies, represented by convolutional neural networks, have made significant progress in the field of medical image classification. However, most of these technologies are based on single-modal data features to complete the task output, requiring a large amount of data samples and richness of features, and suffer from the drawback of capturing only a single dimension of data features. For femoral neck fractures in children, which have a low incidence rate, the construction of traditional deep learning models faces significant challenges due to the limited sample size.

[0005] The strategy of building deep learning models based on image-text fusion offers a new approach to solving the above problems. Combining image information with corresponding textual feature descriptions can enrich image features from multiple perspectives, making up for the shortcomings of capturing image features solely based on the visual dimension, and is expected to improve the accuracy and robustness of the model. Currently, although some studies have attempted to apply multimodal data to medical image classification, there are no research reports on building a deep learning model for intelligent and accurate assessment of femoral neck fracture image classification in children based on image-text fusion.

[0006] Therefore, this invention aims to provide an intelligent classification method and device for pediatric femoral neck fractures based on image-text fusion. By constructing an image-text dataset of children with different subtypes of femoral neck fractures combined with corresponding text feature descriptions, a convolutional neural network is applied to extract image features and a category semantic prototype library is applied to extract text features. After image-text feature alignment and fusion, a deep learning model is trained to achieve intelligent and accurate classification of pediatric femoral neck fractures, thereby facilitating the rapid and standardized diagnosis and treatment of pediatric femoral neck fractures. Summary of the Invention

[0007] The purpose of this application is to provide a method and device for intelligent classification of femoral neck fracture images in children based on image-text fusion. It constructs an image-text dataset of children with different subtypes of femoral neck fractures combined with corresponding text feature descriptions. It applies a convolutional neural network to extract image features from the dataset and applies a category semantic prototype library to extract text features from the dataset. After image-text feature alignment and fusion, a deep learning model is trained to fully capture the rich features of a limited sample size, enabling the model to distinguish differences in image features from multiple dimensions, thereby achieving intelligent and accurate assessment of femoral neck fracture image classification in children.

[0008] This application provides an intelligent classification method for pediatric femoral neck fracture images based on image-text fusion, including the following steps:

[0009] Construct a graphic-text dataset of images of children with different subtypes of femoral neck fractures in anteroposterior and lateral views, combined with corresponding textual feature descriptions;

[0010] The image feature extraction module is used to extract image features from the image-text dataset;

[0011] The application category semantic prototype library is used to extract text features from image and text datasets;

[0012] The application feature alignment and classification module fuses image and text features and outputs classification probabilities.

[0013] The model is trained using the application training module;

[0014] The application diagnostic module and report generation module generate structured diagnostic reports.

[0015] Furthermore, the construction of the image-text dataset of children with different subtypes of femoral neck fractures in anteroposterior and lateral views, combined with corresponding text feature descriptions, includes:

[0016] A dataset of images of children with different subtypes of femoral neck fractures was constructed, combined with corresponding textual descriptions. Femoral neck fractures were divided into two subtypes based on whether the medial column of the neck was comminuted. The standardized textual descriptions for the different subtypes were "Anterior X-ray shows femoral neck fracture, no comminuted bone fragments in the medial column of the fracture end" and "Anterior X-ray shows femoral neck fracture, comminuted bone fragments are visible in the medial column of the fracture end".

[0017] A pictorial dataset combining lateral X-ray images of children with different subtypes of femoral neck fractures with corresponding textual feature descriptions was constructed. Based on the displacement direction of the distal fracture fragment relative to the proximal fracture fragment in the sagittal view and the stability of the posterior column of the neck, femoral neck fractures were divided into four subtypes. The standardized textual descriptions corresponding to different subtypes are as follows: "Lateral X-ray shows femoral neck fracture, with no significant displacement of the distal fracture fragment relative to the proximal fracture fragment", "Lateral X-ray shows femoral neck fracture, with the distal fracture fragment displaced anteriorly relative to the proximal fracture fragment", "Lateral X-ray shows femoral neck fracture, with the distal fracture fragment displaced posteriorly relative to the proximal fracture fragment", and "Lateral X-ray shows femoral neck fracture, with comminuted bone fragments visible in the posterior column of the fracture end".

[0018] Furthermore, the image feature extraction module extracts image features from the image-text dataset, and the specific steps are as follows:

[0019] A pre-trained convolutional neural network is used as the backbone network, and its last fully connected layer is removed to output image feature vectors; among them, a pre-trained VGG16 model is used to extract image features from the dataset.

[0020] Furthermore, the application category semantic prototype library extracts text features from the image-text dataset, specifically through the following steps:

[0021] Using a pre-trained CLIP model text encoder, the standardized descriptive text corresponding to each femoral neck fracture subtype in the anteroposterior or lateral views is encoded to obtain the category semantic prototype vector.

[0022] Furthermore, the application feature alignment and classification module fuses image and text features and outputs classification probabilities. The specific steps are as follows:

[0023] Image feature vectors and category semantic prototype vectors are mapped to the same shared semantic space through learnable linear projection layers and then normalized.

[0024] Calculate the cosine similarity between the normalized image feature vector and the semantic prototype vector of each category, and multiply it by a learnable scaling factor to obtain the classification logical value for each category.

[0025] By maximizing the cosine similarity between image features and their corresponding text prototype vectors, semantic alignment and fusion of cross-modal features are achieved.

[0026] The classification probability is output using the softmax function.

[0027] Furthermore, the shared semantic space is set to 512 dimensions; the learnable scaling factor is initialized to log(1 / 0.07) and optimized along with the network parameters during training.

[0028] Furthermore, the application training module trains the model, specifically through the following steps:

[0029] The parameters of the image feature extraction module, the learnable linear projection layer, and the learnable scaling factor are optimized using the backpropagation algorithm based on the weighted cross-entropy loss function.

[0030] Furthermore, the application diagnostic module and report generation module generate a structured diagnostic report, specifically through the following steps:

[0031] The diagnostic module outputs two-class classification results for anteroposterior view images of femoral neck fractures in children and four-class classification results for lateral view images of femoral neck fractures in children.

[0032] The application report generation module integrates the corresponding fracture subtypes from the above anteroposterior and lateral images and generates a structured diagnostic report.

[0033] Based on the same inventive concept, this application also provides an intelligent classification device for pediatric femoral neck fracture images based on image-text fusion, including a memory and a processor; the memory is used to store instructions executable by the processor, and the processor is used to call the program instructions stored in the memory to realize the intelligent classification steps for pediatric femoral neck fracture images based on image-text fusion as described above.

[0034] The beneficial effects of this invention are:

[0035] This invention provides an intelligent image-based classification method and device for pediatric femoral neck fractures, based on image-text fusion. By constructing image-text datasets combining images of children with different subtypes of femoral neck fractures with corresponding textual feature descriptions, the model can capture rich image features from both visual and textual perspectives. Furthermore, it can accurately distinguish differences in image features based on multi-dimensional features, overcoming the limitations of traditional unimodal models that suffer from low accuracy due to the single feature dimension of the data captured. The classification method and device provided by this invention help inexperienced physicians quickly and accurately assess the classification of pediatric femoral neck fractures, breaking through the reliance on physician experience in traditional fracture classification methods. This facilitates rapid and standardized diagnosis and treatment for children with femoral neck fractures, thereby reducing the risk of subsequent serious complications, and its clinical significance is significant.

[0036] Furthermore, this invention trains a deep learning model using an image-text dataset, enabling the model output to correspond to textual feature descriptions of image classification results. This overcomes the limitation of traditional unimodal models, which can only output image classification probabilities, helping doctors better understand the model's evaluation criteria and apply its output in clinical practice. In addition, the implementation strategy of this invention will provide new ideas for constructing intelligent diagnosis and treatment models for other rare diseases, and has broad application prospects. Attached Figure Description

[0037] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0038] Figure 1 The above is an embodiment of the method for classifying anteroposterior images of femoral neck fractures in children, where PFNF is an abbreviation for femoral neck fracture in children.

[0039] Figure 2 The following is an embodiment of the present invention: a method for classifying lateral femoral neck fracture images in children. PFNF is an abbreviation for femoral neck fracture in children, ANT is an abbreviation for anterior side, and POST is an abbreviation for posterior side.

[0040] Figure 3 This is a general flowchart illustrating the construction of intelligent classification models for unimodal and multimodal pediatric femoral neck fracture images in embodiments of the present invention. Detailed Implementation

[0041] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. The components of the embodiments of this application described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0042] Therefore, the following detailed description of the embodiments of this application provided in the accompanying drawings is not intended to limit the scope of the claimed application, but merely to illustrate selected embodiments of the application. All other embodiments obtained by those skilled in the art based on the embodiments of this application without inventive effort are within the scope of protection of this application.

[0043] The following are specific examples.

[0044] This embodiment provides an intelligent classification method for pediatric femoral neck fracture images based on image-text fusion, including image dataset construction and preprocessing, category semantic prototype library construction, multimodal fusion classification model design, model training, cascade diagnosis and result output, and model testing and evaluation steps.

[0045] The specific steps are as follows:

[0046] 1. Image Dataset Construction and Preprocessing

[0047] This embodiment collected a large number of anteroposterior and lateral hip X-rays of children with femoral neck fractures at the time of injury from multiple hospital imaging systems. Subsequently, two senior pediatric orthopedic surgeons with more than 20 years of experience in diagnosing and treating pediatric hip fractures assessed the fracture subtypes corresponding to the above X-rays according to the classification method for pediatric femoral neck fractures proposed by Wang et al. The classification proposed by Wang et al. divides fractures into different subtypes based on the direction of distal displacement in the sagittal plane and the stability of the medial posterior column of the neck. Specifically, these include no distal displacement in the sagittal plane, anterior displacement of the distal fragment in the sagittal plane, posterior displacement of the distal fragment in the sagittal plane, and comminuted fracture of the medial posterior column of the neck. Since comminuted fractures of the medial posterior column of the neck require anteroposterior and lateral imaging for diagnosis, the two senior surgeons needed to assess whether a comminuted fracture of the medial column of the neck was present on the anteroposterior plane. Figure 1 ), and assess the direction of distal fracture displacement and the stability of the posterior column of the neck on lateral imaging ( Figure 2 ).

[0048] X-ray images with consistent assessments from two senior physicians were used in the next stage of the study. X-ray images with inconsistent assessments from senior physicians were excluded. Ultimately, a total of 711 anteroposterior (AP) X-ray images and 674 lateral X-ray images were included. Among the AP X-ray images, 576 images showed no comminuted fracture of the medial column of the neck (“AP Category I”), and 135 images showed a comminuted fracture of the medial column of the neck (“AP Category II”). Among the lateral X-ray images, 132 images showed no distal displacement of the sagittal fracture (“Lateral Category I”), 329 images showed anterior displacement of the distal sagittal fracture (“Lateral Category II”), 91 images showed posterior displacement of the distal sagittal fracture (“Lateral Category III”), and 122 images showed a comminuted fracture of the posterior column of the neck (“Lateral Category IV”).

[0049] The dataset was divided into training and test sets in approximately a 4:1 ratio. For the anteroposterior images, the training set contained 569 images (462 for "anteroposterior Class I" and 107 for "anteroposterior Class II"), while the test set contained 142 images (114 for "anteroposterior Class I" and 28 for "anteroposterior Class II"). For the lateral images, the training set contained 538 images (107 for "lateral Class I", 260 for "lateral Class II", 72 for "lateral Class III", and 99 for "lateral Class IV"), while the test set contained 136 images (25 for "lateral Class I", 69 for "lateral Class II", 19 for "lateral Class III", and 23 for "lateral Class IV").

[0050] Using Photoshop, all X-ray images with clearly defined fracture types were uniformly adjusted to 224×224 pixels and normalized (mean = [0.485, 0.456, 0.406], standard deviation = [0.229, 0.224, 0.225]). The effective information from the X-ray images was then cropped using a rectangular clipping tool in Photoshop for further research. The effective information from the X-ray images only needed to include the intact acetabulum, fracture ends, and proximal femur on the affected side.

[0051] During the model training phase, data augmentation strategies are employed, including random horizontal flipping, random vertical flipping, and color jittering (brightness, contrast, and saturation perturbations) to enhance the information in the X-ray data. During the model testing phase, only X-ray size adjustment and normalization are performed.

[0052] 2. Construction of Category Semantic Prototype Library

[0053] The imaging features of different types of fractures on anteroposterior and lateral views are described as standardized text. To facilitate accurate understanding by the CLIP model's text encoder, the descriptions avoid directly using category names and instead describe their radiographic features.

[0054] 2.1 The standardized textual descriptions of different types of fracture characteristics on anteroposterior images are as follows:

[0055] Orthopedic Class I: A pediatric femoral neck fracture without a comminuted medialcolumn on anteroposterior X-ray of the hip joint.

[0056] Orthopedic Class II: A pediatric femoral neck fracture with a comminuted medialcolumn on anteroposterior X-ray of the hip joint.

[0057] 2.2 The standardized textual descriptions of different types of fracture characteristics on lateral images are as follows:

[0058] Lateral Class I: This is a lateral X-ray of a child's femur, showing discontinuity in the cortical bone of the femoral neck, indicating a femoral neck fracture, with no obvious displacement of the distal fracture fragment relative to the proximal one.

[0059] Lateral Class II: This is a lateral X-ray of a child's femur, showing discontinuity in the cortical bone of the femoral neck, indicating a femoral neck fracture, with the distal fracture fragment displaced anteriorly relative to the proximal one.

[0060] Lateral Class III: This is a lateral X-ray of a child's femur, showing discontinuity in the cortical bone of the femoral neck, indicating a femoral neck fracture, with the distal fracture fragment displaced posteriorly relative to the proximal one.

[0061] Lateral Class IV: This image is a lateral X-ray of a child's femur, showing discontinuity in the cortical bone of the femoral neck, indicating a femoral neck fracture with a comminuted posterior column.

[0062] 2.3 Text Feature Extraction

[0063] The above text description is encoded using a pre-trained CLIP model (ViT-B / 32) text encoder to obtain the category semantic prototype vector. Where C is the number of categories (C=2 for binary classification, C=4 for four-class classification), and D is the encoding dimension (512). Subsequently, the encoder is frozen, and the prototype vector is registered in the model as a buffer, without participating in gradient updates.

[0064] 3. Design of Multimodal Fusion Classification Model

[0065] The core of this invention is a multimodal fusion classifier, the structure of which is as follows: Figure 3 As shown, the model comprises three main parts: an image feature extractor, a feature alignment mapping module, and a similarity calculation and classification module. This embodiment includes two tasks: a binary classification task for frontal images and a four-class classification task for lateral images. In each image classification task, the image features are extracted using an independent image feature extractor based on the currently mainstream VGG16 single-modal model.

[0066] 3.1 Image Feature Extraction

[0067] For each subtype of image dataset, a VGG16 model network pre-trained on ImageNet is applied as an image feature extractor. The final classifier of this network is removed, allowing the model to output high-dimensional features. The feature map output by the VGG16 model network convolutional architecture is then subjected to adaptive average pooling (AdaptiveAvgPool2d) and flattening to obtain a fixed-dimensional feature vector.

[0068] During model training, some or all of the parameters of the image encoder can be fine-tuned as needed.

[0069] 3.2 Feature Alignment Mapping

[0070] Image feature vectors and text prototype vectors are mapped to the same shared semantic space (dimension) through a learnable linear projection layer. ), and perform L2 normalization:

[0071]

[0072] in , These are the parameters for the image projection layer; , For text projection layer parameters; This indicates L2 normalization, i.e. .

[0073] 3.3 Similarity Calculation and Classification

[0074] Calculate the cosine similarity between the normalized image feature vector and each text prototype vector, and multiply it by a learnable scaling factor. This yields the logits for each category:

[0075]

[0076] in This represents the inner product (since the vectors have been normalized, it is equivalent to cosine similarity). As a learnable scalar, initialized to The final classification probability is given by the softmax function:

[0077]

[0078] 4. Model Training

[0079] The frontal binary classification model and the lateral four-class classification model were trained separately. The training strategy is as follows:

[0080] Loss function: Weighted cross-entropy loss is used to mitigate class imbalance. The weights are calculated and normalized based on the reciprocal of the number of samples from each class in the training set. Let the training set... The number of samples in each class is Then the weight The weighted cross-entropy loss is:

[0081]

[0082] in For batch size, For the first The true label of each sample.

[0083] Optimizer: Stochastic Gradient Descent (SGD), momentum 0.9, weight decay 5e-4.

[0084] Learning rate: The initial learning rate is set to 5e-3. The StepLR scheduler is used, and the learning rate is multiplied by 0.5 every 100 epochs.

[0085] Training epochs: 100 epochs.

[0086] Batch size: 16 (can be adjusted according to video memory).

[0087] Other: Only update the image encoder, projection layer parameters, and learnable scaling factor. The CLIP text encoder is kept frozen to prevent the model from overfitting.

[0088] During training, the accuracy is evaluated on the validation set after each epoch, and the optimal model parameters are saved.

[0089] 5. Cascaded Diagnosis and Result Output

[0090] In practical applications, for each child with a femoral neck fracture, anteroposterior and lateral X-rays of the hip joint are first obtained. Then, these X-rays are input into a binary classification task model (anteroposterior view) and a four-class classification task model (lateral view). Next, the outputs of these two models are converted into corresponding textual feature descriptions and combined into a structured diagnostic report. An example of a structured diagnostic report is: "Anterior view: A pediatric femoral neck fracture with a comminuted medial column on anteroposterior X-ray of the hip joint; Lateral view: This is a lateral X-ray of a child's femur, showing discontinuity in the thecortical bone of the femoral neck, indicating a femoral neck fracture, with the distal fracture fragment displaced posteriorly relative to the proximalone."

[0091] 6. Model Testing and Evaluation

[0092] The model's performance was evaluated on the test set. For each classification task (anteroposterior or lateral), the accuracy, recall, positive predictive value, negative predictive value, and F1 score of the unimodal VGG16 model and the model of this invention were calculated. Simultaneously, receiver operating characteristic (ROC) curve analysis was used to calculate the area under the curve (AUC) and corresponding 95% confidence intervals for each classification task for different models. Higher values ​​for accuracy, recall, positive predictive value, negative predictive value, F1 score, and AUC indicate better model performance. Specific experimental results are shown in Tables 1-4 below.

[0093] Table 1. Comparison of test results of different models in the frontal image classification task.

[0094] Model accuracy Recall rate Positive predictive value Negative predictive value F1 score VGG16 model 90.14% 85.78% 83.99% 83.99% 84.83% The model constructed in this invention 96.48% 92.42% 96.29% 96.29% 94.20%

[0095] Table 2 Comparison of test results of different models in lateral view image classification task

[0096] Model accuracy Recall rate Positive predictive value Negative predictive value F1 score VGG16 model 54.41% 50.48% 46.97% 83.26% 47.75% The model constructed in this invention 70.59% 63.17% 68.76% 88.71% 65.17%

[0097] Table 3 Comparison of test results of different models in the frontal image classification task

[0098] Model Area under the curve Corresponding to 95% confidence interval VGG16 model 91.64% (83.86%,97.42%) The model constructed in this invention 95.21% (89.57%,99.36%)

[0099] Table 4 Comparison of test results of different models in the lateral image classification task

[0100] Model Area under the curve Corresponding to 95% confidence interval VGG16 model 74.87% (68.80%,80.69%) The model constructed in this invention 85.08% (79.32%,90.12%)

[0101] As can be seen from the results above (Tables 1-4), the model constructed in this invention outperforms the VGG16 monomodal model in both anteroposterior and lateral image classification tasks. Although the VGG16 model is currently the mainstream monomodal model, its performance is still inferior to the model constructed in this invention. These results suggest that the image-text fusion-based model construction strategy applied in this invention is of significant importance, overcoming the technical bottleneck of the monomodal VGG16 model in assessing femoral neck fracture images in children, demonstrating its significant innovation.

[0102] The results show that the accuracy of this invention in classifying pediatric femoral neck fractures in both anteroposterior and lateral views is significantly higher than that of the single-modal image deep learning model (VGG16 model). This is related to the fact that the image-text dataset constructed in this invention provides richer data features than single-dimensional image features. This invention combines an image feature extraction module and a category semantic prototype library to capture and fuse rich image-text features from both visual and textual perspectives, compensating for the limited data features provided by the single-modal dataset, thereby effectively improving the model's accuracy and demonstrating significant innovation. Based on limited sample features, this invention not only successfully achieves intelligent and accurate classification of pediatric femoral neck fractures based on image analysis but also provides new ideas for constructing intelligent diagnostic and treatment models for other rare diseases. Its clinical significance is significant and its application prospects are broad.

[0103] The above are merely preferred embodiments of this application and are not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A method for intelligent classification of pediatric femoral neck fractures based on image-text fusion, characterized in that, Includes the following steps: Construct a graphic-text dataset of images of children with different subtypes of femoral neck fractures in anteroposterior and lateral views, combined with corresponding textual feature descriptions; The image feature extraction module is used to extract image features from the image-text dataset; The application category semantic prototype library is used to extract text features from image and text datasets; The application feature alignment and classification module fuses image and text features and outputs classification probabilities. The model is trained using the application training module; The application diagnostic module and report generation module generate structured diagnostic reports.

2. The intelligent classification method for pediatric femoral neck fracture images based on image-text fusion according to claim 1, characterized in that, The constructed image-text dataset of children with different subtypes of femoral neck fractures in anteroposterior and lateral views, combined with corresponding text feature descriptions, includes: A dataset of images of children with different subtypes of femoral neck fractures in the anteroposterior view was constructed, combined with corresponding text feature descriptions. Femoral neck fractures were divided into two subtypes based on whether the medial column of the neck was comminuted. The standardized text descriptions corresponding to the different subtypes were "anteroposterior X-ray shows femoral neck fracture, no comminuted bone fragments in the medial column of the fracture end" and "anteroposterior X-ray shows femoral neck fracture, comminuted bone fragments are visible in the medial column of the fracture end". A pictorial dataset of images of children with different subtypes of femoral neck fractures in lateral views was constructed, combined with corresponding textual feature descriptions. Based on the displacement direction of the distal fracture fragment relative to the proximal fracture fragment in the sagittal view and the stability of the posterior column of the neck, femoral neck fractures were divided into four subtypes. The standardized textual descriptions corresponding to different subtypes are as follows: "Lateral X-ray shows femoral neck fracture, with no obvious displacement of the distal fracture fragment relative to the proximal fracture fragment", "Lateral X-ray shows femoral neck fracture, with the distal fracture fragment displaced anteriorly relative to the proximal fracture fragment", "Lateral X-ray shows femoral neck fracture, with the distal fracture fragment displaced posteriorly relative to the proximal fracture fragment", and "Lateral X-ray shows femoral neck fracture, with comminuted bone fragments visible in the posterior column of the fracture end".

3. The intelligent classification method for pediatric femoral neck fracture images based on image-text fusion according to claim 1, characterized in that, The image feature extraction module extracts image features from the image-text dataset. The specific steps are as follows: A pre-trained convolutional neural network is used as the backbone network, and its last fully connected layer is removed to output image feature vectors; among them, a pre-trained VGG16 model is used to extract image features from the dataset.

4. The intelligent classification method for pediatric femoral neck fracture images based on image-text fusion according to claim 1, characterized in that, The application category semantic prototype library extracts text features from the image and text dataset, and the specific steps are as follows: Based on the category semantic prototype library, a pre-trained CLIP model text encoder is used to encode the standardized descriptive text corresponding to each femoral neck fracture subtype in the anteroposterior or lateral view, thereby obtaining the category semantic prototype vector.

5. The intelligent classification method for pediatric femoral neck fracture images based on image-text fusion according to claim 1, characterized in that, The application feature alignment and classification module fuses image and text features and outputs classification probabilities. The specific steps are as follows: Image feature vectors and category semantic prototype vectors are mapped to the same shared semantic space through learnable linear projection layers and then normalized. Calculate the cosine similarity between the normalized image feature vector and the semantic prototype vector of each category, and multiply it by a learnable scaling factor to obtain the classification logical value for each category. By maximizing the cosine similarity between image features and their corresponding text prototype vectors, semantic alignment and fusion of cross-modal features are achieved. The classification probability is output using the softmax function.

6. The intelligent classification method for pediatric femoral neck fracture images based on image-text fusion according to claim 5, characterized in that, The shared semantic space is set to 512 dimensions; the learnable scaling factor is initialized to log(1 / 0.07) and optimized along with the network parameters during training.

7. The intelligent classification method for pediatric femoral neck fracture images based on image-text fusion according to claim 1, characterized in that, The application training module trains the model, and the specific steps are as follows: The parameters of the image feature extraction module, the learnable linear projection layer, and the learnable scaling factor are optimized using the backpropagation algorithm based on the weighted cross-entropy loss function.

8. The intelligent classification method for pediatric femoral neck fracture images based on image-text fusion according to claim 1, characterized in that, The application diagnostic module and report generation module generate a structured diagnostic report, and the specific steps are as follows: The diagnostic module outputs two-class classification results for anteroposterior view images of femoral neck fractures in children and four-class classification results for lateral view images of femoral neck fractures in children. The application report generation module integrates the corresponding fracture subtypes from the above anteroposterior and lateral images and generates a structured diagnostic report.

9. A smart imaging classification device for pediatric femoral neck fractures based on image-text fusion, characterized in that, It includes a memory and a processor; the memory is used to store instructions executable by the processor, and the processor is used to call the program instructions stored in the memory to implement the intelligent classification steps of pediatric femoral neck fracture images based on image-text fusion as described in any one of claims 1 to 8.