Femoral head image segmentation method based on OsteoSegNet model
By using the OsteoSegNet model in the femoral head CT image segmentation, a variety of feature extraction modules and feature fusion networks are fused, and the problems of low diagnostic accuracy and efficiency in the existing technology are solved, and efficient and accurate femoral head image segmentation is achieved.
Patent Information
- Application Number
- CN202510067340.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2025-05-13
AI Technical Summary
The prior art is difficult to improve diagnostic accuracy and efficiency in CT imaging segmentation of early ischemic necrosis of the femoral head, especially under complex backgrounds and different equipment conditions.
A femoral head image segmentation method based on OsteoSegNet model is proposed. The feature extraction and segmentation efficiency of the model is optimized through the backbone network fusion of C3K2_REPVGG, SPPF-UniRL and C2PSB modules, combined with multi-scale feature fusion and weighted bidirectional feature pyramid network.
It achieves a significant improvement in segmentation efficiency while ensuring high segmentation accuracy, and can maintain excellent segmentation performance under complex backgrounds and low contrast, providing fast and accurate femoral head segmentation results for clinical practice.
Smart Images

Figure CN119991700A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical image processing, and in particular to a femoral head image segmentation method based on an OsteoSegNet model. Background Art
[0002] Avascular necrosis of the femoral head (ONFH) is a common and severe disabling disease in the field of orthopedics, mainly caused by impaired blood supply to the femoral head, leading to the death of bone cells and bone marrow components. CT scanning has become a powerful tool for evaluating ONFH due to its high spatial resolution and thin-layer cross-sectional imaging technology, especially in the early diagnosis, staging and treatment selection of femoral head necrosis. CT can provide important information about bone sclerosis, subchondral fractures and the scope of lesions, helping doctors to accurately evaluate the progression of ONFH. Traditional femoral head segmentation research mostly relies on time-consuming and labor-intensive methods such as manual annotation and threshold segmentation. This type of research is not only inefficient, but also easily interfered by subjective factors, resulting in poor stability of segmentation results.
[0003] With the rapid development of deep learning technology, automated diagnostic methods based on computer vision and image analysis have gradually become a research hotspot. In recent years, more and more studies have begun to explore the application of deep learning technology in the imaging diagnosis of ONFH, including classification and segmentation of CT images through convolutional neural networks (CNN), accurate extraction of segmented areas using generative adversarial networks (GAN), and feature dimension reduction and unsupervised classification using autoencoders. Although these studies have made some progress, there are still few studies on the application of deep learning technology to the early diagnosis of ONFH, and there are challenges such as how to improve diagnostic accuracy, handle complex backgrounds, and adapt to different equipment conditions.
[0004] Therefore, how to optimize the speed and effect of early diagnosis of ONFH while ensuring high accuracy, especially improving the segmentation accuracy of CT images through deep learning algorithms, remains a research problem that needs to be solved urgently.
[0005] Based on this demand, the present invention proposes a lightweight femoral head instance segmentation network named OsteoSegNet model, which aims to achieve efficient real-time segmentation of multi-scale femoral head areas. Summary of the invention
[0006] In view of the shortcomings of the prior art, the technical problem to be solved by the present invention is to provide a femoral head image segmentation method based on the OsteoSegNet model, aiming to improve the image segmentation accuracy while optimizing the speed and effect of early diagnosis of ONFH, especially to improve the segmentation accuracy of CT images through a deep learning algorithm.
[0007] To solve the above technical problems, in a first aspect, the technical solution adopted by the present invention is a femoral head image segmentation method based on the OsteoSegNet model, comprising the following steps:
[0008] S1: Obtain femoral head CT image data, convert the file into the corresponding jpg image format, traverse each set of CT data, build a data set, and synchronously use ISAT to mark the image;
[0009] S2: Build the OsteoSegNet model, extract features from the data set obtained in step S1, and perform model training;
[0010] S3: Extract femoral head image features using the model obtained in step S2, perform feature fusion, and output the detection results.
[0011] Optionally, the OsteoSegNet model includes a backbone network, a neck network and a head network, wherein the backbone network includes a C3K2_REPVGG module, and the C3K2_REPVGG module is obtained by fusing the RCS-OSA module with the C3K2 module.
[0012] Optionally, the backbone network includes an SPPF-UniRL module, and the SPPF-UniRL module is an SPPF module optimized by an improved spatial pyramid pooling module SPPF and a UniRepLKNet module.
[0013] Optionally, the backbone network includes a C2PSB module, which is a C2PSA module optimized by a dual self-attention mechanism BiFormer module.
[0014] Optionally, the feature maps output by the backbone network are input into a multi-scale feature fusion module ScalSeq, and convolution operations are applied to the three feature maps respectively to adjust their sizes. The adjusted feature maps are converted into three-dimensional tensors and spliced in the third dimension. The spliced three-dimensional tensors are fused through three-dimensional convolution, and the features are refined through maximum pooling operations.
[0015] Optionally, the neck network uses a multi-level feature fusion module SDI and a weighted bidirectional feature pyramid network to optimize the traditional ASF-YOLO model.
[0016] Optionally, the OsteoSegNet model is implemented by PyTorch and executed on a device equipped with NVIDIA GeForce RTX4050 Laptop GPU, 6140 MiB, Intel (R) Core (TM) i7-13700HX 2.10 GHz, Python-3.9.15 torch-1.13.1.
[0017] Optionally, the experimental parameter settings of the OsteoSegNet model are: epochs=200, batch=10, imgsz=640.
[0018] In a second aspect, the present invention further provides an electronic device, comprising:
[0019] one or more processors;
[0020] Memory; and one or more programs stored in the memory, wherein the one or more programs include instructions for executing any of the above-mentioned femoral head image segmentation methods based on the OsteoSegNet model.
[0021] In a third aspect, the present invention further provides a computer-readable storage medium, comprising one or more programs for execution by one or more processors of an electronic device, wherein the one or more programs include instructions for executing any of the above-mentioned femoral head image segmentation methods based on the OsteoSegNet model.
[0022] Compared with the prior art, the technical solution of the present invention has the following beneficial effects:
[0023] 1. The OsteoSegNet model combines the innovative design of contextual lightweight structure and feature balanced network, which can achieve the technical effects of processing image segmentation of lesions of different scales, low consumption of computing resources and high image segmentation speed.
[0024] 2. The OsteoSegNet model can significantly improve segmentation efficiency while ensuring high segmentation accuracy. The model can accurately capture the details of the femoral head area, especially in complex backgrounds or low-contrast conditions, and still maintain excellent segmentation performance, thereby providing faster and more accurate femoral head segmentation results for clinical use. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 A flow chart of a femoral head image segmentation method based on the OsteoSegNet model provided in an embodiment of the present invention;
[0026] Figure 2 A network structure diagram of an OsteoSegNet model provided in an embodiment of the present invention;
[0027] Figure 3 A structural diagram of an RCS module provided by an embodiment of the present invention;
[0028] Figure 4 A structural diagram of an RCS-OSA module provided in an embodiment of the present invention;
[0029] Figure 5 A C3K2_REPVGG module structure diagram provided for an embodiment of the present invention;
[0030] Figure 6 A structure diagram of a SPPF_UniRL module provided in an embodiment of the present invention;
[0031] Figure 7 A C2PSB module structure diagram provided in an embodiment of the present invention;
[0032] Figure 8 A BiFormer module structure diagram provided in an embodiment of the present invention;
[0033] Fig. 9 A structural diagram of a ScalSeq multi-scale feature fusion module provided in an embodiment of the present invention;
[0034] Fig.10 An SDI hierarchical feature graph structure provided by an embodiment of the present invention;
[0035] Fig.11 A schematic diagram of the overall structure of an SDI layer provided by an embodiment of the present invention;
[0036] Fig.12 A C2f-GMD module structure diagram provided in an embodiment of the present invention;
[0037] Fig.13 A data set establishment flow chart provided in an embodiment of the present invention;
[0038] Fig.14 A multi-model segmentation result evaluation provided by an embodiment of the present invention;
[0039] Fig.15 A Grad-CAM heat map provided by an embodiment of the present invention;
[0040] Fig.16 A comparison diagram of OsteoSegNet model training results provided by an embodiment of the present invention;
[0041] Fig.17 This is a graph of image segmentation prediction results of an OsteoSegNet model provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0042] Obviously, many modifications and changes made by those skilled in the art based on the purpose of the present invention belong to the protection scope of the present invention.
[0043] It will be understood by those skilled in the art that, unless expressly stated, the singular forms "a", "an", "said" and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of the present invention refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when an element or component is said to be "connected" to another element or component, it may be directly connected to the other element or component, or there may be an intermediate element or component. The term "and / or" used herein includes any unit and all combinations of one or more associated listed items.
[0044] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0045] The technical solution of the present invention is described in detail below with reference to the accompanying drawings. A femoral head image segmentation method based on the OsteoSegNet model, such as Figure 1 As shown, the following steps are included:
[0046] S1. Obtain femoral head CT image data, convert the file into the corresponding jpg image format, traverse each set of CT data, build a data set, and synchronously use ISAT to mark the image.
[0047] S2. Build the OsteoSegNet model. The network structure of the OsteoSegNet model is shown in the figure below. Figure 2 As shown, including:
[0048] Backbone network (Backbone), neck network (Neck) and head network (Head).
[0049] The backbone network of the OsteoSegNet model: The RCS-OSA module is integrated into the C3K2 module to obtain the C3K2_REPVGG module to improve the processing efficiency and segmentation accuracy of the model, and the C2PSA is improved with the BiFormer module of the bidirectional attention mechanism to obtain the C2PSB module, which enhances the model's perception of the femoral head area. The UniRepLKNet module and the improved spatial pyramid pooling module SPPF are efficiently combined to form the SPPF-UniRL module, which effectively improves the perception ability of the model and avoids the local information loss that may be caused by dilated convolution.
[0050] The segmentation of femoral head necrosis images usually involves extracting clear regional information from complex medical images. The RCS-OSA module can focus on the key features of the femoral head area and its lesions by reducing the number of channels while enhancing spatial attention, so that the model can quickly find and accurately segment the femoral head area in high-resolution medical images.
[0051] RCS (Reparameterized Convolution Based on Channel Shuffle) aims to learn rich feature information through a multi-branch structure during the training phase, and reduce memory consumption and achieve fast reasoning by simplifying it to a single-branch structure during the inference phase. In addition, RCS uses channel segmentation and channel shuffle operations to reduce computational complexity while maintaining information exchange between channels, so that the computational complexity can be reduced by half compared to ordinary 3×3 convolutions during the inference phase. Through structural reparameterization, RCS is able to learn deep representations from input features during the training phase and achieve fast reasoning during the inference phase while reducing memory consumption. In the RCS module, the structure uses multiple branches during the training phase, including 1x1 and 3x3 convolutions, and a direct connection (Identity) to learn rich feature representations. During the inference phase, the structure is reparameterized into a single 3x3 convolution to reduce computational complexity and memory consumption, while maintaining the feature expression capabilities learned during the training phase, which can improve computational efficiency without sacrificing performance. The structure of RCS is shown in the figure below. Figure 3 As shown in the figure, it is divided into a training phase and an inference phase. In the training phase, the input is split by channels, one part of the input passes through the RepVGG block, and the other part remains unchanged. The output of the RepVGG block is then processed by 1x1 convolution and 3x3 convolution, and the channel is shuffled and connected with the other part of the input. In the inference phase, the original multi-branch structure is simplified to a single 3x3RepConv block. This design allows complex features to be learned during training and reduces computational complexity during inference. The rectangles with black borders represent specific module operations, the gradient rectangles represent specific features of the tensor, and the width of the rectangle represents the number of channels of the tensor.
[0052] OSA (One-Shot Aggregation) is a key module designed to improve the efficiency of the network when dealing with dense connections. The OSA module overcomes the inefficiency of dense connections in DenseNet by representing diverse features with multiple receptive fields and aggregating all features only once in the final feature map. The use of the OSA module has two main purposes:
[0053] 1. Improve the diversity of feature representation: OSA increases the network's sensitivity to different scales by aggregating features with different receptive fields, which helps improve the model's ability to segment objects of different sizes.
[0054] 2. Improved efficiency: By performing feature aggregation only once in the last part of the network, OSA reduces repeated feature computation and storage requirements, thereby improving the computational and energy efficiency of the network.
[0055] The OSA module is further combined with the RCS module to form the RCS-OSA module, such as Figure 4 As shown. This combination not only maintains low-cost memory consumption, but also achieves effective extraction of semantic information, which is particularly important for building lightweight and large-scale object segmenters. In the RCS-OSA module, the input is divided into two parts, one part passes directly, and the other part is processed by the stacked RCS modules. The processed features and the features passed directly are merged after channel shuffle (ChannelShuffle). This structure is designed to enhance the feature extraction and utilization efficiency of the model. A key component in the RCS-YOLO architecture aims to improve the model's ability to process features through one-time aggregation while maintaining computational efficiency.
[0056] Combine the RCS-OSA module and the C3K2 module to obtain the C3K2_REPVGG module, such as Figure 5 As shown, the specific implementation includes the following steps: 1. Create RCS-OSA module file to implement RCS-OSA module; 2. Create C3K2 module file to implement C3K2 module; 3. Modify the backbone network of OsteoSegNet model to integrate C3K2 module and RCS-OSA module; 4. Model training and verification.
[0057] In the femoral head instance segmentation task, the edge of the femoral head may be blurred, especially in low-quality or low-contrast images. The SPPF-UniRL module is obtained by combining the improved spatial pyramid pooling module SPPF and the UniRepLKNet module, as shown in Figure 6As shown in the figure, the image feature map (B, C1, H, W) enters the module, and then the feature map is subjected to three maximum pooling operations. Each pooling operation will obtain information of different scales. Finally, all the pooled feature maps are combined by splicing, and then the spliced feature maps are further processed by the UniRepLKNet module. This module will combine the Dilated Reparam Block and GRN to further enhance the capture of details and improve the local feature extraction ability of the model to achieve enhanced local details, so that the model can better capture these details. The size of the femoral head varies greatly, especially through multi-scale pooling operations in different patients and different scanning angles. The model can adapt to these changes and perform segmentation more accurately.
[0058] C2PSB: The instance segmentation task requires accurate segmentation of different objects at the pixel level, and for complex structures (such as the femoral head), it is particularly important to be able to effectively extract their multi-scale features. Traditional convolutional neural networks (CNNs) can usually only effectively capture local features and may encounter difficulties when faced with complex object shapes. In this context, we propose to improve the C2PSA (Channel-wise Position Attention) module by combining the BiFormer module (dual attention mechanism). BiFormer introduces a dual self-attention mechanism to enhance the network's ability to model multi-scale features. Combining the BiFormer module with the C2PSA module can effectively improve the accuracy of femoral head instance segmentation, especially when dealing with subtle boundaries and complex shapes.
[0059] By introducing PSA (Position-Sensitive Attention), it aims to enhance feature extraction capabilities through multi-head attention mechanism and feedforward neural network. It can selectively add residual structure (shortcut) to optimize gradient propagation and network training effect. At the same time, using FFN can map input features to a higher-dimensional space, capture the complex nonlinear relationship of input features, and allow the model to learn richer feature representations. Specifically, we integrate the BiFormer module containing BRA into the C2PSA of the OsteoSegNet architecture to obtain the C2PSB module, as shown in Figure 7 As shown in the figure, the input feature map is processed by the first convolutional layer (cv1) and divided into two parts: one part is used to maintain the original channel information, and the other part is enhanced by the self-attention mechanism. After being processed by the BiFormer module, the enhanced features are restored to the original number of channels through the second convolutional layer (cv2). Finally, the network merges the features of the two parts and outputs the feature map enhanced by BiFormer.
[0060] like Figure 8As shown in the figure, BiFormer introduces a dual attention mechanism based on the traditional self-attention mechanism. It adopts a two-layer routing strategy to achieve dynamic, query-aware sparsity. By collecting the key-value pairs in the first k relevant windows, the sparsity is used to skip the calculation process of the least relevant area, and only dense matrix multiplication suitable for GPU is performed. In the traditional attention mechanism, all key-value pairs are calculated globally, resulting in high computational complexity. In BiFormer, through the two-layer routing attention mechanism, only the first k windows related to the query are paid attention to, and only dense matrix multiplication suitable for GPU is performed. This approach takes advantage of sparsity and avoids redundant calculations in the least relevant areas, thereby improving computational efficiency. Only the key-value pairs related to the query participate in the dense matrix multiplication operation, reducing the amount of calculation and memory usage.
[0061] Furthermore, the feature map output by the backbone network is input into the multi-scale feature fusion module ScalSeq, as Fig. 9 As shown in the figure, convolution operations are applied to the three feature maps respectively to extract higher-level features. The dimensions are adjusted to ensure effective fusion of the feature maps. Then, the adjusted feature maps are converted into three-dimensional tensors and concatenated in the third dimension. The concatenated three-dimensional tensors are fused through three-dimensional convolution to capture multi-scale features. Finally, the features are refined through the maximum pooling operation to enhance the expressiveness of the features.
[0062] The neck network of the OsteoSegNet model: combines multi-level feature fusion and weighted bidirectional feature pyramid to improve the accuracy and speed of image processing, and optimizes the convolution operation of C2f through GhostModule to reduce redundant calculations, so that the model can better adapt to different imaging conditions (such as different resolutions, device types, etc.).
[0063] The femoral head instance segmentation task not only involves the recognition and positioning of the target, but also requires the model to accurately segment the specific area of each femoral head, especially in complex backgrounds or when multiple femoral heads overlap. To this end, based on the traditional ASF-YOLO model, a multi-level feature fusion module (SDI) and a weighted bidirectional feature pyramid network (Bi-FPN) are used, and dynamic convolution is introduced to replace the traditional convolution layer, thereby reducing the computational complexity and memory consumption, while maintaining or even improving network performance. The innovative Neck structure greatly enhances the model's ability to handle femoral head instance segmentation tasks.
[0064] SDI combined with ASF-YOLO: Accurate boundary recognition is one of the core challenges in the bone instance segmentation task. The boundary of the femoral head may be relatively fuzzy or complex, especially in the presence of cracks, lesions, etc. The SDI module can help the model better capture and segment the complex boundaries of the femoral head by fusing low-level detail features (such as edges and textures) with high-level semantic features (such as overall shape and position). Fig.10 As shown in the figure. The size and shape of the femoral head may vary greatly in different images, especially in CT images. The SDI module can effectively handle this scale inconsistency, allowing the model to accurately segment the femoral head targets at different scales, especially small targets or targets with complex shapes (such as femoral head cracks or lesions) can be better identified. The femoral head segmentation task requires both capturing fine local structures (such as small lesions, cracks, etc.) and understanding global semantic information (such as the approximate position and overall shape of the femoral head). Fig.11 This is a schematic diagram of the overall structure of SDI. SDI achieves an organic combination of low-level details and high-level semantics through multi-level feature fusion, making segmentation more accurate.
[0065] ASF-YOLO uses effective feature fusion and multi-scale processing, which is particularly important for femoral head instance segmentation, especially when segmenting the details of the femoral head, to ensure accurate segmentation without missing or mis-segmenting. It has a significant advantage in processing speed. For femoral head instance segmentation, especially on large-scale medical image data sets, it can significantly improve segmentation efficiency and shorten segmentation time, making it suitable for real-time medical image analysis.
[0066] In the neck part, C2f introduces dynamic convolution to replace the traditional convolution layer. Through more efficient feature fusion and more refined feature extraction, the model can more accurately segment the details and boundaries of the femoral head, especially in complex backgrounds or multi-target scenes. For tiny lesions, cracks and other details, the dynamic convolution mechanism of GhostModule can provide better small target segmentation capabilities and improve the processing accuracy of details. Through the dynamic convolution mechanism, feature fusion and computational efficiency improvement, it can not only improve the computational efficiency while maintaining high accuracy, but also enhance the processing capabilities of complex backgrounds and small targets. This solves the problem of poor segmentation effect of femoral head instance segmentation under low resolution and complex backgrounds. The original C2f structure continues the advantages of the ELAN structure multi-gradient classification, and adds the residual branch of BottleNeck, so that the model can learn richer feature representations. The present invention is further improved on the basis of the Ghost module of C2f, replacing BottleNeck with Ghost BottleNeck, and then obtaining the C2f-GMD module, such as Fig.12 shown.
[0067] S3. Use the model obtained in step S2 to extract femoral head image features, perform feature fusion, and output the detection results.
[0068] Experimental environment: All experiments are implemented using PyTorch and executed on a system equipped with NVIDIA GeForce RTX4050Laptop GPU, 6140MiB, Intel(R) Core(TM) i7-13700HX 2.10GHz, Python-3.9.15torch-1.13.1.
[0069] Experimental data set: see attached Fig.13 ,The CT data of the participants are exported by the hospital imaging workstation, and the files need to be converted into the corresponding jpg image format. Then, each set of CT data is traversed to construct a data set, and ISAT is used to mark the images synchronously.
[0070] Experimental parameter settings of the OsteoSegNet model: epochs=200, batch=10, imgsz=640.
[0071] Experimental results: Tests were conducted based on different models using the same dataset and experimental environment. Fig.14The segmentation results of different models (%) show that the OsteoSegNet model has the highest scores under different evaluation indicators, including accuracy (P), recall rate (R), mean average precision mAP (50), and mAP (95). The bounding box accuracy (Box(P)) can reach 97.1, which can accurately locate the target area in the femoral head CT image. Compared with other models, its positioning is more accurate. The mask accuracy (Mask(P)) is 97.1, which can accurately outline the contour of the femoral head and provide accurate mask information for subsequent medical analysis. The recall rate (R) reaches 97.3, which means that the model can detect as many femoral head targets in the image as possible and reduce the occurrence of missed detection. In terms of mask prediction, the recall rate (R) is also 97.3, which can completely recall the mask information of the femoral head, making the segmentation result more complete. In terms of mean average precision (mAP) indicators, OsteoSegNet performs well under different intersection-over-union (IoU) thresholds. The mAP(50) is as high as 98.8, and the mAP(95) is 95.4, which is much higher than models such as YOLOV5-seg and MaskR-CNN. In terms of the average precision of mask prediction, the mAP(50) is 98.0, and the mAP(95) is 73.9, which can provide accurate segmentation results under different mask IoU thresholds. Therefore, OsteoSegNet is the best model, which can accurately locate and segment, provide high-quality basis for medical diagnosis, and has relatively stable performance. It will not fluctuate greatly due to slight changes in data, providing reliable technical support for medical diagnosis.
[0072] The Grad-CAM heat map shows that the model’s decisions are mainly activated by the regions of interest, such as Fig.15 As shown, the heat map shows that the decision of the convolutional neural network is mainly activated by the located femoral head area, which proves that the developed OsteoSegNet model is reliable and can use important image features of interest to make decisions.
[0073] Fig.16This is the segmentation curve of OsteoSegNet. In the figure, whether it is train / box_loss, train / seg_loss, train / cls_loss or train / dfl_loss, they all show a trend of gradually decreasing with the increase of the number of training iterations. This shows that the model can effectively learn the characteristics of the data during the training process and continuously optimize its own parameters, so that the gap between the predicted results and the true values gradually decreases, reflecting the good learning ability and adaptability of the model. The loss curve gradually stabilizes in the later stage, and the value is relatively low. This shows that after a certain round of training, the model can maintain good performance on unseen verification data without overfitting. That is, the model not only performs well on the training set, but also has good generalization ability and can adapt to new data. The metrics / precision(B) and metrics / recall(B) (accuracy and recall of bounding box prediction) curves as well as metrics / precision(M) and metrics / recall(M) (accuracy and recall of mask prediction) curves are all rising steadily, indicating that the model can find the target more and more accurately and recall the target information completely in the bounding box localization and mask segmentation tasks, which is crucial for the accuracy and completeness of the model in practical applications. The rise of the mean average precision curves such as metrics / mAP50(B), metrics / mAP50-95(B), metrics / mAP50(M) and metrics / mAP50-95(M) indicates that the comprehensive performance of the model under different intersection-over-union (IoU) thresholds is constantly improving. In particular, the rise of mAP50 and mAP50-95 indicates that the model can perform well under both looser and stricter IoU threshold requirements, reflecting the stability and reliability of the model in detection and segmentation tasks.
[0074] Fig.17 The OsteoSegNet segmentation prediction result is shown. The segmentation prediction results of multiple femoral head CT images are shown. The blue highlighted area in the image represents the part of the femoral head predicted and segmented by the model, and different confidence levels and other information are indicated above some highlighted areas. We can see that the model can accurately locate and segment the femoral head area. In different CT images, the shape and position of the femoral head are different, but the model can capture the contour of the femoral head well, and the segmented area is basically consistent with the actual position and shape of the femoral head. For example, in some images, the boundary of the femoral head is clear, and the model can accurately segment along the edge of the femoral head; in images with high contrast between the femoral head and surrounding tissues, the segmentation effect is more obvious and accurate.
[0075] Through the above steps, the OsteoSegNet model combines the innovative design of contextual lightweight structure and feature balanced network to solve the problems of traditional methods in the segmentation process, such as difficulty in simultaneously processing lesions of different scales, excessive consumption of computing resources and slow segmentation speed. By introducing a variety of improved modules, such as bidirectional attention mechanism (BiFormer optimized C2PSA), multi-scale spatial pyramid pooling (SPPF), UniRepLKNet module, etc., the OsteoSegNet model can significantly improve the segmentation efficiency while ensuring high accuracy, and can effectively capture the details of the femoral head area, especially in complex backgrounds or low contrast conditions, and can still maintain excellent segmentation effects, thereby providing more accurate and fast femoral head area segmentation results for clinical use.
[0076] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0077] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0078] Finally, it should be noted that, in this article, relationships such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprises" or any other variations are intended to cover non-exclusive inclusion, so that a process, method, article or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device.
Claims
1. A femoral head image segmentation method based on the OsteoSegNet model, characterized in that: include: S1: Obtain femoral head CT image data, convert the file into the corresponding jpg image format, traverse each set of CT data, build a data set, and synchronously use ISAT to mark the image; S2: Build the OsteoSegNet model, extract features from the data set obtained in step S1, and perform model training; S3: Extract femoral head image features using the model obtained in step S2, perform feature fusion, and output the detection results.
2. According to the femoral head image segmentation method based on the OsteoSegNet model in claim 1, the OsteoSegNet model comprises a backbone network, a neck network and a head network, characterized in that: The backbone network includes a C3K2_REPVGG module, and the C3K2_REPVGG module is obtained by integrating an RCS-OSA module with a C3K2 module.
3. The femoral head image segmentation method based on the OsteoSegNet model according to claim 2, characterized in that: The backbone network includes an SPPF-UniRL module, and the SPPF-UniRL module is an SPPF module optimized by an improved spatial pyramid pooling module SPPF and a UniRepLKNet module.
4. The femoral head image segmentation method based on the OsteoSegNet model according to claim 2, characterized in that: The backbone network includes a C2PSB module, which is a C2PSA module optimized by a dual self-attention mechanism BiFormer module.
5. The femoral head image segmentation method based on the OsteoSegNet model according to claim 2, characterized in that: The feature maps output by the backbone network are input into the multi-scale feature fusion module ScalSeq, and convolution operations are applied to the three feature maps respectively to adjust the sizes. The adjusted feature maps are converted into three-dimensional tensors and spliced in the third dimension. The spliced three-dimensional tensors are fused through three-dimensional convolution, and the features are refined through maximum pooling operations.
6. The femoral head image segmentation method based on the OsteoSegNet model according to claim 2, characterized in that: The neck network uses a multi-level feature fusion module SDI and a weighted bidirectional feature pyramid network to optimize the traditional ASF-YOLO model.
7. The femoral head image segmentation method based on the OsteoSegNet model according to claim 1, characterized in that: The OsteoSegNet model was implemented by PyTorch and executed on a device equipped with NVIDIA GeForce RTX 4050 Laptop GPU, 6140 MiB, Intel (R) Core (TM) i7-13700HX 2.10 GHz, Python-3.9.15 torch-1.13.
1.
8. The femoral head image segmentation method based on the OsteoSegNet model according to claim 1, characterized in that: The experimental parameter settings of the OsteoSegNet model are: epochs=200, batch=10, imgsz=640.
9. An electronic device, characterized in that: include: one or more processors; Memory; and one or more programs stored in a memory, wherein the one or more programs include instructions for executing the femoral head image segmentation method based on the OsteoSegNet model as described in any one of claims 1-8.
10. A computer-readable storage medium, characterized in that: It comprises one or more programs for execution by one or more processors of an electronic device, wherein the one or more programs comprise instructions for executing the femoral head image segmentation method based on the OsteoSegNet model as described in any one of claims 1 to 8.