Fusion feature enhanced fracture image classification method and system
Through the deep learning model combined with multiple feature enhancement strategies, the fusion problem of local details and global structure in fracture image classification is solved, and efficient and accurate fracture image classification is achieved, which is suitable for medical image analysis.
Patent Information
- Application Number
- CN202510633459.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-16
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-05-16
AI Technical Summary
The prior art is difficult to efficiently integrate local details and global structural information in fracture image classification, and computational resources are wasted and classification accuracy is insufficient, especially in complex fracture image processing.
The deep learning model is adopted, combined with multiple feature enhancement strategies, including dense blocks, feature fusion enhancement modules, multi-scale expansion convolution, adaptive significance guidance and low-rank fusion technology, and optimize feature expression and computing efficiency through multi-scale expansion convolution and channel attention mechanisms of local and global features.
It improves the accuracy and robustness of fracture image classification, can accurately identify fracture areas in complex fracture images, reduce calculation complexity, and meet the efficient classification needs of medical scenarios.
Smart Images

Figure CN120259780A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer vision and artificial intelligence, and particularly relates to a fracture image classification method and system with enhanced fusion features. Background Art
[0002] The statements in this part merely provide background technical information related to the present invention and do not necessarily constitute prior art.
[0003] Fracture classification in medical image analysis is a key challenge. Traditional methods rely on artificial feature extraction techniques such as edge detection and morphological transformation, but perform poorly on complex fractures and low-quality images, making it difficult to meet clinical needs. In recent years, convolutional neural networks (CNNs) have made significant progress in medical image analysis and can automatically learn multi-level features from raw images. However, although standard CNNs perform well in local feature extraction, they have limited ability to capture long-range dependencies and global context information, while fracture diagnosis often requires considering both the details of tiny cracks and the changes in the overall bone structure simultaneously.
[0004] To overcome the limitations of CNNs, researchers have proposed various improvement methods. Dense connection networks (DenseNets) enhance feature propagation and reuse through dense connections between layers, improving the model's expressiveness and parameter utilization efficiency. However, even advanced CNN architectures still face challenges in processing the complex global structures and tiny details of fracture images, especially in scenarios that require integrating multi-scale and multi-modal information.
[0005] Salience-guided techniques aim to help the model identify key regions with diagnostic value. Traditional salience detection identifies prominent regions based on low-level visual features, but the diseased regions in medical images often do not differ significantly from the surrounding healthy tissues, resulting in poor performance. Deep learning methods can generate salience maps that are more in line with the needs of medical diagnosis, but often require additional annotated data or a two-stage process, increasing the computational burden and complexity. Existing salience-guided methods mostly adopt a spatial weighting form, ignoring the mutual relationships between feature channels and not fully utilizing multi-dimensional information.
[0006] Attention mechanisms have been introduced into visual models to enhance the attention to important regions. Techniques such as self-attention and channel attention allow the model to dynamically adjust the attention to different positions and channels, improving the ability to identify key features. However, when dealing with complex medical images, a simple attention mechanism often needs to be combined with other techniques to achieve the best results.
[0007] In terms of feature fusion, existing multi-branch models extract different levels of features through different network paths and then fuse them through simple concatenation or weighted summation. Although these methods can integrate different levels of features, they often ignore the high-order interaction relationships between features, resulting in information redundancy and waste of computational resources, which is not conducive to medical deployments with limited resources.
[0008] Feature compression and redundancy reduction techniques such as Tucker decomposition can decompose high-dimensional feature tensors into a compact core tensor and factor matrices, effectively reducing the number of parameters and computational complexity. However, traditional applications are often carried out in a static manner, lacking adaptability to the complexity of input data and unable to dynamically adjust the complexity according to the image content.
[0009] For the second-order feature extraction of fracture images, existing methods mainly rely on the first-order features extracted by global pooling or convolutional operations, ignoring the covariance relationship between features, and these high-order statistical information is crucial for identifying subtle fracture features. Although some studies have begun to explore techniques such as second-order pooling and bilinear pooling, the balance between computational efficiency and feature expression ability still needs to be optimized.
[0010] In summary, the existing technologies still face challenges such as efficiently integrating local and global information, adaptively adjusting feature strategies, enhancing attention to key regions, and reducing computational redundancy while maintaining high classification accuracy. Summary of the Invention
[0011] To solve the above problems, the present invention proposes a fracture image classification method and system with fused feature enhancement. By adopting a deep learning model and introducing multiple feature enhancement strategies, the present invention can take into account both local detailed features and overall structural information while extracting local detailed features, thereby effectively improving the accuracy and robustness of fracture classification and providing reliable technical support for intelligent medical image analysis.
[0012] According to some embodiments, the first solution of the present invention provides a fracture image classification method with fused feature enhancement, adopting the following technical solutions: A fracture image classification method with fused feature enhancement, comprising: Obtain fracture medical images and perform image preprocessing; Based on the preprocessed fracture medical images, use a pre-trained fracture image classification model for classification, specifically: After the preprocessed fracture images are initially convolved, they are sequentially processed through four dense blocks and a feature fusion and enhancement module to obtain the final fused features, and then processed through a transition layer and a global pooling layer to obtain corrected fused features. Based on the corrected fused features, fracture image classification is performed to obtain a classification result; Among them, the processing process of each dense block and the feature fusion and enhancement module is: Extract primary features based on the dense block; Enhance the primary features and then split them into local features and retained features along the channel dimension. After enhancing the local features, fuse them with the retained features to obtain locally enhanced features; Perform multi-scale dilated convolution and adaptive saliency guidance processing on the locally enhanced features, fuse the processing results with the locally enhanced features to obtain globally attention features; Concatenate the locally enhanced features and the globally attention features to obtain a feature tensor, based on the second-order feature optimized tensor decomposition method of the feature tensor, and use the optimized tensor decomposition method to decompose and reconstruct the feature tensor to obtain fused features.
[0013] According to some embodiments, the second solution of the present invention provides a fracture image classification system with fused feature enhancement, and adopts the following technical solution: A fracture image classification system with fused feature enhancement, comprising: An image preprocessing module configured to obtain a fracture medical image and perform image preprocessing; An image classification module configured to classify based on the preprocessed fracture medical image by using a pre-trained fracture image classification model, specifically: After the preprocessed fracture image undergoes initial convolution processing, it sequentially passes through four dense blocks and a feature fusion and enhancement module to obtain final fused features, and then passes through a transition layer and a global pooling layer to obtain corrected fused features, and classify the fracture image based on the corrected fused features to obtain a classification result; Among them, the processing process of each dense block and the feature fusion and enhancement module is: Extract primary features based on the dense block; Enhance the primary features and then split them into local features and reserved features along the channel dimension, enhance the local features and then fuse them with the reserved features to obtain locally enhanced features; Perform multi-scale dilated convolution and adaptive saliency guidance processing on the locally enhanced features, fuse the processing results with the locally enhanced features to obtain globally attention features; Concatenate the locally enhanced features and the globally attention features to obtain a feature tensor, based on the second-order feature optimized tensor decomposition method of the feature tensor, and use the optimized tensor decomposition method to decompose and reconstruct the feature tensor to obtain fused features.
[0014] According to some embodiments, the third solution of the present invention provides a computer-readable storage medium.
[0015] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in a fracture image classification method with fused feature enhancement as described in the first aspect above.
[0016] According to some embodiments, the fourth solution of the present invention provides a computer device.
[0017] A computer device includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a fracture image classification method with enhanced fusion features as described in the first aspect above.
[0018] According to some embodiments, a fifth aspect of the present invention provides a computer program product or a computer program.
[0019] The present invention provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the steps in a fracture image classification method with enhanced fusion features as described in the first aspect above.
[0020] Compared with the prior art, the beneficial effects of the present invention are as follows: In the fracture image classification, the present invention takes into account both local details and global structure information. By embedding the LCAD (Local Context and Dependency Attention) module and the low-rank fusion module in deep backbone networks such as DenseNet, precise localization of the fracture area and efficient fusion of multi-scale features are achieved. The module uses multi-branch local convolution to strengthen fine-grained information such as micro-cracks and edges, and at the same time combines multi-scale dilated convolution and channel attention mechanism to capture long-range dependencies, enabling the model to still have high discriminative power under complex fracture morphologies. Second-order features are extracted through bilinear descriptor coding, enhancing the expression ability for subtle fracture features; based on the saliency guidance mechanism, the model can automatically focus on the key fracture areas; using the low-rank fusion technology of Tucker decomposition and adaptive complexity adjustment, while improving the classification performance, the computational amount is significantly reduced, meeting the dual requirements of efficiency and accuracy in actual medical scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings forming a part of this invention are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention.
[0022] Figure 1 is a flowchart of a fracture image classification method with enhanced fusion features in an embodiment of the present invention; Figure 2 is an algorithm flowchart of a fracture image classification model in an embodiment of the present invention; Figure 3 is a flowchart of the operation of a feature fusion enhancement module in an embodiment of the present invention; Figure 4 It is the workflow diagram of the low-rank fusion and second-order feature enhancement layer in the embodiments of the present invention. Specific embodiments
[0023] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0024] It should be noted that the following detailed description is illustrative and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.
[0025] In the case of no conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0026] Embodiment 1 As Figure 1 shown, this embodiment provides a fracture image classification method for fusing feature enhancement. In this embodiment, the method includes the following steps: Obtain fracture medical images and perform image preprocessing; Based on the preprocessed fracture medical images, use a pre-trained fracture image classification model for classification, specifically: After the preprocessed fracture images are initially convolutionally processed, they sequentially pass through four dense blocks and a feature fusion and enhancement module to obtain the final fused features, and then pass through a transition layer and a global pooling layer to obtain corrected fused features. Based on the corrected fused features, fracture image classification is performed to obtain a classification result; Among them, the processing process of each dense block and the feature fusion and enhancement module is as follows: Extract primary features based on the dense block; Enhance the primary features and then split them into local features and retained features along the channel dimension. After enhancing the local features, fuse them with the retained features to obtain locally enhanced features; Perform multi-scale dilated convolution and adaptive saliency guidance processing on the locally enhanced features, and fuse the processing results with the locally enhanced features to obtain globally focused features; Concatenate the locally enhanced features and the globally focused features to obtain a feature tensor. Based on the second-order feature optimization tensor decomposition method of the feature tensor, use the optimized tensor decomposition method to decompose and reconstruct the feature tensor to obtain fused features.
[0027] The present invention provides a medical image fusion method based on a large kernel attention mechanism. As Figure 1 shown, its process includes the following steps: 1. Input fracture medical images Obtain the publicly available fracture dataset as fracture medical images. This dataset contains different types of fracture images and non-fracture images, and is suitable for fracture detection and classification tasks in medical imaging.
[0028] 2. Image preprocessing The image preprocessing steps mainly include image resizing, data augmentation, and normalization. First, all images are uniformly scaled to dimensions to ensure input consistency. For the training set, data augmentation methods such as random rotation, flipping, and cropping are also used to increase data diversity and improve the generalization ability of the model. All images are normalized, and the mean and standard deviation of the ImageNet dataset are used to normalize pixel values, thereby accelerating model convergence and enhancing training stability. The dataset is divided into a training set, a validation set, and a test set to ensure model evaluation and tuning at different stages. This technology uses a standard data loading method to represent the dataset as a set containing image samples and corresponding class labels, and inputs them in a batch processing manner to improve computational efficiency. The mathematical expression of data input is as follows: The data input in this embodiment can be represented as: (1); Among them, represents the dataset, which consists of the image sample and its corresponding class label . represents the input fracture or non-fracture image, with a size of , , . Among them, and are the height and width of the image, respectively, fixed at 224, that is, , is the number of channels, taking the value of 3 ( image). is the class label of the image, represents a fracture image, 0 represents a non-fracture image. represents the total number of images in the dataset, including the training set, the validation set, and the test set. Through this data input method, medical imaging data can be efficiently organized and loaded for training and inference by deep learning models.
[0029] As Figure 2 shown, this embodiment proposes a fracture image classification model including a backbone network module and a feature fusion and enhancement module.
[0030] 3. Fracture image classification model, processing flow, including: The fracture image classification model includes a backbone network module, and the backbone network module includes an initial convolutional layer, four dense blocks, a transition layer, and a global pooling layer; a feature fusion and enhancement module is connected behind each dense block.
[0031] Specifically, the fracture image classification model uses a deep neural network as the backbone network, selects DenseNet201 as the backbone network, and optimizes it on this basis. DenseNet201 adopts a dense connection mechanism, enabling the features of each layer to be directly transmitted to subsequent layers, enhancing information circulation, and improving the feature expression ability of the model. Among them, the backbone network module mainly consists of an initial convolutional layer, multiple dense blocks, a transition layer, and a global pooling layer. The layers within each dense block adopt a short connection method to improve the feature transmission efficiency and reduce the problem of gradient disappearance.
[0032] Based on the Dense Block structure, a feature fusion module is additionally embedded to enhance the modeling ability for local details and global dependencies. The primary features output by each dense block will be used as inputs and transmitted to the feature fusion and enhancement module to further optimize the feature expression, enabling the model to more accurately complete the fracture image classification task. Through multi-level convolutional and pooling operations, multi-scale and high-level feature representations are extracted from the input image.
[0033] The preprocessed fracture image is input into the fracture image classification model. First, feature extraction is performed through the initial convolutional layer, and the input image for the dense block is , then the feature extraction process of the dense block can be expressed as follows: (2); Among them, is the feature map output by the dense block, that is, the primary feature, represents the DenseNet201 calculation map, are the network parameters.
[0034] The feature fusion and enhancement module includes a local feature enhancement sub-module and a local-global feature perception and fusion sub-module; In the fracture image classification task, it is necessary to pay attention to both local fine-grained features and capture the overall structure of the bone and its long-range dependencies. If traditional convolutional models are only based on a fixed convolutional kernel scale, it is often difficult to balance these two requirements. Therefore, a Local Context and Dependency Attention (LCDA) feature fusion and enhancement mechanism is proposed to integrate local and global information in a multi-stage and multi-modal manner, and reduce redundancy and computational burden during the fusion process by means of Tucker decomposition.
[0035] The local feature enhancement sub-module includes a convolutional layer, a channel splitting layer, a multi-branch local convolutional layer, and a fusion layer.
[0036] After enhancing the primary features, they are split into local features and retained features along the channel dimension. After enhancing the local features, they are fused with the retained features to obtain locally enhanced features. Specifically: After performing local context enhancement on the primary features, primary extended features are obtained; The primary extended features are split into local features and retained features along the channel dimension according to a set ratio; Multi-branch local convolution is used to divide the local features into two parts to obtain local features of different scales, and the local features of different scales are concatenated in the channel dimension to obtain locally enhanced features after feature enhancement; The locally enhanced features after feature enhancement are concatenated and fused with the retained features to obtain locally enhanced features.
[0037] Specifically, after receiving the primary features from the fracture image local context enhancement is first performed. The core idea is to expand the number of channels from to through convolution, so that the features have a larger expression space.
[0038] Specifically, define the convolutional kernel and the bias , and use a non-linear activation function (such as GELU) for feature mapping. Specifically: (3); (4); Among them, Split represents splitting the feature tensor into two parts along the channel dimension according to a specified ratio, occupies of the number of channels for local feature enhancement and is the local feature; occupies , most of the information of the original features is retained for subsequent global information processing, which are the retained features. This unbalanced segmentation ratio ensures that while enhancing local details, sufficient global context information is retained.
[0039] This embodiment adopts a multi-branch local convolution structure to divide the local features further into two parts, and respectively uses and convolution kernels for processing, enabling the model to capture local features at different scales simultaneously: (5); (6); (7); Among them, and are respectively the two branch features of the local features, and are respectively and convolution kernels, represents the concatenation operation of the channel dimension, and are respectively the enhanced features after the and functions process, is Sigmoid the activation function. This multi-branch structure enables the model to capture the edges, textures and morphological features of the fracture area more precisely, improving the detection ability for micro cracks and subtle structural changes.
[0040] Subsequently, is concatenated with along the channel dimension to obtain the local enhanced feature : (8); Through this step, while the network is strengthened in local detail mining, sufficient global information is still retained for subsequent processing.
[0041] The local-global feature perception and fusion sub-module includes a multi-scale dilated convolution layer, an adaptive saliency guidance layer, and a low-rank fusion and second-order feature enhancement layer.
[0042] Perform multi-scale dilated convolution on the local enhanced feature to obtain the global initial feature, then apply adaptive saliency guidance processing to the global initial feature, and fuse the processing result with the local enhanced feature to obtain the globally focused feature. Specifically: Split the local enhanced feature along the channel dimension to obtain two local enhanced sub-features; After performing small-scale conventional convolution and large-scale dilated convolution on two local enhancer features respectively, local enhancer features of different scales are obtained and concatenated to obtain the global initial feature; The global initial feature is processed separately from the channel and space in the global initial feature using adaptive saliency guidance and then concatenated to obtain the global attention feature.
[0043] Specifically, after completing local context enhancement, the perception of the overall structure of the bone and the long-range dependence relationship is strengthened through a multi-scale dilated convolution structure. In this embodiment, the multi-scale dilated convolution layer simultaneously uses small-scale conventional convolution and large-scale dilated convolution, significantly expanding the receptive field range: (9); (10); (11); (12); Among them, Split is the splitting operation of the channel dimension, and are two parts of the local enhancement feature, that is, the local enhancement feature sub-features, is the small-scale convolution kernel, is the large-scale convolution kernel, represents the dilated convolution operation, and the dilation rate is , is the corresponding bias term, is is the corresponding bias term, is the small-scale feature obtained by processing through small-scale convolution; is the large-scale feature obtained by processing through large-scale dilated convolution; is and the result after concatenation, that is, the global initial feature; is Sigmoid the activation function. This multi-scale dilated convolution structure enables the model to effectively capture the position and context information in the overall structure of the bone, provides a more comprehensive basis for judging the type and severity of fractures, and enhances the modeling ability for long-range dependence relationships.
[0044] In this embodiment, by designing a multi-branch local convolution structure (such as the 3×3 and 5×5 convolution branches in the LCAD module) and multi-scale dilated convolutions (small and large branches), the model can simultaneously focus on local features at different scales, and integrate these features through a low-rank fusion module, thereby improving the recognition ability of various fracture morphologies.
[0045] The adaptive saliency guidance layer processes the global initial features separately in the channel and spatial dimensions using adaptive saliency guidance and then concatenates them to obtain global attention features. Specifically: Use global average pooling to capture the global context information of each channel of the global initial features, and learn the dependencies between all channels through non-linear transformation based on the global context information to obtain channel attention weights; Calculate the channel-guided features using the channel attention weights and the global initial features; Downsample the global initial features to obtain downsampled features, and calculate the spatial covariance matrix of the downsampled features; Normalize the diagonal elements of the spatial covariance matrix of the downsampled features to obtain spatial saliency weights, and perform a broadcast multiplication operation based on the spatial saliency weights and the downsampled features to obtain spatially modulated features; Upsample the spatially modulated features and then concatenate them with the local enhanced features and the channel-guided features to obtain global attention features.
[0046] Specifically, this embodiment designs a saliency guidance mechanism based on second-order statistics and adaptive attention, enabling the model to automatically focus on the key diagnostic regions in fracture images. This mechanism acts simultaneously in the channel dimension and the spatial dimension to form a dual saliency guidance, enabling the model to automatically focus on the key diagnostic regions in fracture images. Different from traditional saliency detection methods, this embodiment is realized through the dual dimensions of the importance of feature channels and the key nature of spatial positions.
[0047] Channel saliency guidance: In the channel dimension, this embodiment adopts a channel attention mechanism based on statistical characteristics, which can learn the importance of different channels for fracture recognition and dynamically adjust the channel weights accordingly to highlight the expression of key feature channels. First, capture the global context information of each channel through global average pooling : (13); Among them, is the coordinate index on the feature map, represents the channel index, , is the number of channels, and respectively represent the spatial position indices (in the height and width directions) of the feature maps; Then, the dependencies between channels are learned through non-linear transformation: (14); (15); where, and are the fully connected layer parameters for dimensionality reduction and dimensionality increase, is the dimensionality reduction ratio, is ReLU the activation function, is Sigmoid the activation function, is the channel descriptor vector obtained through global average pooling. The finally obtained channel attention weight is used to modulate the importance of each channel to obtain the channel-guided feature : (16); This channel saliency mechanism can adaptively emphasize the channels containing key fracture features while suppressing the influence of irrelevant or redundant channels.
[0048] Spatial saliency guidance: In the spatial dimension, spatial saliency guidance based on bilinear descriptor coding (BDC) is introduced. The BDC module first downsamples the feature to a fixed size to reduce the computational burden: (17); where, is a downsampling operation including adaptive pooling and convolution, which reduces the feature map size to a predefined size (such as ), is the downsampled feature. Then, the spatial covariance matrix of the downsampled feature is calculated to capture the interdependencies between spatial positions: (18); (19); (20); (21); where, and are the sizes of the downsampled feature maps, is the reshaped tensor that reshapes into the shape [B, C, -1], is the number of channels, is the batch size, and -1 indicates automatic calculation of the remaining dimension (i.e., ); is a tensor shape reconstruction function that reorganizes the three-dimensional feature map into a two-dimensional form for easy covariance calculation. is the mean of the downsampled features at all spatial positions; is the centered feature, that is, the reshaped tensor of the original feature minus the mean of the downsampled features. is the covariance matrix of the downsampled features, which reflects the correlation between different channels. By extracting the diagonal elements of the covariance matrix of the downsampled features, the spatial saliency distribution of each channel is obtained: (22); (23); Among them, is the function to extract the diagonal elements of the covariance matrix, representing the variance within each channel, which reflects the degree of spatial variation of the channel; is through softmax The normalized channel weight, that is, the spatial saliency weight, is used to highlight the channels with high spatial variability; is the dimension representing the covariance matrix as , is the number of channels.
[0049] These spatial saliency weights are applied to the BDC features to emphasize the regions with high spatial variability, which usually correspond to the fracture locations: (24); Among them, represents the appropriate broadcast multiplication operation, is through The BDC features modulated by the weights, that is, the spatially modulated features. In the spatial dimension, the BDC module can capture the saliency information of spatial positions by calculating the feature covariance and extracting the diagonal weights, enabling the model to pay more attention to the feature expression of the fracture region. This dual saliency guidance mechanism enables the model to automatically focus on the key diagnostic regions when processing complex fracture images, improving the classification accuracy and interpretability. Especially for those cases with unclear fracture features or located in unconventional positions, it can more accurately locate and identify the fracture region, reducing the risk of missed diagnosis and misdiagnosis.
[0050] Subsequently, the features with enhanced spatial saliency are restored to the original size through upsampling and fused with the locally enhanced features: (25); (26); Among them, is the upsampling operation, is a balance parameter that controls the contribution degree of the BDC feature, is the upsampled BDC feature, which is used to fuse with the local enhanced feature to form the global attention feature . By calculating the covariance matrix of the downsampled features and extracting the diagonal elements as channel weights, the model can capture the correlations between channels, which often correspond to the unique patterns of different types of fractures. This second-order feature representation significantly enhances the model's ability to identify subtle fracture features and can distinguish fracture types that are not obvious in spatial distribution but have unique patterns in feature correlations. Through downsampling and covariance calculation, the second-order statistical information of the features is further extracted, forming a multi-level second-order representation of the fracture features, greatly improving the model's ability to represent complex fracture patterns.
[0051] The low-rank fusion and second-order feature enhancement layer works based on a second-order feature optimization tensor decomposition method for the feature tensor. The feature tensor is decomposed and reconstructed using the optimized tensor decomposition method to obtain the fusion feature. Specifically: Calculate the covariance matrix based on the second-order features of the feature tensor to obtain the second-order covariance matrix; Perform adaptive average pooling and flattening on the feature tensor to obtain the flattened tensor feature, and determine the complexity score of the tensor feature based on the flattened tensor feature; Determine the adaptive rank based on the complexity score of the tensor feature, and obtain the optimized tensor decomposition method with the adaptive rank; Extract the diagonal elements of the second-order covariance matrix as channel weights, and map and reshape the channel weights to obtain the modulation coefficient tensor; Adjust the core tensor using the modulation coefficient tensor, and decompose the modulated core tensor based on the optimized tensor decomposition method to obtain the cropped modulated core tensor; Use the cropped modulated core tensor for reconstruction and fusion to obtain the fusion feature.
[0052] Specifically, after obtaining the local enhanced feature and the global attention feature , a low-rank fusion module is introduced, and the final feature fusion and compression are completed based on Tucker decomposition. For this purpose, first concatenate them in the channel dimension to obtain the feature tensor , and then adopt the form of trilinear tensor decomposition.
[0053] Second-order feature covariance representation: The low-rank fusion module first calculates the second-order statistical information of the input features, that is, the second-order features of the feature tensor, to capture the high-order correlations between features. Specifically, for the input features (feature tensor) , it is downsampled through an adaptive pooling operation with a fixed size, and then its covariance representation is calculated: (27); (28); (29); (30); (31); Wherein, is the feature processed by the adaptive average pooling operation ; is the adaptive average pooling operation; is reshaped into a two-dimensional tensor of shape reshape by the function, where is the number of feature channels, and -1 indicates automatic calculation of the remaining dimensions (all spatial positions); is the spatial size of the feature map after pooling; is the feature mean, and the calculation method is similar to formula (19), which is to average the feature values at all spatial positions ; is to access the value of all channels at position (i, j) through indexing; is the result after centering (subtracting the mean ); is the calculated covariance matrix, indicating the correlation between different channels; is the corresponding matrix.
[0054] This second-order covariance representation can capture the mutual relationships and high-order dependencies between feature channels, contains rich structural information and discriminative features, can provide richer structural information and texture descriptions, and is particularly suitable for processing medical images with complex structures and fine textures such as fractures. For fracture images, this high-order information is crucial for distinguishing subtle fracture textures and morphological changes.
[0055] For fracture images, the correlation and covariance information between feature channels often contain important diagnostic clues, especially for minor or atypical fractures. In this embodiment, by introducing a second-order feature processing module, such as the covariance calculation and diagonal weight extraction functions in the low-rank fusion module, and the BDC (Bilinear Descriptor Coding) bilinear feature coding module in the LCAD module, the model can capture the mutual relationships between feature channels and enhance the ability to identify subtle fracture features.
[0056] Feature complexity assessment and adaptive rank adjustment: In this embodiment, a feature complexity evaluator is innovatively designed, which adaptively adjusts the rank of Tucker decomposition by analyzing the statistical characteristics of the input features: (32); (33); (34); Among them, is the adaptive average pooling operation, and are the parameters of the fully connected layer, is the ReLU activation function, is Sigmoid activation function. is the tensor pooling feature, the feature information extracted from the feature tensor X through adaptive average pooling, which is used to evaluate the complexity of the feature; is the operation of flattening the multi-dimensional tensor into a one-dimensional vector, which is convenient for subsequent fully connected layer processing, is the tensor flattening feature, is the complexity score of the tensor feature.
[0057] The value output by the complexity evaluator ranges from 0 to 1, reflecting the complexity of the input features.
[0058] Based on the complexity score, the model dynamically adjusts the rank of Tucker decomposition: (35); Among them, is the base rank, is the minimum rank ratio (such as 0.5), is the adaptive rank, which can not only achieve efficient compression and fusion of features, but also realize the adaptive allocation of computing resources through the complexity evaluator; this enables the model to dynamically allocate computing resources according to the complexity of the input image, using a lower tensor rank for simple fracture images to improve computing efficiency, and using a higher tensor rank for complex fracture images to ensure feature expression ability, achieving a balance between computing efficiency and classification accuracy. This innovation is particularly important for medical image processing, because the fracture complexity varies greatly among different cases. Introducing the complexity evaluator and the complexity-based dynamic rank adjustment mechanism enables the model to adaptively adjust the used tensor rank and computing resources according to the complexity of the input image, improving computing efficiency while maintaining recognition accuracy.
[0059] Covariance-modulated Tucker decomposition: The present invention combines second-order feature representation with Tucker decomposition, and modulates the core tensor through covariance information to enhance the expression ability of key features. First, the diagonal elements of the covariance matrix are extracted as channel weights : (36); Then, these channel weights are mapped to the modulation coefficients of the core tensor through a fully connected layer : (37); (38); wherein, is the rank of the three dimensions of Tucker decomposition, are the parameters of the fully connected layer, is Sigmoid the activation function, and formula (38) reshapes the one-dimensional modulation coefficient into a modulation coefficient tensor with the same shape as the core tensor , facilitating subsequent element-wise modulation of the core tensor. Based on the Tucker decomposition theory, the input feature tensor can be decomposed into the product of a core tensor and three mode matrices: (39); wherein, is the core tensor, are the three mode matrices, represents the tensor product along the n th dimension.
[0060] The core tensor is adjusted through the modulation coefficient tensor to enhance the expression ability of key features, and the modulated core tensor is obtained: (40); wherein, represents element-wise multiplication, is the modulation intensity coefficient (such as 0.1), is the original core tensor. This covariance modulation mechanism enables the core tensor to adaptively emphasize the parts containing key fracture information according to the second-order statistical characteristics of the input features, improving the classification accuracy.
[0061] According to the adaptive rank , the model dynamically prunes the factor matrices and the core tensor: (41); (42); (43); where , , is the actual used rank dynamically calculated according to the adaptive rank and , , is the corresponding cropped mode matrix (factor matrix); is the cropped modulation core tensor.
[0062] Finally, the fused feature is calculated through Tucker reconstruction: (44); (45); where is the result of projecting the feature tensor X into the low-dimensional space; is the reconstructed fused feature, which contains the compressed representation of local and global information, and through second-order feature enhancement and adaptive rank adjustment, retains the most discriminative feature components and suppresses redundant channels.
[0063] After the first dense block and the feature fusion enhancement module are processed, it is input into the combination of the next dense block and the feature fusion enhancement module, and the final fused feature is output until all four dense blocks and the feature fusion enhancement modules are processed in sequence; Then, it is processed through the transition layer and the global pooling layer in sequence to obtain the corrected fused feature .
[0064] In this embodiment, by introducing a low-rank fusion mechanism, the high-dimensional feature tensor is decomposed into a core tensor and a series of factor matrices, and redundancy is reduced through orthogonal initialization constraints while key information is retained. In addition, the module modulates the core tensor through feature covariance, making the fusion process pay more attention to the part related to the second-order feature, thereby improving the feature expression ability while reducing the computational complexity. This embodiment can not only focus on capturing the fine-grained structure of the fracture area, but also construct the long-range dependence under the overall pattern of the bone; at the same time, through second-order feature enhancement and saliency guidance, the perception ability of the key fracture area is improved; low-rank fusion and adaptive rank adjustment significantly reduce the computational burden while ensuring the feature expression ability, enabling the model to have higher inference efficiency and generalization performance in the practical application of medical image analysis.
[0065] In the classification and prediction stage, the model adopts a multi-task learning strategy and uses the main classification head and the auxiliary classification head simultaneously: (46); (47); (48); (49); Wherein, is the main classification feature vector obtained by global average pooling (GAP) of the corrected fusion feature, and are the weight and bias parameters of the main classification head, is the predicted output of the main classification head. Similarly, is the feature vector of the auxiliary classification head obtained by global average pooling of the local enhanced feature in the last feature fusion enhancement module, and are the weight and bias parameters of the auxiliary classification head, is the predicted output of the auxiliary classification head. represents the global average pooling operation (Global Average Pooling). The final loss function is the weighted combination of the main classification head loss and the auxiliary classification head loss : (50); Wherein, is the weight coefficient (such as 0.4) used to balance the contributions of the two classification heads.
[0066] This embodiment introduces a multi-task learning strategy for the main classification head and the auxiliary classification head. Through the joint supervision of features at different levels, the feature learning ability and generalization performance of the model are enhanced. The model optimizes the two classification heads simultaneously during the training phase. The main classification head classifies based on the corrected fusion feature that has undergone complete feature fusion enhancement, while the auxiliary classification head directly uses the local enhanced feature in the intermediate processing process to provide additional gradient signals. The losses of the two are combined in a weight ratio of 0.4 to ensure that the main classification task has a higher priority while the intermediate features can also be effectively supervised. Through this multi-task learning strategy, the model can simultaneously learn useful information from features at different levels, enhancing the understanding ability of fracture features, improving the classification accuracy and model robustness. Especially when dealing with medical data sets with limited sample sizes, this multi-task learning can more effectively utilize the limited training data. In the actual inference phase, the model will intelligently use only the prediction results of the main classification head. This "rich training, lean inference" design not only improves the feature learning ability and generalization performance of the network but also maintains the efficiency and intuitiveness of the inference process, which is particularly suitable for medical image classification tasks such as fractures that require simultaneous attention to local details and global structures.
[0067] 4. Update network parameters through backpropagation for learning: The backpropagation mechanism is used to iteratively update the network parameters, aiming to make the fracture classification model continuously approach the optimal solution. Specifically, during the network training process, the initial learning rate is set to 1e-4, the batch size is 64, and the number of training epochs is 10. In each batch, the model calculates the prediction results through forward propagation and compares them with the true labels to obtain the cross-entropy loss. Then, the chain rule is used to solve the gradients of the loss with respect to the learnable parameters of each layer, and the Adam optimizer is used to update the parameters. After 10 training epochs, the model weights with the best performance on the validation set are selected as the final result, which can effectively improve the accuracy and robustness of fracture classification while maintaining a relatively fast convergence speed. In terms of the recognition rate of fracture classification, compared with the original DenseNet201 (80% accuracy), the present invention has increased by about 11 percentage points, and finally reaches a recognition rate of 91%, which is basically close to the human recognition level.
[0068] There are a wide variety of fracture types, including transverse fractures, spiral fractures, compression fractures, etc. The manifestation forms of each type of fracture are significantly different in images. Existing models often have difficulty effectively distinguishing different types of fractures when facing these complex fractures, resulting in a decrease in classification accuracy. In this embodiment, by combining multiple feature enhancement techniques, such as channel attention mechanism, multi-scale convolution processing, and low-rank fusion, the adaptability to different types of fractures is improved, enabling the model to more comprehensively capture the features of different types of fractures and improve classification accuracy and model robustness.
[0069] Efficiently integrating the feature fusion enhancement module into the existing backbone network while maintaining the training stability and inference efficiency of the network is an important challenge. In this embodiment, by selectively embedding the LCAD module in the Dense Block of DenseNet (by checking the number of features in the block) and combining the multi-task learning strategy of the main classification head and the auxiliary classification head, seamless integration of the feature enhancement module and the backbone network is achieved, while improving the model performance, maintaining training stability and computational efficiency. This design enables the feature enhancement module to make full use of the dense connection characteristics of DenseNet and effectively improve the feature expression ability without significantly increasing the model complexity. Compared with simply stacking additional modules or replacing the entire network structure, this embedding strategy not only retains the knowledge transfer advantage of the pre-trained model but also realizes the enhancement of key features, which is a design that takes into account both computational efficiency and classification performance.
[0070] In this embodiment, by innovatively combining techniques such as low-rank fusion, multi-branch local convolution, multi-scale dilated convolution, second-order feature processing, and adaptive complexity adjustment, the above technical problems are effectively solved, and an efficient and accurate fracture image classification method is proposed, providing a new technical solution for medical image analysis.
[0071] The present invention proposes a fracture classification method based on deep learning, aiming to improve the accuracy and robustness of fracture image classification through innovative feature extraction and fusion mechanisms. Based on the DenseNet201 backbone network, an innovative feature fusion enhancement module is embedded in its DenseBlock, focusing on multi-scale local feature extraction and global information capture. Through multi-branch local convolution and multi-scale dilated convolution, efficient extraction of fracture features at different scales is achieved. Through the Tucker decomposition technique, the high-dimensional feature tensor is decomposed into a core tensor and factor matrices, realizing feature compression and redundancy removal. At the same time, through covariance calculation and core tensor modulation, the second-order feature representation ability is enhanced, and the calculation resource allocation is dynamically adjusted according to the image complexity. This architecture design enables the model to simultaneously focus on tiny local fracture features and the overall bone structure when processing fracture images, and can adaptively adjust the feature extraction and fusion strategies according to the characteristics of different types of fractures, improving the classification accuracy. In addition, by introducing a multi-task learning mode with an auxiliary classification head and a main classification head, the feature learning ability and generalization performance of the model are further enhanced.
[0072] It can be understood that this embodiment is not limited to a specific backbone network. Although DenseNet201 is exemplarily used as the base network in the technical implementation process for initially extracting multi-level features of medical images, this backbone network is replaceable. That is to say, any deep network with good representation ability (such as ResNet, Inception, EfficientNet, etc.) can be combined with the LCDA (Local Context and Dependency Attention) module proposed by the present invention on the premise of keeping the core idea unchanged, so as to play an equivalent or similar effect in local feature enhancement, global dependency capture, and feature compression and fusion. The DenseNet201 network is only an exemplary selection, and the present invention is equally applicable to any other convolutional neural network or deep network with feature extraction ability, all within the protection scope of the present invention.
[0073] The fracture image classification technology of the present invention can be widely applied to products in the fields of medical image diagnosis, surgical navigation and evaluation, telemedicine, and big data education and research. For example, it can be embedded in the software systems of digital X-ray machines, CT machines, or MRI devices to achieve real-time assisted diagnosis of fracture detection and classification; it can also be integrated into telemedicine platforms or hospital information systems (PACS) to provide cloud-based fracture recognition and classification services for remote areas or inter-hospital collaborations; at the same time, it can be used for preoperative image evaluation and intraoperative assistance in orthopedic surgical planning systems to improve the accuracy of fracture reduction and internal fixation surgeries; and by combining with medical image management software or big data research platforms, it can provide rapid batch annotation and automated analysis support for pathological research and medical teaching, thus effectively promoting innovation and development in the field of intelligent medical imaging.
[0074] Embodiment 2 This embodiment provides a fracture image classification system with enhanced fusion features, including: An image preprocessing module configured to obtain fracture medical images and perform image preprocessing; An image classification module configured to classify based on the preprocessed fracture medical images using a pre-trained fracture image classification model. Specifically: After the preprocessed fracture images undergo initial convolution processing, they sequentially pass through four dense blocks and a feature fusion and enhancement module to obtain the final fusion features, and then pass through a transition layer and a global pooling layer to obtain corrected fusion features. Based on the corrected fusion features, fracture image classification is performed to obtain classification results; Among them, the processing process of each dense block and the feature fusion and enhancement module is as follows: Extract primary features based on the dense block; Enhance the primary features and then split them into local features and retained features along the channel dimension. After enhancing the local features, fuse them with the retained features to obtain locally enhanced features; Perform multi-scale dilated convolution and adaptive saliency guidance processing on the locally enhanced features, and fuse the processing results with the locally enhanced features to obtain globally focused features; Concatenate the locally enhanced features and the globally focused features to obtain a feature tensor. Based on the second-order feature optimization tensor decomposition method of the feature tensor, use the optimized tensor decomposition method to decompose and reconstruct the feature tensor to obtain fusion features.
[0075] The examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above Embodiment 1. It should be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.
[0076] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0077] The proposed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the above modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.
[0078] Embodiment III This embodiment provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps in a fracture image classification method with enhanced fusion features as described in Embodiment I above.
[0079] Embodiment IV This embodiment provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps in a fracture image classification method with enhanced fusion features as described in Embodiment I above.
[0080] Embodiment V This embodiment provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, which are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the steps in a fracture image classification method with enhanced fusion features as described in Embodiment I above.
[0081] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a hardware embodiment, a software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories and optical memories, etc.) containing computer-usable program code.
[0082] Although the specific implementation manners of the present invention are described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present invention. Those skilled in the art should understand that, based on the technical solutions of the present invention, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present invention.
Claims
1. A fracture image classification method with enhanced fusion features, characterized in that Including: Obtain fracture medical images and perform image preprocessing; Based on the preprocessed fracture medical images, use a pre-trained fracture image classification model for classification. Specifically: After the preprocessed fracture images are initially convolved, they are sequentially processed through four dense blocks and a feature fusion enhancement module to obtain the final fused features, and then processed through a transition layer and a global pooling layer to obtain corrected fused features. Based on the corrected fused features, fracture image classification is performed to obtain classification results; Among them, the processing process of each dense block and the feature fusion enhancement module is as follows: Extract primary features based on the dense block; Enhance the primary features and then split them into local features and retained features along the channel dimension. After enhancing the local features, fuse them with the retained features to obtain locally enhanced features; Perform multi-scale dilated convolution and adaptive saliency guidance processing on the locally enhanced features, and fuse the processing results with the locally enhanced features to obtain globally attentive features; Concatenate the locally enhanced features and the globally attentive features to obtain a feature tensor. Based on the second-order feature optimization tensor decomposition method of the feature tensor, use the optimized tensor decomposition method to decompose and reconstruct the feature tensor to obtain fused features.
2. The fracture image classification method with enhanced fusion features according to claim 1, characterized in that, The step of enhancing the primary features and then splitting them into local features and retained features along the channel dimension, enhancing the local features and then fusing them with the retained features to obtain locally enhanced features is specifically: After performing local context enhancement on the primary features, obtain primary extended features; Split the primary extended features into local features and retained features along the channel dimension according to a set ratio; Use multi-branch local convolution to divide the local features into two parts to obtain local features of different scales, and concatenate the local features of different scales in the channel dimension to obtain locally enhanced features after feature enhancement; Concatenate and fuse the locally enhanced features after feature enhancement with the retained features to obtain locally enhanced features.
3. The fracture image classification method with enhanced fusion features according to claim 1, wherein, Perform multi-scale dilated convolution and adaptive saliency guidance processing on the locally enhanced features, and fuse the processing results with the locally enhanced features to obtain globally attentive features. Specifically: Split the locally enhanced features along the channel dimension to obtain two locally enhanced sub-features; After performing small-scale conventional convolution and large-scale dilated convolution processing on the two locally enhanced sub-features respectively, obtain local enhanced sub-features of different scales and concatenate them to obtain global initial features; Use adaptive saliency guidance to process and then concatenate the global initial features separately from the channel and space to obtain globally attentive features.
4. The fracture image classification method with enhanced fusion features according to claim 3, wherein Use adaptive saliency guidance to process and then concatenate the global initial features separately from the channel and space to obtain globally attentive features. Specifically: Use global average pooling to capture the global context information of each channel of the global initial features, and learn the dependency relationships between all channels through non-linear transformation based on the global context information to obtain channel attention weights; Calculate channel-guided features using the channel attention weights and the global initial features; Downsample the global initial features to obtain downsampled features, and calculate the spatial covariance matrix of the downsampled features; Normalize the diagonal elements of the spatial covariance matrix based on the downsampled features to obtain spatial saliency weights, and perform a broadcast multiplication operation based on the spatial saliency weights and the downsampled features to obtain spatially modulated features; After upsampling the spatially modulated features, concatenate them with the local enhanced features and the channel-guided features to obtain globally attended features.
5. The fracture image classification method with enhanced fusion features as described in claim 1, characterized in that, Based on the second-order feature optimization tensor decomposition method of the feature tensor, use the optimized tensor decomposition method to decompose and reconstruct the feature tensor to obtain fused features. Specifically: Calculate the covariance matrix of the feature tensor based on its second-order features to obtain the second-order covariance matrix; Perform adaptive evaluation pooling and flattening on the feature tensor to obtain a flattened tensor feature, and determine the complexity score of the tensor feature based on the flattened tensor feature; Determine the adaptive rank based on the complexity score of the tensor feature, and obtain the optimized tensor decomposition method with the adaptive rank; Extract the diagonal elements of the second-order covariance matrix as channel weights, and map and reshape the channel weights to obtain a modulation coefficient tensor; Use the modulation coefficient tensor to adjust the core tensor, and decompose the modulated core tensor based on the optimized tensor decomposition method to obtain a cropped modulated core tensor; Use the cropped modulated core tensor for reconstruction and fusion to obtain fused features.
6. The method for classifying fracture images with enhanced fusion features according to claim 1, wherein, The feature fusion enhancement module includes a local feature enhancement sub-module and a local-global feature perception and fusion sub-module; The local feature enhancement sub-module includes a convolutional layer, a channel split layer, a multi-branch local convolutional layer, and a fusion layer; The local-global feature perception and fusion sub-module includes a multi-scale dilated convolutional layer, an adaptive saliency guidance layer, and a low-rank fusion and second-order feature enhancement layer.
7. A fracture image classification system with enhanced fusion features, characterized in that Comprises: An image preprocessing module configured to obtain a fracture medical image and perform image preprocessing; An image classification module configured to perform classification on the preprocessed fracture medical image using a pre-trained fracture image classification model. Specifically: After the preprocessed fracture image undergoes initial convolution processing, it sequentially passes through four dense blocks and a feature fusion enhancement module to obtain the final fused features, and then passes through a transition layer and a global pooling layer to obtain corrected fused features. Based on the corrected fused features, fracture image classification is performed to obtain a classification result; Among them, the processing process of each dense block and the feature fusion enhancement module is as follows: Extract primary features based on the dense block; Enhance the primary features and then split them into local features and retained features along the channel dimension. After enhancing the local features, fuse them with the retained features to obtain locally enhanced features; Perform multi-scale dilated convolution and adaptive saliency guidance processing on the locally enhanced features, and fuse the processing results with the locally enhanced features to obtain globally attended features; Concatenate the locally enhanced features and the globally attended features to obtain a feature tensor. Based on the second-order feature optimization tensor decomposition method of the feature tensor, use the optimized tensor decomposition method to decompose and reconstruct the feature tensor to obtain fused features.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the steps in a fracture image classification method with fused feature enhancement as described in any one of claims 1-6.
9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the steps in a fracture image classification method with enhanced fusion features as described in any one of claims 1-6 are implemented.
10. A computer program product, characterized in that, The computer program product includes a computer program which, when executed by a processor, implements the steps in a fracture image classification method with enhanced fusion features as described in any one of claims 1-6.
Citation Information
Patent Citations
New coronal pneumonia X-ray image classification method and system based on lightweight model
CN111931867A
Spine CT image classification method and system based on bidirectional covariance
CN118552799A
High-precision logging lithology intelligent identification method and system based on DenseNet-Transform depth fusion, and storage medium
CN119862475A
Medical image segmentation method based on multi-scale feature fusion
US20250095828A1
Cited By
Adaptive frequency domain enhanced electroencephalogram classification method and system
CN121465610A