A PR-GAN-based method and system for converting lumbar vertebral fracture images

By using an image conversion method based on PR-GAN, X-ray images are converted into CT images, which solves the problems of low efficiency in X-ray diagnosis and high radiation in CT, enabling rapid and accurate diagnosis of lumbar vertebral fractures, improving diagnostic efficiency and patient health.

CN120147450BActive Publication Date: 2026-03-06YICHANG CENT PEOPLES HOSPITAL
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510207141.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2026-03-06
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

In the diagnosis of lumbar spine fractures, current technologies cannot detect microfractures or fracture lines at different sectional angles using X-ray imaging, while CT imaging suffers from high radiation and high cost, resulting in low diagnostic efficiency and potential risks to patient health.

Method used

A PR-GAN-based image conversion method was adopted to convert lumbar vertebral fracture X-ray images into CT images. By constructing a PR-GAN model and utilizing generative adversarial networks and position attention mechanisms, the conversion from X-ray images to CT images was achieved, thereby improving diagnostic efficiency and accuracy.

Benefits of technology

It improves the efficiency of CT imaging in radiology departments, assists spinal surgeons in quickly diagnosing lumbar fractures, reduces radiation exposure, lowers medical costs, and improves patient prognosis and quality of life.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147450B_ABST
    Figure CN120147450B_ABST
Patent Text Reader

Abstract

This invention provides a PR-GAN-based method and system for converting lumbar spine fracture images. First, X-ray and CT images of the lumbar spine fracture are acquired from the patient. Then, the acquired X-ray and CT images are preprocessed to construct a training set. Next, a PR-GAN model is constructed. The training set is input into the PR-GAN model for training. Finally, untrained X-ray images of the lumbar spine fracture are input into the trained PR-GAN model. The PR-GAN model then converts these untrained X-ray images into target CT images of the lumbar spine fracture. By combining the advantages of X-ray and CT examinations, the PR-GAN model not only improves the efficiency of CT imaging in radiology departments but also assists spinal surgeons in the diagnosis and treatment selection of lumbar spine fractures. With the support and application of this invention, rapid diagnosis and treatment can be achieved, which has significant positive implications for improving patient prognosis and quality of life.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of image processing technology, and in particular to a method and system for converting lumbar vertebral fracture images based on PR-GAN. Background Technology

[0002] Lumbar vertebral fracture is a relatively common type of spinal injury, usually caused by external force acting on the spine. It may cause severe pain, limited movement, or even nerve damage, and in severe cases, it may even affect vital functions.

[0003] The causes of lumbar vertebral fractures include the following:

[0004] 1. Traumatic causes:

[0005] Trauma is one of the most common causes of lumbar vertebral fractures. Car accidents, falls, falls from heights, sports accidents, or work-related injuries can all cause varying degrees of spinal damage. In these cases, the pressure or force on the spine exceeds its tolerance, leading to a fracture.

[0006] 2. Osteoporosis:

[0007] Osteoporosis is a condition that affects bone strength and density, making bones brittle and prone to injury. Even small forces during daily activities can be enough to cause fractures, especially in important support structures like the lumbar spine.

[0008] 3. Tumors or infections:

[0009] Spinal tumors or infections can erode the vertebral bodies of the lumbar vertebrae, making them fragile and prone to fracture. Tumor growth can also disrupt normal bone structure, increasing the risk of bone damage.

[0010] The clinical manifestations of lumbar vertebral fractures vary depending on the type, location, and severity of the fracture. The following are possible symptoms and characteristics:

[0011] 1. Severe pain:

[0012] Fractures are usually accompanied by severe pain, especially when exercise or stress is applied, which may worsen the pain. Lumbar vertebral fractures are often accompanied by lower back pain.

[0013] 2. Limited mobility:

[0014] Pain and lumbar instability can lead to limited mobility, including restricted movements such as bending over and rotating the torso.

[0015] 3. Spinal deformity:

[0016] Severe fractures can lead to changes in the shape of the spine, resulting in obvious deformities or deformities. Movements during daily activities can worsen these deformities, and over time, may lead to lumbar lordosis or kyphosis.

[0017] 4. Sensory and motor impairments:

[0018] In cases of severe burst fractures, bone fragments may compress the spinal cord or corresponding nerve roots, potentially causing neurological damage. This could manifest as sensory abnormalities, paralysis, or muscle dysfunction. Symptoms may include weakness, numbness, or pain in the lower limbs, perineal paresthesia, urinary incontinence, or bowel dysfunction.

[0019] 5. Increased risk of complications:

[0020] Fractures can lead to other systemic complications, especially in older adults or people with other health problems, such as deep vein thrombosis, pressure sores, and lung infections.

[0021] Therefore, lumbar vertebral fractures pose a significant risk, making their diagnosis and treatment crucial. Currently, the diagnosis of lumbar vertebral fractures requires the use of various imaging methods, with common methods including X-rays, CT scans, and MRI. These methods each have their own imaging principles, advantages and disadvantages, and imaging manifestations, allowing for a deeper understanding of lumbar vertebral fractures and their imaging characteristics.

[0022] X-ray imaging works by the absorption of radiation through human tissue to varying degrees, which is then captured by a detector to form an image. Bone tissue absorbs more X-rays and appears as a bright area, while soft tissue absorbs less and appears as a dark area. The appearance of lumbar vertebral fractures on X-ray images varies depending on the type and severity of the fracture. Common imaging findings include: a decrease in vertebral body height at the fracture site, a wedge-shaped deformation of the anterior edge of the vertebral body, and in severe cases, a significant posterior protrusion of the lumbar curvature. Transverse fractures present as a linear fracture of the vertebral body, with irregular anterior and posterior edges, and a visible transverse fracture line. Burst fractures show multiple fragments of the fractured vertebral body, which may protrude posteriorly and compress the spinal cord.

[0023] The advantages of X-ray imaging are: it is fast (a scan can be completed within one minute), convenient, and low in cost, and is routinely used for preliminary assessment of fractures; its disadvantages are: it is not sensitive enough for early or atypical fractures, it cannot show the condition of soft tissues, it cannot assess spinal cord compression, and it cannot assess whether the fracture site is fresh or old; DR requires exposure to low-dose radiation, and may not be suitable for certain groups such as pregnant women.

[0024] Currently, the clinical applications of X-ray imaging are as follows: DR can quickly diagnose or rule out lumbar vertebral fractures, preliminarily assess the fracture type and lumbar stability, and provide a reference for subsequent treatment.

[0025] CT imaging, on the other hand, uses X-ray imaging technology, employing a rotating X-ray source and detector to create images. The specific imaging principle is as follows: X-ray source: A rotating X-ray source irradiates the patient's body at different angles; Detector: The detector receives X-ray signals that have been absorbed and attenuated by body tissues at a relative position; Computer reconstruction: The computer uses mathematical algorithms to calculate and reconstruct the data collected, generating cross-sectional (tomographic) images; Multiplanar reconstruction: CT scans not only provide cross-sectional images, but can also present a more comprehensive anatomical structure through multiplanar reconstruction (MPR) or three-dimensional reconstruction (3D); The appearance of lumbar vertebral fractures on CT images varies depending on the type and severity of the fracture. The following are common imaging manifestations: Changes in vertebral body height: Compression fractures lead to a decrease in vertebral body height, presenting as a wedge-shaped deformation of the anterior edge of the vertebral body. In severe cases, the lumbar curvature can be significantly posteriorly protruding; Transverse fractures present as transverse linear fractures of the vertebral body, which may be accompanied by pedicle fractures or lamina cracks; Burst fractures can show multiple fragments of the vertebral body. Sometimes, bone fragments can protrude and compress surrounding structures, and in severe cases, they can compress the spinal cord, causing serious impact; Fractures can be accompanied by intervertebral disc damage, leading to disc bulging or fissures.

[0026] The advantages of CT imaging are: CT provides high-resolution three-dimensional images, showing fracture details more comprehensively, and has higher sensitivity and specificity in diagnosing fractures. It can assess damage to the anterior and posterior margins of the vertebral body, intervertebral discs, and pedicles, aiding in the assessment of spinal stability; 3D anatomical visualization: Compared to DR, CT can perform multi-planar reconstruction, displaying three-dimensional anatomical structures from multiple angles such as transverse, coronal, and sagittal planes, which helps assess fracture location and stability. Furthermore, CT scans are fast (completed within one minute), with short imaging time, making them very useful for rapid fracture diagnosis in emergency situations. However, its disadvantages are: Compared to DR, CT scans have a relatively higher radiation dose, and frequent, long-term examinations may have some impact on patient health; it cannot display soft tissue conditions, assess spinal cord compression, or determine the freshness or age of the fracture site; contrast agent use: contrast agent injection is required for some patients, which may cause allergic reactions or kidney damage in patients with renal insufficiency; higher cost: compared to X-ray examinations, CT scans are more expensive, potentially increasing medical costs.

[0027] Currently, the clinical applications of CT imaging are as follows: CT scans play an important role in the diagnosis and treatment of lumbar spine fractures. They can help doctors determine the type, location, and severity of the fracture, assess the fracture healing process, develop treatment plans, and monitor treatment effectiveness. CT imaging plays an important role in pre- and post-operative assessments, the diagnosis of complex fractures, and the rehabilitation process.

[0028] Therefore, for a patient with a spinal fracture, clinicians typically use X-rays for initial diagnosis. While X-rays are fast, convenient, and efficient, image overlap and organ artifacts prevent them from providing information on small fractures and the direction of fracture lines at different sections, and also fail to show the impact of fracture fragments on the spinal canal. In severe lumbar fractures, fracture fragments may compress the posterior spinal cord, leading to severe lower extremity neurological symptoms. If not relieved promptly, this can result in serious, irreversible consequences, severely impacting the patient's prognosis and future quality of life. In such cases, CT imaging is necessary. CT offers higher image resolution and can generate three-dimensional images, allowing clinicians to visualize the fracture at different sections and make a more accurate diagnosis. Simultaneously, CT clearly displays soft tissue, helping clinicians determine if fracture fragments protrude into the spinal canal, improving the efficiency of surgical strategy development. However, CT also has disadvantages such as higher examination costs, longer examination times, higher radiation doses than X-rays, and the potential for increased radiation risks with repeated examinations over a long period.

[0029] This application provides a PR-GAN-based method and system for converting lumbar fracture images into CT images, thereby combining the advantages of both X-ray and CT. This facilitates rapid diagnosis and accurate selection of treatment options for spinal surgeons of lumbar fractures, and also benefits postoperative recovery and long-term prognosis for patients. Summary of the Invention

[0030] The purpose of this invention is to provide a PR-GAN-based method and system for converting lumbar vertebral fracture images, in order to solve the problems existing in the prior art.

[0031] To achieve the above objectives, the present invention provides the following solution:

[0032] This invention provides a method for converting lumbar vertebral fracture images based on PR-GAN, comprising the following steps:

[0033] S1. Acquire X-ray and CT images of the patient's lumbar vertebral fracture;

[0034] S2. Preprocess the acquired X-ray and CT images of lumbar vertebral fractures to construct a training set;

[0035] S3. Construct the PR-GAN model;

[0036] S4. Input the training set into the PR-GAN model for training;

[0037] S5. Input the untrained lumbar vertebral fracture X-ray images into the trained PR-GAN model, and use the PR-GAN model to convert the untrained lumbar vertebral fracture X-ray images into target lumbar vertebral fracture CT images.

[0038] Preferably, in step S2, the preprocessing includes:

[0039] S21. Cropping the acquired X-ray and CT images of lumbar vertebral fractures;

[0040] S22. Unify the brightness and contrast of X-ray and CT images of lumbar fractures from different patients, and use the two-dimensional median filtering algorithm in OpenCV to denoise the images;

[0041] S23. Pair the unified X-ray and CT images of the patient's lumbar fracture to align the structures.

[0042] Preferably, in step S21, the standard for cropping the image is: cropping the image with one vertebra above and below the fractured vertebra as the center, and placing the spinal canal in the middle of the image, with the cropped image size uniformly set to 512*512dip.

[0043] Preferably, step S3 includes:

[0044] S31. Construct the residual network to obtain the objective function of the PR-GAN model:

[0045]

[0046] S32. Introduce a positional attention mechanism into the PR-GAN model to form a positional attention mechanism-residual network model;

[0047] S33. Optimize the PR-GAN model.

[0048] Preferably, step S32 is as follows: First, a local feature map X is provided, and three feature maps Q, K, and V are generated through a convolutional layer. By performing matrix multiplication on Q and K, the corresponding attention scores are obtained. Then, the softmax activation function is applied to calculate the corresponding spatial attention map and its related weights. Finally, residual combination learning is applied, and the output result is multiplied by the learnable parameter gamma to make its attention score 0 to 1. Then, the output is obtained by multiplying it with V and summing the results. Finally, the output is combined with the original input feature map X to obtain the final output feature map.

[0049] Preferably, in step S33, the positional attention mechanism-residual network model is added between the third and fourth downsampling layers of the discriminator, and spectral normalization is added at the convolutional layer of the discriminator.

[0050] Preferably, step S33 is as follows: First, the X-ray image to be detected is input into the generator. The downsampling structure of each layer of the generator is convolution, IN regularization, and Leaky ReLU activation function. Before downsampling, image enhancement is performed by introducing a ReflectionPad2d layer. After the first downsampling, the data becomes 64×256×256, after the second downsampling, the data becomes 128×128×128, and after the third downsampling, the data becomes 256×64×64. At this point, the bottom layer is reached. Nine residual modules are introduced between downsampling and upsampling to deepen the network and enhance the data. Then, upsampling is performed. The upsampling structure is deconvolution, IN normalization, and ReLU activation function. The liveness function restores the image size, and the ReflectionPad2d layer is used for data augmentation. The first upsampling receives data from the Resnet_block, transforming the data into 256×64×64. The second upsampling transforms the data into 128×128×128, and the third upsampling transforms the data into 64×256×256. The discriminator takes the data from the generator as input, passes it through a convolutional layer and LeakyReLU, transforming the data into 64×128×128. Through three convolutions, IN regularization, and the Leaky ReLU activation function, the data is transformed from 64×128×128 into 128×64×64 and then into 256×32×32. At this point, PositionalAttention Resnet is added to label and focus on the fracture site features in the image, generating an image that conforms to CT features, finally reaching 512×31×31. Then, through a convolutional layer, the data is transformed into 1×30×30.

[0051] Preferably, step S4 employs an alternating training method.

[0052] Preferably, step S4 includes:

[0053] S41. Generator training: First, the input image is used to generate the corresponding output image; then, the loss function between the generated image and the target image is calculated, and the error is backpropagated to update the generator parameters; finally, during the training process, stochastic gradient descent is used to continuously update the generator parameters.

[0054] S42. To train the discriminator, firstly, a set of real images and a set of generated images are randomly selected; then, these images are input into the discriminator, and the probabilities of the real images and generated images belonging to the real data are calculated; the loss functions corresponding to the real images and generated images, as well as the overall loss function of the discriminator, are calculated, and the discriminator parameters are updated by backpropagating the error; finally, during the training process, the discriminator parameters are continuously updated using stochastic gradient descent.

[0055] This invention also provides a PR-GAN-based image conversion system for lumbar vertebral fractures, comprising:

[0056] The image acquisition module is used to acquire X-ray images and CT images of lumbar vertebral fractures in patients.

[0057] The image preprocessing module is used to preprocess the acquired X-ray and CT images of lumbar vertebral fractures to construct a training set.

[0058] The model building module is used to build PR-GAN models;

[0059] The model training module is used to input the training set into the PR-GAN model for training;

[0060] The image simulation module inputs untrained lumbar vertebral fracture X-ray images into the trained PR-GAN model, and the PR-GAN model converts the untrained lumbar vertebral fracture X-ray images into target lumbar vertebral fracture CT images.

[0061] The present invention achieves the following beneficial technical effects compared to the prior art:

[0062] This invention provides a PR-GAN-based method and system for converting lumbar spine fracture images. First, X-ray and CT images of the lumbar spine fracture are acquired from the patient. Then, the acquired X-ray and CT images are preprocessed to construct a training set. Next, a PR-GAN model is constructed. The training set is input into the PR-GAN model for training. Finally, untrained X-ray images of the lumbar spine fracture are input into the trained PR-GAN model, which then converts these untrained X-ray images into target CT images of the lumbar spine fracture. By combining the advantages of X-ray and CT examinations using the PR-GAN model, not only is the CT imaging efficiency of radiology improved, but it also assists spinal surgeons in the diagnosis and treatment selection of lumbar spine fractures. With the support and application of this invention, rapid diagnosis and treatment can be initiated, which has significant positive implications for improving patient prognosis and quality of life. Attached Figure Description

[0063] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0064] Figure 1 This is a preprocessed image of the present invention;

[0065] Figure 2This is a structural diagram of the PR-GAN of the present invention;

[0066] Figure 3 Comparison of synthetic CT and real CT images of sagittal lumbar vertebral fractures;

[0067] Figure 4 This is a comparison image of a synthetic CT scan and a real CT scan for a coronal lumbar vertebral fracture. Detailed Implementation

[0068] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0069] The purpose of this invention is to provide a PR-GAN-based method and system for converting lumbar vertebral fracture images, in order to solve the problems existing in the prior art.

[0070] This invention is a specific application of deep learning in medical imaging. Deep learning is a machine learning technique that mimics the structure and function of the human brain's neural networks, enabling computers to automatically learn and understand data. It uses multi-layered neural network models to process complex data, performing high-level abstraction and understanding, thereby achieving efficient processing and analysis of data such as images, audio, and text. Common neural network structures used in deep learning include: Convolutional Neural Networks (CNN), Recurrent Neural Networks (RNN), Generative Adversarial Networks (GAN), Long Short-Term Memory Networks (LSTM), Autoencoders (AE), and Attention Mechanisms (AM).

[0071] Deep learning is a machine learning technique that mimics the structure and function of the human brain's neural networks, enabling computers to automatically learn and understand data. It uses multi-layered neural network models to process complex data, performing high-level abstraction and understanding, thereby achieving efficient processing and analysis of data such as images, speech, and text. Common neural network structures in deep learning include: Convolutional Neural Networks (CNNs), Recurrent Neural Networks (RNNs), Generative Adversarial Networks (GANs), Long Short-Term Memory Networks (LSTMs), Autoencoders (AEs), and Attention Mechanisms (AMs).

[0072] Deep learning has the following advantages:

[0073] Automatic feature learning: Deep learning models can automatically learn features from data without the need for manual feature extraction, making them more adaptable.

[0074] Processing large-scale data: It can efficiently process large-scale data, and its learning ability increases with the amount of data. It is suitable for large amounts of medical imaging data in the medical field.

[0075] Improved prediction accuracy: In medical image analysis, deep learning models can provide predictions with higher accuracy and precision.

[0076] Potential versatility: Applicable to different types of medical imaging data, such as X-rays, MRI, CT, etc., it has a certain degree of versatility.

[0077] Interpretability: Some deep learning models have a certain degree of interpretability, which can help doctors understand the diagnostic process.

[0078] Currently, deep learning is used in medical imaging in the following ways:

[0079] Lesion segmentation: In medical imaging, deep learning can be used to locate and segment lesion areas. For example, in spinal surgery, there are papers reporting the segmentation of normal or diseased vertebrae.

[0080] Image reconstruction: Deep learning can be used in medical imaging to reconstruct three-dimensional images from two-dimensional images. For example, in spinal surgery, there are already papers reporting the ability to reconstruct three-dimensional CT images of the lumbar spine from DR images.

[0081] Image simulation: There are already papers reporting the ability to convert and simulate single-modal or multi-modal imaging images, such as converting uncontrast CT images of the carotid artery and aorta into contrast-enhanced CT images, and converting normal lumbar spine transverse CT images into MR images.

[0082] Disease-aided diagnosis: Deep learning can be used in medical imaging to diagnose tumors, lung diseases, cardiovascular diseases, and more. Examples include mammogram image analysis for breast cancer screening, assisted detection of lung nodules, and tumor detection and classification for CT or MRI scans.

[0083] This invention specifically employs Generative Adversarial Network (GAN). The purpose of this invention is to convert X-ray images of lumbar vertebral fractures into CT images for auxiliary diagnosis. As X-ray and CT images are based on two different imaging principles, the conversion between them is naturally suitable for learning and training with GAN.

[0084] A GAN network consists of two parts: a generator and a discriminator. The generator transforms a random noise vector into an image or other type of data similar to the training data, while the discriminator distinguishes between the generated image and real data. Through repeated game-like interactions, the generator and discriminator gradually optimize their network parameters until an equilibrium is reached—the generator can produce highly realistic data, while the discriminator cannot accurately distinguish between real and generated data. During GAN training, the generator and discriminator compete against each other, continuously optimizing their performance in this game. Specifically, the generator attempts to confuse the discriminator, making it unable to accurately distinguish between generated and real data, while the discriminator tries to identify generated and real data as accurately as possible. During backpropagation of errors, the generator can adjust its data generation strategy based on feedback from the discriminator, gradually generating more realistic data. The objective function of the GAN network is shown in the figure below.

[0085]

[0086] in and This represents the probabilities of real and generated data. As seen in the loss function above, the loss function calculations are all performed within D (the discriminator). Since the output of D is generally a True / Fake judgment, a binary cross-entropy function is used overall. The left side contains two parts: minG and maxD. For maxD, we want 1-D(G(z)) to be as close to 1 as possible, so that the discriminator can accurately distinguish the generated image. For minG, because we want the generated data to pass the discriminator's judgment, D(G(z)) needs to be close to 1, meaning the generated data must pass the discriminator. Therefore, for... It should be as small as possible.

[0087] The generator is a neural network that takes random noise (a latent space vector) as input and attempts to generate realistic data. It learns the characteristics of the data distribution repeatedly, trying to generate data that can "fool" the discriminator.

[0088] The discriminator is another neural network, similar to a binary classifier, that receives real data and fake data generated by the generator and attempts to classify them correctly. The goal of the discriminator is to accurately distinguish between real data and fake data generated by the generator.

[0089] The generator and discriminator compete against each other. During training, the generator attempts to produce data realistic enough to "fool" the discriminator, while the discriminator strives to learn to accurately distinguish between real and fake data. This adversarial training process iterates until the data generated by the generator is realistic enough that the discriminator cannot distinguish it.

[0090] The training process of GAN is as follows:

[0091] Initialize network parameters: The weights of the generator and discriminator are usually initialized randomly.

[0092] Adversarial training loop: In each training iteration, the generator and discriminator are trained alternately. For each training iteration: the generator receives random noise and generates data samples. The discriminator receives real data and fake data generated by the generator, and classifies them accordingly. The discriminator calculates a loss (accuracy in distinguishing real and fake data), and the generator calculates a loss (trying to "fool" the discriminator). The weights of the generator and discriminator are updated using the loss function to improve the generator's generation ability and the discriminator's discrimination ability.

[0093] Convergence and Generation: As training progresses, the generator gradually learns the characteristics of the data distribution, and the generated data becomes increasingly realistic. When the data generated by the generator is realistic enough that the discriminator cannot accurately distinguish it, the model reaches convergence.

[0094] After convergence, the generator, having learned sufficient distribution features, has reached its optimal state. We input the original image into the generator, which will then use the learned features to output the target image.

[0095] GANs tend to overgeneralize, learning the characteristics of every vertebra and its surrounding tissues. This is detrimental to clinical needs, as we require a focused approach to fractured vertebrae and surrounding soft tissue injuries to aid in diagnosis and treatment selection. Therefore, this invention introduces a novel GAN ​​model that incorporates an attention mechanism to focus on changes within the fractured vertebrae.

[0096] Attention mechanisms are a way of thinking that mimics human visual attention, used to enhance the degree to which neural networks pay attention to different parts of the input. Its application in deep learning allows models to focus more on important parts when processing sequential data (such as natural language processing and image processing), improving model performance and generalization ability.

[0097] Attention mechanisms are a technique frequently used in machine learning and deep learning to allow models to focus on different parts of the input and assign them different weights based on their importance.

[0098] In traditional neural networks, each input feature is treated equally, regardless of its importance to the task. However, in many tasks, different input sequences or features may have different importance, and attention mechanisms can help the model dynamically learn and focus on these important parts. Attention mechanisms are commonly used to process sequential data, such as text or audio data in Natural Language Processing (NLP), and image data in Computer Vision. The basic principle of attention mechanisms is as follows:

[0099] Input representation: First, the input data is transformed into a representation vector in some way. For text data, word embeddings or character embeddings can be used to represent words or characters; for image data, convolutional neural networks can be used to extract image features.

[0100] Generating attention weights: Next, an attention weight is assigned to each input location or feature by calculating the similarity between the input representation and the attention weight. Common methods include using dot products, additive methods, and multilayer perceptrons (MLPs) to calculate similarity.

[0101] Weight normalization: To ensure that the sum of the attention weights is 1, the weights need to be normalized. This can be achieved using the softmax function, which makes each weight value between 0 and 1, and their sum equal to 1.

[0102] Context vector generation: Attention weights are added to the input representation to generate a weighted context vector. This context vector can be viewed as a weighted summary of the input, where positions or features with larger weights correspond to more important information.

[0103] Output generation: Based on the context vector and the requirements of the specific task, it is input into subsequent models for further processing. The most common approach is to input the context vector into fully connected layers or other models to generate the final output.

[0104] By using attention mechanisms, models can flexibly focus on the importance of different parts of the input, thereby improving model performance. It has achieved significant success in many domains, including machine translation, text summarization, and image caption generation. Attention mechanisms enable models to automatically learn and focus on important parts of the input and assign them appropriate weights, thus improving model performance and generalization ability.

[0105] The above describes the principles and methods upon which this invention is based. To make the above-mentioned objectives, features, and advantages of this invention more apparent and understandable, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0106] Example 1:

[0107] This embodiment provides a PR-GAN-based image conversion method for lumbar vertebral fractures, including the following steps:

[0108] S1. Acquire X-ray and CT images of the patient's lumbar vertebral fracture;

[0109] S2. The acquired X-ray and CT images of lumbar vertebral fractures were preprocessed to construct a training set. To eliminate interference from surrounding soft tissues and allow the model to more accurately learn the characteristics of lumbar vertebral fractures, the X-ray and CT images of lumbar vertebral fractures were cropped. The cropping criteria were as follows: one vertebra above and one below the fractured vertebra, with the spinal canal placed as centrally as possible in the image, and the cropped images were uniformly 512*512dip. In addition, to ensure pixel values ​​were as uniform as possible for model learning and image conversion, the brightness and contrast of the X-ray and CT images of different patients were unified, and the images were denoised using a two-dimensional median filter algorithm in OpenCV. Finally, the unified X-ray and CT images of patients were paired to align the structures as much as possible. Figure 1 The image after preprocessing;

[0110] S3. Construct the PR-GAN model; specifically:

[0111] PR-GAN generators typically employ an encoder-decoder architecture that combines convolutional and deconvolutional layers, enabling them to capture global context while preserving details such as... Figure 2As shown, it consists of 9 ResNet blocks, enhancing the network's ability to learn complex features by introducing skip connections between each pair of convolutional layers. ResNet has an advantage over the UNet structure in processing complex images, making it more suitable for models to identify fractured or non-fractured vertebrae and convert them from X-rays to CT scans. Residual networks introduce the concept of shortcut connections, which greatly increases the network depth and effectively solves the gradient vanishing problem caused by excessive depth. In traditional convolutional neural networks, the output of a network layer is directly transformed by the activation function and then passed to the next layer; however, in ResNet, the input of each network layer is not only passed to the next layer, but also passed through shortcut connections. The input is directly passed to subsequent layers via a direct path, forming a residual block. The advantage of this is that even as the number of network layers increases, the information from the original input can more easily propagate to later layers due to the direct path, avoiding the rapid decay and gradient explosion problems in deep networks. The residual block in ResNet consists of two or three convolutional layers, including an identity mapping and a non-linear activation function (such as ReLU). The identity mapping directly passes the input to the output, while the non-linear activation function introduces a non-linear transformation. This design allows the network to learn the identity mapping when needed, that is, to directly pass the input to the output, thus better adapting to tasks of varying complexity and difficulty.

[0112] PR-GAN is a variant of the CycleGAN model, and the objective function of the CycleGAN model can be expressed as:

[0113]

[0114] The objective of G is to minimize this objective function, and the objective of D is to counteract G. To avoid explicit ambiguity, the L1 distance metric is chosen as the objective function, which can be expressed as:

[0115]

[0116] Therefore, the final objective function of PR-GAN can be expressed as:

[0117]

[0118] Inspired by NLP, attention mechanisms, proposed by Vaswani et al., have been widely applied in various fields such as natural language reasoning, text representation, and image translation. Positional attention, an extension of self-attention, specifically handles positional information in sequential data. Unlike traditional self-attention, which only focuses on the relationships between elements in the input sequence, positional attention effectively integrates the positional information of elements. First, a local feature map X is provided, which generates three feature maps: Q, K, and V through convolutional layers. Matrix multiplication of Q and K yields the corresponding attention scores. Then, a softmax activation function is applied to calculate the corresponding spatial attention map and its associated weights. Finally, residual combination learning is applied, multiplying the output by the learnable parameter gamma to make its attention score between 0 and 1. This is then multiplied by V and summed to obtain the output, which is then combined with the original input feature map X to obtain the final output feature map.

[0119] To find the model with the best optimization effect, we added PositionalAttention ResNet between the third and fourth downsampling layers of the discriminator. This helps the network model to better capture the features in the input image, improving the model's performance and imaging effect. To further improve the model's performance and reduce computational overhead, spectral normalization was added to the convolutional layers of the discriminator.

[0120] First, the X-ray image to be detected is input into the generator. The downsampling structure of each layer of the generator is convolution, IN regularization, and Leaky. The ReLU activation function is used. Before downsampling, an image enhancement operation is performed by introducing a ReflectionPad2d layer, which symmetrically adjusts the image along the edges to increase resolution. After the first downsampling, the data becomes 64×256×256; after the second downsampling, it becomes 128×128×128; and after the third downsampling, it becomes 256×64×64. At this point, nine residual modules are introduced between downsampling and upsampling to deepen the network and enhance the data. Then, upsampling is performed to magnify and restore the extracted features. The upsampling structure consists of deconvolution, IN normalization, and the ReLU activation function to restore the image size, and the ReflectionPad2d layer is also used for data enhancement. The first upsampling receives data from the Resnet_block, making the data 256×64×64; the second upsampling makes the data 128×128×128; and the third upsampling makes the data 64×256×256.

[0121] The discriminator takes the data from the generator as input, passes it through convolutional layers and Leaky ReLU to transform it into 64×128×128. Then, through three convolutions, IN regularization, and the Leaky ReLU activation function, the data is transformed from 64×128×128 to 128×64×64, then to 256×32×32 (at this point, Positional Attention ResNet is added to label and focus on the fracture site features in the image, enabling it to better generate images that conform to CT features), finally reaching 512×31×31. Finally, a convolutional layer transforms the data into 1×30×30. Our model uses Spectral Normalization in the discriminator's network structure, which improves the discriminator network's performance, reduces instability during training, enhances model convergence, and thus improves the model's training effect and the quality of generated samples.

[0122] S4. Input the training set into the PR-GAN model for training; specifically:

[0123] For two different networks, discriminator and generator, an alternating training method is used, which divides the entire training process into two parts: generator training and discriminator training.

[0124] The generator training process involves taking a given input image and generating a corresponding output image. The generated image and the target image are then fed into the discriminator for training. The training process begins by generating the corresponding output image from the input image. Then, the loss function between the generated and target images is calculated, and the error is backpropagated to update the generator's parameters. Finally, during training, stochastic gradient descent is used to continuously update the generator's parameters, thereby gradually improving the generator's performance.

[0125] The discriminator is trained by distinguishing between generated and target images to determine whether they are real. The training process begins by randomly selecting a set of real images and a set of generated images. These images are then input into the discriminator, and the probabilities of each image belonging to the real data are calculated (i.e., the discriminator's output). The discriminator's parameters are updated by calculating the loss functions for the real and generated images, as well as the overall loss function of the discriminator, and backpropagating the error. Finally, during training, stochastic gradient descent is used to continuously update the discriminator's parameters, thereby improving its performance.

[0126] During alternating training, the generator and discriminator are trained repeatedly, and their parameters are continuously optimized until the model converges or reaches the preset number of training iterations. This alternating training method can effectively balance the performance of the generator and discriminator during training, enabling the final model to better complete the image translation task.

[0127] S5. Input the untrained lumbar vertebral fracture X-ray images into the trained PR-GAN model, and use the PR-GAN model to convert the untrained lumbar vertebral fracture X-ray images into target lumbar vertebral fracture CT images.

[0128] This embodiment also provides a PR-GAN-based lumbar vertebral fracture image conversion system, including:

[0129] The image acquisition module is used to acquire X-ray images and CT images of lumbar vertebral fractures in patients.

[0130] The image preprocessing module is used to preprocess the acquired X-ray and CT images of lumbar vertebral fractures to construct a training set.

[0131] The model building module is used to build PR-GAN models;

[0132] The model training module is used to input the training set into the PR-GAN model for training;

[0133] The image simulation module inputs untrained lumbar vertebral fracture X-ray images into the trained PR-GAN model, and the PR-GAN model converts the untrained lumbar vertebral fracture X-ray images into target lumbar vertebral fracture CT images.

[0134] Figure 3 and Figure 4 This demonstrates the effect of synthesizing images using the present invention. Figure 3 Group A consisted of L1 fractures, Group B consisted of L2 fractures, Group C consisted of L3 fractures, Group D consisted of L4 fractures, and Group E consisted of L3 fractures. Figure 4 Group A and Group B had L1 fractures, and Group C had L3 fractures.

[0135] The synthesized images showed good similarity in both non-fractured and non-fractured vertebrae. In fractured vertebrae, the degree of vertebral compression, whether the fracture protruded into the spinal canal, and the extent of protrusion were clearly shown. Both subjective and objective evaluation methods were used to assess image quality. The objective evaluation employed structural similarity (SSIM), with values ​​ranging from 0 to 1; values ​​closer to 1 indicated higher image similarity quality. The subjective evaluation involved three senior spine surgeons, defining the fractured vertebrae in sagittal CT as regions of interest (SROIs).

[0136] In the objective evaluation results, the SSIM of PR-GAN synthesized sagittal CT compared to real CT reached 0.773, and the SSIM of PR-GAN synthesized coronal CT compared to real CT reached 0.649. In the subjective evaluation results, the sensitivity and specificity of PR-GAN synthesized CT in diagnosing sagittal fractures reached 96.23%, and the sensitivity and specificity in diagnosing coronal fractures reached 95.12% and 95.2%, respectively. The sensitivity and specificity of PR-GAN synthesized CT in diagnosing fractures protruding into the spinal canal reached 86.67% and 98.33%, respectively.

[0137] This invention has illustrated its principles and implementation methods using specific examples. The descriptions of these embodiments are merely illustrative of the method and its core ideas; furthermore, those skilled in the art will recognize that modifications may be made to the specific implementation methods and application scope based on the principles of this invention. Therefore, the content of this specification should not be construed as limiting the invention.

Claims

1. A PR-GAN-based lumbar vertebra fracture image conversion method, characterized in that: The method comprises the following steps: S1. Collecting X-ray images and CT images of lumbar vertebra fracture of a patient; S2. Preprocessing the collected X-ray images and CT images of lumbar vertebra fracture to construct a training set; S3. Constructing a PR-GAN model; step S3 comprises: S31. Constructing a residual network, and obtaining a target function of the PR-GAN model as follows: ; S32. Introducing a position attention mechanism into the PR-GAN model to form a position attention mechanism-residual network model; S33. Optimizing the PR-GAN model; In step S32, a local feature mapping X is provided first, three feature mappings Q, K and V are generated through a convolution layer, and corresponding attention scores are obtained by performing matrix multiplication on Q and K; then a softmax activation function is applied to calculate corresponding spatial attention mappings and related weights; finally, residual combination learning is applied, the output result is multiplied by a learnable parameter gamma, so that the attention score becomes 0-1, then the output is obtained by performing point multiplication and summation on V again, and the final output feature mapping is obtained by combining the output with the original input feature mapping X; In step S33, the position attention mechanism-residual network model is added between the third and fourth down-sampling layers of the discriminator, and spectral normalization is added at the convolution layer of the discriminator. Step S33 is: first, the X-ray picture to be detected is input into the generator, the down-sampling structure of each layer of the generator is convolution, IN regularization and Leaky ReLU activation function, image enhancement is performed before down-sampling, a ReflectionPad2d layer is introduced, after the first down-sampling, the data becomes 64*256*256, after the second down-sampling, the data becomes 128*128*128, after the third down-sampling, the data becomes 256*64*64, at this time, the bottom is reached, 9 residual modules are introduced between down-sampling and up-sampling to deepen the network while enhancing the data; then, up-sampling operation is performed, the up-sampling structure is deconvolution, IN standardization, ReLU activation function to restore the size of the image, and the ReflectionPad2d layer is used for data enhancement, the first up-sampling receives data from the Resnet_block, and the data becomes 256*64*64, the second up-sampling changes the data to 128*128*128, and the third up-sampling changes the data to 64*256*256; the discriminator takes the data from the generator as input, and changes the data to 64*128*128 through a convolution layer and LeakyReLU, and changes the data from 64*128*128 to 128*64*64 and then to 256*32*32 through three convolution, IN regularization and Leaky ReLU activation functions, at this time, the Positional Attention Resnet is added to mark and pay attention to the image fracture part features, so as to generate an image conforming to the CT features, and finally to 512*31*31, and then to 1*30*30 through a convolution layer; S4. Input the training set into the PR-GAN model for training; S5. Input the lumbar fracture X-ray image not participating in the training into the PR-GAN model after the training, and convert the lumbar fracture X-ray image not participating in the training into a target lumbar fracture CT image through the PR-GAN model.

2. The PR-GAN-based lumbar vertebra fracture image conversion method of claim 1, wherein: In step S2, the preprocessing includes: S21. The collected lumbar fracture X-ray images and lumbar fracture CT images are cropped; S22. The brightness and contrast of the lumbar fracture X-ray images and lumbar fracture CT images of different patients are unified, and the images are denoised using the two-dimensional median filtering algorithm in open CV; S23. The unified lumbar fracture X-ray images and lumbar fracture CT images of the patients are paired to align the structure.

3. The PR-GAN-based lumbar vertebra fracture image conversion method according to claim 2, characterized in that: In step S21, the standard for cropping the image is: taking the fractured vertebral body as the center, retaining one vertebral body above and below the fractured vertebral body, and placing the spinal canal in the middle of the image, and the size of the cropped image is unified to 512*512dip.

4. The PR-GAN-based lumbar vertebra fracture image conversion method of claim 1, wherein: Step S4 adopts an alternating training method for training.

5. The PR-GAN-based lumbar vertebra fracture image conversion method according to claim 4, characterized in that: Step S4 includes: S41. The training of the generator, first, the input image is generated by the generator to generate the corresponding output image; then the loss function between the generated image and the target image is calculated, and the error is propagated to update the parameters of the generator; finally, in the training process, the parameters of the generator are constantly updated in the way of stochastic gradient descent; S42. The training of the discriminator, first, a set of real images and a set of generated images are randomly selected; then the images are input into the discriminator, and the probability that the real image and the generated image belong to the real data is calculated; the loss function corresponding to the real image and the generated image, and the loss function of the whole discriminator are calculated, and the parameters of the discriminator are updated by backpropagation error; finally, in the training process, the parameters of the discriminator are constantly updated in the way of stochastic gradient descent. 6.A PR-GAN based lumbar vertebrae fracture image translation system, characterized in that: It comprises: An image acquisition module for acquiring X-ray images and CT images of lumbar vertebral fractures of patients; An image preprocessing module for preprocessing the acquired X-ray images and CT images of lumbar vertebral fractures to construct a training set; A model construction module for constructing a PR-GAN model; the construction method comprises: Constructing a residual network to obtain the objective function of the PR-GAN model: ; Introducing a position attention mechanism into the PR-GAN model to form a position attention mechanism-residual network model; Optimizing the PR-GAN model; Introducing a position attention mechanism into the PR-GAN model to form a position attention mechanism-residual network model is: first, a local feature mapping X is provided, three feature mappings: Q, K, and V are generated through a convolution layer, the corresponding attention scores are obtained by matrix multiplication on Q and K; then the corresponding spatial attention mapping and its related weights are calculated by applying the softmax activation function; finally, the residual combination learning is applied, the output result is multiplied by the learnable parameter gamma, so that its attention score becomes 0~1, then the output is obtained by point multiplication and summation with V, and then combined with the original input feature mapping X to obtain the final output feature mapping; In the optimization of the PR-GAN model, the position attention mechanism-residual network model is added between the third and fourth down-sampling layers of the discriminator, and the spectral normalization is added at the convolution layer of the discriminator; The PR-GAN model is optimized as follows: first, input the X-ray picture to be detected into the generator, the down-sampling structure of each layer of the generator is convolution, IN regularization and Leaky ReLU activation function, before down-sampling, image enhancement operation is performed first, a ReflectionPad2d layer is introduced, after the first down-sampling, the data becomes 64*256*256, after the second down-sampling, the data becomes 128*128*128, after the third down-sampling, the data becomes 256*64*64, at this time, the bottom is reached, 9 residual modules are introduced between down-sampling and up-sampling to deepen the network while enhancing the data; then, up-sampling operation is performed, the up-sampling structure is deconvolution, IN standardization, ReLU activation function to restore the size of the image, and ReflectionPad2d layer is used for data enhancement, the first up-sampling receives data from the Resnet_block, the data becomes 256*64*64, the second up-sampling changes the data to 128*128*128, and the third up-sampling changes the data to 64*256*256; the discriminator takes the data from the generator as input, changes the data to 64*128*128 through a convolution layer and LeakyReLU, changes the data from 64*128*128 to 128*64*64 and then to 256*32*32 through three convolution, IN regularization and Leaky ReLU activation functions, at this time, the Positional Attention Resnet is added to mark and pay attention to the image fracture part features, so as to generate an image conforming to the CT features, finally to 512*31*31, then through a convolution layer, the data is changed to 1*30*30; The model training module is used for inputting the training set into the PR-GAN model for training; The image simulation module inputs the lumbar fracture X-ray image not participating in the training into the PR-GAN model trained, and converts the lumbar fracture X-ray image not participating in the training into a target lumbar fracture CT image through the PR-GAN model.