Lumbar fracture image conversion method and system based on PR-GAN

Through the PR-GAN-based lumbar fracture image conversion method, X-ray images are converted into CT images, which solves the problem that X-ray imaging cannot provide detailed fracture information and the high radiation and cost of CT imaging in the prior art, and achieves efficient and accurate diagnosis and treatment options for lumbar fractures.

CN120147450AActive Publication Date: 2025-06-13YICHANG CENT PEOPLES HOSPITAL
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510207141.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-25
Publication Date
2025-06-13
Estimated Expiration
2045-02-25

AI Technical Summary

Technical Problem

In the prior art, when diagnosing lumbar fractures, X-ray imaging cannot provide the direction of micro fractures and different tomographic fracture lines, and cannot see the impact of fracture fragments on the spinal canal. CT imaging has problems such as high radiation dose and high cost.

Method used

Using the PR-GAN-based lumbar fracture image transformation method, the X-ray images of lumbar fractures are converted into CT images. By constructing a PR-GAN model, using the adversarial training of the generator and the discriminator, realistic CT images are generated, thereby combining the advantages of X-ray and CT to assist spinal surgeons in diagnostic and treatment choices.

Benefits of technology

It improves the efficiency of CT imaging in the imaging department, assists spinal surgeons in making accurate diagnosis and treatment options, reduces the radiation risk of patients and reduces medical expenses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147450A_ABST
    Figure CN120147450A_ABST
Patent Text Reader

Abstract

The invention provides a PR-GAN-based lumbar fracture image conversion method and system. The PR-GAN-based lumbar fracture image conversion method comprises the following steps: firstly, acquiring a lumbar fracture X-ray image and a lumbar fracture CT image of a patient; then preprocessing the acquired lumbar fracture X-ray image and the acquired lumbar fracture CT image, and constructing a training set; then, a PR-GAN model is constructed; inputting the training set into a PR-GAN model for training; and finally, inputting the lumbar fracture X-ray image which does not participate in training into the trained PR-GAN model, and converting the lumbar fracture X-ray image which does not participate in training into a target lumbar fracture CT image through the PR-GAN model. The PR-GAN model is combined with the advantages of X-ray and CT examination, so that the CT imaging efficiency of the imaging department is improved, and a spine surgeon can be assisted in diagnosis and treatment selection of lumbar vertebra fracture; through support and application of the kit, rapid diagnosis and rapid treatment can be realized, and the kit has very important positive significance on improvement of prognosis and life quality of patients.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a lumbar fracture image conversion method and system based on PR-GAN. Background Art

[0002] Lumbar fracture is a relatively common type of spinal injury, usually caused by external forces acting on the spine, which may lead to severe pain, limited movement, and even nerve damage. In severe cases, it may even affect life functions.

[0003] The causes of lumbar fractures are as follows:

[0004] 1. Traumatic causes:

[0005] Trauma is one of the most common causes of lumbar fractures. Car accidents, falls, high falls, sports accidents, or work injuries may all cause varying degrees of damage to the spine. In these cases, the pressure or force exerted on the spine exceeds its tolerance, resulting in fractures.

[0006] 2. Osteoporosis:

[0007] Osteoporosis is a disease that affects bone strength and density, making bones brittle and prone to injury. Even small forces during daily activities are sufficient to cause fractures, especially in important support areas such as the lumbar spine.

[0008] 3. Tumors or infections:

[0009] Spinal tumors or infections may erode the lumbar vertebrae, making them fragile and prone to fractures. The growth of tumors destroys the normal bone tissue structure, increasing the risk of bone damage.

[0010] The clinical manifestations of lumbar fractures vary depending on the type, location, and severity of the fracture. The following are the possible symptoms and characteristics:

[0011] 1. Severe pain:

[0012] The fracture site is usually accompanied by severe pain, especially when moving or applying pressure, and the pain may intensify. Lumbar fractures are often accompanied by low back pain.

[0013] 2. Limited movement:

[0014] Pain and lumbar instability may lead to limited movement, including restricted movements such as bending and rotating the torso.

[0015] 3. Spinal deformity:

[0016] Severe fractures may lead to changes in the spinal column's shape, presenting obvious deformities or malformations. Movements during daily activities may exacerbate the deformation, which may, in the long term, result in lumbar lordosis or kyphosis.

[0017] 4. Sensory and motor disturbances:

[0018] In the case of severe burst fractures, fracture fragments may compress the spinal cord or corresponding nerve roots at the back, potentially causing corresponding damage to the nervous system. For example, sensory abnormalities, paralysis, or muscle motor dysfunction. This may manifest as symptoms such as weakness, numbness, or pain in the lower limbs, or sensory abnormalities, urinary incontinence, bowel dysfunction, etc. in the perineal area.

[0019] 5. Increased risk of complications:

[0020] Fractures may lead to complications in other systems, especially in the elderly or those with other health problems, such as deep vein thrombosis, pressure ulcers, pulmonary infections, etc.

[0021] Thus, the hazards of lumbar fractures are significant, so their diagnosis and treatment are very important. Currently, the diagnosis of lumbar fractures requires the use of different imaging methods. Common diagnostic methods include X-rays, CT scans, and MRIs. These methods have their own imaging principles, advantages and disadvantages, and imaging manifestations, which can provide in-depth understanding of lumbar fractures and their imaging characteristics.

[0022] Among them, X-ray imaging is achieved by X-rays penetrating human tissues and being absorbed to varying degrees, and then captured by detectors to form images. Bone tissue absorbs more X-rays and appears as bright areas, while soft tissue absorbs less and appears as dark areas; the manifestations of lumbar fractures on X-ray images vary depending on the fracture type and severity. The following are common imaging manifestations: at the fracture site, there is a decrease in vertebral body height, and the anterior edge of the vertebral body shows a wedge-shaped deformation. In severe cases, the lumbar curvature can significantly protrude backward. Transverse fractures appear as horizontal linear fractures of the vertebral body, and the anterior and posterior edges of the vertebral body may be irregular, and the transverse fracture line can be seen. In burst fractures, the fractured vertebral body can be seen as multiple fragments, and the bone fragments can protrude backward to compress the spinal cord.

[0023] The advantages of X-ray imaging are: X-ray imaging is fast (scanning can be completed within one minute), convenient, and low-cost, and is commonly used for the preliminary assessment of fractures; its disadvantages are: it is not sensitive enough for early or atypical fractures, cannot show the soft tissue situation, cannot evaluate spinal cord compression, cannot assess whether the fracture site is fresh or old; DR requires exposure to low-dose radiation, which may not be suitable for specific populations such as pregnant women.

[0024] The current clinical application of X-ray imaging is: DR can quickly diagnose or rule out lumbar fractures, preliminarily assess the fracture type and lumbar stability, and provide a reference for subsequent treatment.

[0025] CT imaging is achieved through X-ray imaging technology, using a rotating X-ray source and detectors for imaging. The specific imaging principle is as follows: X-ray source: The rotating X-ray source irradiates the patient's body at different angles; Detectors: The detectors receive the X-ray signals that have been absorbed and attenuated by the body tissues at relative positions; Computer reconstruction: The computer performs calculations and reconstructions using mathematical algorithms based on the large amount of data collected, generating cross-sectional (tomographic) images; Multi-planar reconstruction: CT scans not only provide cross-sectional images but can also present a more comprehensive anatomical structure through multi-planar reconstruction (MPR) or three-dimensional reconstruction (3D); The manifestations of lumbar fractures on CT images vary depending on the fracture type and severity. The following are the common imaging findings: Changes in vertebral body height, compression fractures can lead to a decrease in vertebral body height, presenting as a wedge-shaped deformation of the anterior edge of the vertebral body. In severe cases, the lumbar curvature can significantly protrude backward; Transverse fractures are manifested as linear fractures across the vertebral body, which may be accompanied by pedicle fractures or lamina cracks; Burst fractures can show multiple fragments of the vertebral body. Sometimes, bone fragments can protrude and compress surrounding structures, and in severe cases, the spinal cord can be compressed, causing serious effects; Fractures can be accompanied by disc injuries, resulting in disc bulges or fissures.

[0026] The advantages of CT imaging are as follows: CT provides high-resolution three-dimensional images, showing the fracture conditions in more detail. It has higher sensitivity and specificity in diagnosing fractures, can evaluate the damage conditions of the anterior and posterior edges of the vertebral body, intervertebral discs, and pedicles, and helps in judging spinal stability; Three-dimensional anatomical display: Compared with DR, CT can perform multi-planar reconstruction, showing three-dimensional anatomical structures at multiple angles such as cross-section, coronal plane, and sagittal plane, which is helpful for evaluating the fracture location and stability. In addition, the CT scan speed is fast (it can be completed within one minute), and the imaging time is short, which is very useful for the rapid diagnosis of fractures in emergency situations; However, its disadvantages are: Compared with DR, the radiation dose of CT scans is relatively high, and long-term and frequent examinations may have a certain impact on the patient's health; It cannot show the soft tissue conditions, cannot evaluate the spinal cord compression situation, and cannot evaluate whether the fracture site is fresh or old; Contrast agent use: Some patients need to be injected with contrast agents, which may cause allergic reactions or kidney damage to those with renal insufficiency; Higher cost: Compared with X-ray examinations, CT scans are more expensive, which may increase medical expenses.

[0027] The current clinical applications of CT imaging are as follows: CT scans play an important role in the diagnosis and treatment of lumbar fractures. It can help doctors determine the fracture type, location, severity, evaluate the fracture healing situation, formulate treatment plans, and monitor the treatment effects. CT imaging plays an important role in the evaluation before and after surgery, the diagnosis of complex fractures, and the rehabilitation process.

[0028] It can be seen that for a patient with spinal fracture, clinicians usually use X-ray for preliminary diagnosis. Although X-ray has the advantages of being fast, convenient and efficient, due to image overlap and organ artifact interference, it cannot provide the situation of minor fractures and the orientation of fracture lines in different tomographies, nor can it show the impact of fracture fragments on the spinal canal. In the case of severe lumbar fractures, the fracture fragments may compress the posterior spinal cord, which can lead to severe lower limb nerve symptoms. If not relieved in time, it may lead to serious and irreversible consequences, having a serious impact on the prognosis and future quality of life of the patient. At this time, the support of CT images is needed. CT has higher image resolution and can generate three-dimensional images, which can help clinicians see the situation of fractures in different tomographies and make a more accurate diagnosis. At the same time, CT can clearly show soft tissues, which can help clinicians judge whether the fracture fragments protrude into the spinal canal and improve the efficiency of doctors in formulating surgical strategies. However, CT also has the disadvantages of high examination cost, long examination time, higher radiation dose than X-ray, and the potential increase in radiation risk with long-term multiple examinations.

[0029] This application provides a method and system for lumbar fracture image conversion based on PR-GAN under the above background, which converts lumbar fracture X-ray images into CT images, so as to combine the advantages of X-ray and CT, facilitating the rapid diagnosis of lumbar fractures by spinal surgeons, accurately selecting treatment plans, and being beneficial to the postoperative recovery and long-term prognosis of patients. Summary of the Invention

[0030] The purpose of the present invention is to provide a method and system for lumbar fracture image conversion based on PR-GAN to solve the problems existing in the above-mentioned prior art.

[0031] To achieve the above purpose, the present invention provides the following solutions:

[0032] The present invention provides a method for lumbar fracture image conversion based on PR-GAN, including the following steps:

[0033] S1. Collect lumbar fracture X-ray images and lumbar fracture CT images of patients;

[0034] S2. Preprocess the collected lumbar fracture X-ray images and lumbar fracture CT images to construct a training set;

[0035] S3. Construct a PR-GAN model;

[0036] S4. Input the training set into the PR-GAN model for training;

[0037] S5. Input the lumbar fracture X-ray images that have not participated in the training into the trained PR-GAN model, and convert the lumbar fracture X-ray images that have not participated in the training into target lumbar fracture CT images through the PR-GAN model.

[0038] Preferably, in step S2, the preprocessing includes:

[0039] S21. Crop the collected lumbar fracture X-ray images and lumbar fracture CT images;

[0040] S22. Unify the brightness and contrast of the lumbar fracture X-ray images and lumbar fracture CT images of different patients, and use the two-dimensional median filtering algorithm in open CV to perform noise reduction processing on the images;

[0041] S23. Pair the unified lumbar fracture X-ray images and lumbar fracture CT images of the patients to align the structures.

[0042] Preferably, in step S21, the standard for cropping the images is: Centering on the fractured vertebra, retain one vertebra above and below for cropping the image, and place the spinal canal in the middle position of the image. The size of the cropped picture is uniformly 512*512 dip.

[0043] Preferably, step S3 includes:

[0044] S31. Construct a residual network, and the objective function of the PR-GAN model is:

[0045]

[0046] S32. Introduce a position attention mechanism into the PR-GAN model to form a position attention mechanism-residual network model;

[0047] S33. Optimize the PR-GAN model.

[0048] Preferably, step S32 is: First, provide a local feature map X, generate three feature maps through a convolutional layer: Q, K, V, obtain the corresponding attention scores by performing matrix multiplication on Q and K; then apply the softmax activation function to calculate the corresponding spatial attention map and its related weights; finally, apply residual combination learning, multiply the output result by the learnable parameter gamma to make its attention score become 0 to 1, then perform dot product with V and sum to get the output, and then combine it with the original input feature map X to obtain the final output feature map.

[0049] Preferably, in step S33, the position attention mechanism-residual network model is added between the third and fourth downsampling layers of the discriminator, and spectral normalization is added at the convolutional layer of the discriminator.

[0050] Preferably, step S33 is as follows: First, input the X-ray image to be detected into the generator. The structure of each layer of downsampling in the generator is convolution, IN regularization, and the Leaky ReLU activation function. Before downsampling, an image enhancement operation is performed, and the ReflectionPad2d layer is introduced. After the first downsampling, the data becomes 64×256×256. After the second downsampling, the data becomes 128×128×128. After the third downsampling, the data becomes 256×64×64. At this time, it reaches the bottom. Nine residual modules are introduced between downsampling and upsampling to deepen the network and enhance the data at the same time. Then, an upsampling operation is performed. The structure of upsampling is transposed convolution, IN normalization, and the ReLU activation function to restore the size of the image, and the ReflectionPad2d layer is used for data enhancement. The first upsampling receives the data from the Resnet_block and changes the data to 256×64×64. The second upsampling changes the data to 128×128×128. After the third upsampling, the data becomes 64×256×256. The discriminator takes the data transmitted from the generator as input, and through the convolutional layer and LeakyReLU, changes the data to 64×128×128. Through three convolutions, IN regularization, and the Leaky ReLU activation function, the data is changed from 64×128×128 to 128×64×64 and then to 256×32×32. At this time, the PositionalAttention Resnet is added to mark and focus on the characteristics of the fracture part of the image, so as to generate an image that conforms to the CT characteristics. Finally, it reaches 512×31×31, and then the data is changed to 1×30×30 through a convolutional layer.

[0051] Preferably, step S4 is trained by an alternating training method.

[0052] Preferably, step S4 includes:

[0053] S41. Training of the generator: First, generate the corresponding output image by passing the input image through the generator. Then, calculate the loss function between the generated image and the target image, and backpropagate the error to update the parameters of the generator. Finally, during the training process, use the method of stochastic gradient descent to continuously update the parameters of the generator.

[0054] S42. Training of the discriminator: First, randomly select a set of real images and a set of generated images. Then, input these images into the discriminator, and calculate the probabilities that the real images and the generated images belong to the real data respectively. By calculating the loss functions corresponding to the real images and the generated images, as well as the overall loss function of the discriminator, and backpropagating the error to update the parameters of the discriminator. Finally, during the training process, use the method of stochastic gradient descent to continuously update the parameters of the discriminator.

[0055] The present invention also provides a lumbar fracture image conversion system based on PR-GAN, including:

[0056] An image acquisition module, configured to acquire X-ray images and CT images of lumbar fractures of patients;

[0057] An image preprocessing module, configured to preprocess the acquired X-ray images and CT images of lumbar fractures to construct a training set;

[0058] A model construction module, configured to construct a PR-GAN model;

[0059] A model training module, configured to input the training set into the PR-GAN model for training;

[0060] An image simulation module, which inputs the X-ray images of lumbar fractures that have not participated in training into the trained PR-GAN model, and converts the X-ray images of lumbar fractures that have not participated in training into target CT images of lumbar fractures through the PR-GAN model.

[0061] The present invention has achieved the following beneficial technical effects compared with the prior art:

[0062] A lumbar fracture image conversion method and system based on PR-GAN provided by the present invention first acquire X-ray images and CT images of lumbar fractures of patients; then preprocess the acquired X-ray images and CT images of lumbar fractures to construct a training set; then construct a PR-GAN model; input the training set into the PR-GAN model for training; finally, input the X-ray images of lumbar fractures that have not participated in training into the trained PR-GAN model, and convert the X-ray images of lumbar fractures that have not participated in training into target CT images of lumbar fractures through the PR-GAN model; by combining the advantages of X-ray and CT examinations through the PR-GAN model, not only the CT imaging efficiency of the imaging department is improved, but also it can assist spinal surgeons in the diagnosis and treatment selection of lumbar fractures; with the support and application of the present invention, rapid diagnosis and treatment can be carried out, which has very important positive significance for improving the prognosis and quality of life of patients. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0064] Figure 1 It is a picture after preprocessing of the present invention;

[0065] Figure 2This is the structural diagram of the PR-GAN of the present invention;

[0066] Figure 3 It is a comparison diagram of the synthetic CT and the real CT of a sagittal lumbar fracture;

[0067] Figure 4 It is a comparison diagram of the synthetic CT and the real CT of a coronal lumbar fracture. Specific embodiments

[0068] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0069] The purpose of the present invention is to provide a method and system for lumbar fracture image conversion based on PR-GAN to solve the problems existing in the prior art.

[0070] The present invention is a specific application of deep learning in medical imaging. Deep learning is a machine learning technology that enables a computer to automatically learn and understand data by mimicking the structure and function of the human brain neural network. It processes complex data through a multi-layer neural network model, performs high-level abstraction and understanding of the data, and then realizes the efficient processing and analysis of data such as images, voices, and texts. Common neural network structures in deep learning include: convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), long short-term memory network (LSTM), autoencoder (AE), attention mechanism (AM), etc.

[0071] Deep learning is a machine learning technology that enables a computer to automatically learn and understand data by mimicking the structure and function of the human brain neural network. It processes complex data through a multi-layer neural network model, performs high-level abstraction and understanding of the data, and then realizes the efficient processing and analysis of data such as images, voices, and texts. Common neural network structures in deep learning include: convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), long short-term memory network (LSTM), autoencoder (AE), attention mechanism (AM), etc.

[0072] Deep learning has the following advantages:

[0073] Automatic feature learning: The deep learning model can automatically learn features from data without manual feature extraction, and has stronger adaptability.

[0074] Processing large-scale data: It can efficiently process large-scale data, and its learning ability improves with the increase in the amount of data, making it suitable for a large number of medical imaging data in the medical field.

[0075] Improving prediction accuracy: In medical image analysis, deep learning models can provide prediction results with higher accuracy and precision.

[0076] Potential generality: It is applicable to different types of medical imaging data, such as X-rays, MRIs, CTs, etc., and has a certain degree of generality.

[0077] Interpretability: Some deep learning models have a certain degree of interpretability, which can help doctors understand the diagnosis process.

[0078] Currently, the applications of deep learning in medical imaging include:

[0079] Lesion segmentation: In medical imaging, deep learning can be used to locate and segment lesion areas. For example, in spinal surgery, there are papers reporting the segmentation of normal or diseased vertebral bodies.

[0080] Image reconstruction: Deep learning in medical imaging can be used to reconstruct three-dimensional images from planar images. For example, in spinal surgery, there are papers reporting that lumbar three-dimensional CT can be reconstructed from lumbar DR.

[0081] Image simulation: Currently, there are papers reporting the realization of single-modal or multi-modal imaging picture conversion simulation. For example, converting non-contrast CT pictures of the carotid artery and aorta into contrast CT pictures, and converting normal lumbar cross-sectional CT into MR pictures.

[0082] Disease-assisted diagnosis: Deep learning in medical imaging can be used to diagnose tumors, lung diseases, cardiovascular diseases, etc. For example, for the analysis of mammogram images for breast cancer screening, the auxiliary detection of lung nodules, and the detection and classification of tumors using CT or MRI scans.

[0083] The present invention specifically adopts a generative adversarial network (GAN). The purpose of the present invention is to convert lumbar fracture X-ray pictures into CT pictures and perform auxiliary diagnosis. X-rays and CTs are pictures with two different imaging principles. To achieve mutual conversion, it is naturally suitable to apply GAN for learning and training.

[0084] The GAN network consists of two parts: the Generator and the Discriminator. Among them, the Generator is responsible for converting a random noise vector into an image or other types of data similar to the training data, while the Discriminator is responsible for distinguishing the differences between the images generated by the Generator and the real data. The Generator and the Discriminator gradually optimize their network parameters through repeated games until they reach a balanced state - the Generator can generate highly realistic data, and the Discriminator cannot accurately distinguish whether the generated data is real or not. During the training process of the GAN, the Generator and the Discriminator confront each other and continuously optimize their performance in the game. Specifically, the Generator tries to deceive the Discriminator so that it cannot accurately distinguish between the generated data and the real data, while the Discriminator tries to identify the generated data and the real data as accurately as possible. In this way, during the process of backpropagating the error, the Generator can adjust its strategy of generating data according to the feedback of the Discriminator, thereby gradually generating more realistic data. The objective function in the GAN network is shown in the following figure:

[0085]

[0086] Among them and represent the probabilities of real data and generated data. It can be seen from the above loss function that the calculation of the loss function is generated in D (the Discriminator). Since the output of D is generally a True / Fake judgment, the binary cross-entropy function is generally adopted as a whole. The left side contains two parts: minG and maxD. For maxD, for 1 - D(G(z)), it is hoped that it can be as close as possible to 1, that is, the Discriminator can accurately distinguish the generated pictures. For minG, because it is desired that the generated data can pass the discrimination of the Discriminator, it is necessary to make D(G(z)) close to 1, that is, the generated data has to pass the Discriminator. Therefore, for it has to be as small as possible.

[0087] Among them, the Generator is a neural network that receives a random noise (latent space vector) as input and attempts to generate realistic data. It tries to generate data that can "fool" the Discriminator by repeatedly learning the characteristics of the data distribution.

[0088] The Discriminator is another neural network, similar to a binary classifier, which receives real data and fake data generated by the Generator and attempts to correctly classify them. The goal of the Discriminator is to be able to accurately distinguish between real data and fake data generated by the Generator.

[0089] The generator and the discriminator compete against and with each other. During the training process, the generator attempts to generate data that is realistic enough to "fool" the discriminator, while the discriminator endeavors to learn to accurately distinguish between real and fake data. This adversarial training process is iterated until the data generated by the generator is realistic enough to be indistinguishable by the discriminator.

[0090] The training process of GAN is as follows:

[0091] Initialize network parameters: The weights of the generator and the discriminator are usually randomly initialized.

[0092] Adversarial training loop: In each training iteration, the generator and the discriminator are alternately trained. For each training iteration: The generator receives random noise and generates data samples. The discriminator receives real data and fake data generated by the generator and makes classification judgments respectively. The discriminator calculates the loss (the accuracy of distinguishing real and fake data), and the generator calculates the loss (trying its best to "fool" the discriminator). The weights of the generator and the discriminator are updated through the loss function to improve the generation ability of the generator and the discrimination ability of the discriminator.

[0093] Convergence and generation: As the training progresses, the generator gradually learns the characteristics of the data distribution, and the generated data will become more and more realistic. When the data generated by the generator is realistic enough to be accurately distinguished by the discriminator, the model reaches the convergence state.

[0094] After convergence is completed, the generator that has learned sufficient distribution characteristics has reached the optimal state. We input the original picture into the generator, and the generator will use the learned characteristics to output the target picture.

[0095] The learning of GAN is too general and will widely learn the characteristics of each vertebral body and paravertebral area, which is not conducive to clinical needs because we need to focus on the fractured vertebral body and its surrounding soft tissue injuries to assist clinical diagnosis and treatment selection. Therefore, in this invention, we innovated a new GAN model, in which an attention mechanism is introduced to focus on the changes of the fractured vertebral body:

[0096] The attention mechanism is a way of thinking that mimics human visual attention and is used to enhance the attention degree of the neural network to different parts of the input. Its application in deep learning enables the model to pay more attention to important parts when processing sequence data (such as natural language processing, image processing, etc.), improving the performance and generalization ability of the model.

[0097] The attention mechanism is a technique often used in machine learning and deep learning to model the degree to which different parts of the input are attended to and assign different weights to them according to their importance.

[0098] In traditional neural networks, each input feature is treated equally regardless of its importance for the task. However, in many tasks, different input sequences or features may have different levels of importance, and the attention mechanism can help the model dynamically learn and focus on these important parts. The attention mechanism is commonly used to process sequential data, such as text or audio data in natural language processing (NLP), and image data in computer vision. The basic principle of the attention mechanism is as follows:

[0099] Input representation: First, the input data is transformed into a representation vector in some way. For text data, word embeddings or character embeddings can be used to represent words or characters; for image data, convolutional neural networks can be used to extract image features.

[0100] Generating attention weights: Next, by calculating the similarity between the input representation and the attention weights, an attention weight is assigned to each input position or feature. Common methods are to use dot products, additive methods, multi-layer perceptrons (MLPs), etc. to calculate the similarity.

[0101] Weight normalization: To ensure that the sum of the attention weights is 1, the weights need to be normalized. This can be achieved by using the softmax function, so that each weight value is between 0 and 1 and the sum is 1.

[0102] Generating the context vector: The attention weights are added to the input representation to generate a weighted context vector. This context vector can be regarded as a weighted summary of the input, where the information corresponding to the positions or features with larger weights is more important.

[0103] Output generation: According to the context vector and the requirements of the specific task, it is input into the subsequent model for further processing. The most common way is to input the context vector into a fully connected layer or other models to generate the final output.

[0104] By using the attention mechanism, the model can flexibly focus on the importance of different parts of the input, thereby improving the performance and performance of the model. It has achieved remarkable success in tasks in many fields, including machine translation, text summarization, image caption generation, etc. The attention mechanism enables the model to automatically learn and focus on the important parts of the input and assign appropriate weights to them, thereby improving the performance and generalization ability of the model.

[0105] The above are the principles and methods on which the present invention is based. To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0106] Example 1:

[0107] This embodiment provides a method for converting lumbar fracture images based on PR-GAN, including the following steps:

[0108] S1. Collect X-ray images and CT images of lumbar fractures of patients;

[0109] S2. Preprocess the collected X-ray images and CT images of lumbar fractures to construct a training set; to exclude the interference of surrounding soft tissues and enable the model to more accurately learn the characteristics of lumbar fractures, we cropped the X-ray and CT images of lumbar fractures. The criteria for cropping the images are as follows: one vertebral body above and below the fractured vertebral body as the center, and place the spinal canal in the middle of the image as much as possible. The size of the cropped images is uniformly 512*512 dip; in addition, to make the pixel values as uniform as possible for the learning of the model and the conversion of the images, we unified the brightness and contrast of the X-ray and CT images of different patients respectively, and used the two-dimensional median filtering algorithm in open CV to perform noise reduction processing on the images; finally, pair the X-ray and CT images of the same patient and align the structures as much as possible, Figure 1 as the image after preprocessing;

[0110] S3. Construct a PR-GAN model; specifically:

[0111] The generator of PR-GAN usually adopts an encoder-decoder architecture that combines convolutional layers and transposed convolutional layers, enabling it to capture global context and retain details such as Figure 2As shown, it consists of 9 Resnet blocks, and the ability of the network to learn complex features is enhanced by introducing skip connections between each pair of convolutional layers; ResNet has an advantage over the Unet structure in processing complex images, which is more conducive to the model to identify fractured or non-fractured vertebral bodies and convert them from X-rays to CT; the residual network proposed the concept of shortcut connection, which greatly increases the depth of the network and effectively solves the problem of gradient disappearance caused by excessive depth; in traditional convolutional neural networks, the output of the network layer is directly transformed by the activation function and then passed to the next layer; while in ResNet, the input of each network layer will not only be passed to the next layer, but also directly passed to several subsequent layers through the shortcut connection, forming a residual block; the advantage of this is that even if the number of network layers increases, due to the direct path, the information of the original input can be more easily propagated to the subsequent layers, avoiding the rapid attenuation of gradients in deep networks and the problem of gradient explosion; the residual block in ResNet consists of two or three convolutional layers, including an identity mapping and a non-linear activation function (such as ReLU); the identity mapping directly passes the input to the output, while the non-linear activation function introduces non-linear transformation; this design allows the network to learn the identity mapping when needed, that is, directly pass the input to the output, so as to better adapt to tasks with different complexities and difficulties;

[0112] PR-GAN is a variant of the CycleGAN model, and the objective function of the CycleGAN model can be expressed as:

[0113]

[0114] The goal of G is to minimize this objective function, and the goal of D is to counteract G. To facilitate the generation of clear ambiguities, the L1 distance metric is selected as the objective function, which can be expressed as:

[0115]

[0116] Therefore, the final objective function of PR-GAN can be expressed as:

[0117]

[0118] Inspired by NLP, the attention mechanism has been widely applied to various fields such as natural language inference, text representation, and image translation since it was proposed by Vaswani et al. Positional attention is an extension of self-attention that specifically deals with the positional information in sequential data. Different from traditional self-attention that only focuses on the relationships between elements in the input sequence, positional attention effectively integrates the positional information of elements. First, a local feature map X is provided, and three feature maps, Q, K, and V, are generated through a convolutional layer. By performing matrix multiplication on Q and K, the corresponding attention scores are obtained, and then the softmax activation function is applied to calculate the corresponding spatial attention map and its related weights. Finally, residual combination learning is applied. The output result is multiplied by the learnable parameter gamma to make its attention score range from 0 to 1. Then, it is multiplied by V and summed to obtain the output, which is then combined with the original input feature map X to obtain the final output feature map.

[0119] To find the model with the best optimization effect, we added PositionalAttention Resnet between the third and fourth downsampling layers of the discriminator, which is beneficial for the network model to better capture the features in the input image and improve the performance and imaging effect of the model. To further improve the performance of the model and reduce the computational overhead, spectral normalization was added at the convolutional layer of the discriminator.

[0120] First, the X-ray image to be detected is input into the generator. The structure of each downsampling layer in the generator is convolution, IN normalization, and Leaky ReLU activation function. Before downsampling, an image enhancement operation is first performed, and the ReflectionPad2d layer is introduced, which symmetrically enlarges the image along the edges up, down, left, and right to increase the image resolution. After the first downsampling, the data becomes 64×256×256, after the second downsampling, the data becomes 128×128×128, and after the third downsampling, the data becomes 256×64×64. At this time, it reaches the bottom. Nine residual modules are introduced between downsampling and upsampling to deepen the network and enhance the data at the same time. Then, the upsampling operation is performed. The role of upsampling is to magnify and restore the extracted features. The structure of upsampling is transposed convolution, IN normalization, and ReLU activation function to restore the image size, and the ReflectionPad2d layer is also used for data enhancement. The first upsampling receives the data from the Resnet_block and changes the data to 256×64×64. The second upsampling changes the data to 128×128×128, and after the third upsampling, the data becomes 64×256×256.

[0121] The discriminator takes the data passed from the generator as input, passes it through convolutional layers and LeakyReLU to transform the data into 64×128×128. Through three convolutions, IN normalization, and the Leaky ReLU activation function, the data is transformed from 64×128×128 to 128×64×64 and then to 256×32×32 (at this time, Positional Attention Resnet is added to mark and focus on the features of the fractured part of the image, enabling it to better generate images that conform to CT features), and finally to 512×31×31. Then, through a convolutional layer, the data is transformed into 1×30×30. In the network structure of the discriminator of our model, Spectral Normalization is used, which can improve the performance of the discriminator network, reduce unstable behaviors during training, enhance the convergence effect of the model, and thus improve the training effect of the model and the quality of the generated samples.

[0122] S4. Input the training set into the PR-GAN model for training; specifically:

[0123] For the two different networks of the discriminator and the generator, an alternating training method is used, that is, the entire training process is divided into two parts: the training of the generator and the training of the discriminator.

[0124] Training of the generator. In the training of the generator, the given input image is passed through the generator to generate the corresponding output image, and the generated image and the target image are sent to the discriminator for training together. The training process is first to pass the input image through the generator to generate the corresponding output image; then calculate the loss function between the generated image and the target image, and backpropagate the error to update the parameters of the generator; finally, during the training process, the parameters of the generator are continuously updated in a stochastic gradient descent manner, thereby gradually improving the performance of the generator.

[0125] Training of the discriminator. In the training of the discriminator, the discriminator needs to discriminate between the generated images and the target images to determine whether these images are real. The training process is first to randomly select a set of real images and a set of generated images; then input these images into the discriminator, and calculate the probabilities that the real images and the generated images respectively belong to the real data (i.e., the output values of the discriminator); calculate the loss functions corresponding to the real images and the generated images, as well as the overall loss function of the discriminator, and backpropagate the error to update the parameters of the discriminator; finally, during the training process, the parameters of the discriminator are also continuously updated in a stochastic gradient descent manner to improve its performance.

[0126] During the alternating training process, the generator and discriminator are repeatedly trained, and their parameters are continuously optimized until the model converges or reaches the preset number of training times; this alternating training method can effectively balance the performance of the generator and discriminator during the training process, enabling the final model to better complete the image translation task;

[0127] S5. Input the lumbar fracture X-ray images that have not participated in the training into the trained PR-GAN model, and convert the lumbar fracture X-ray images that have not participated in the training into target lumbar fracture CT images through the PR-GAN model.

[0128] This embodiment also provides a lumbar fracture image conversion system based on PR-GAN, including:

[0129] An image acquisition module for acquiring lumbar fracture X-ray images and lumbar fracture CT images of patients;

[0130] An image preprocessing module for preprocessing the acquired lumbar fracture X-ray images and lumbar fracture CT images to construct a training set;

[0131] A model construction module for constructing a PR-GAN model;

[0132] A model training module for inputting the training set into the PR-GAN model for training;

[0133] An image simulation module that inputs the lumbar fracture X-ray images that have not participated in the training into the trained PR-GAN model, and converts the lumbar fracture X-ray images that have not participated in the training into target lumbar fracture CT images through the PR-GAN model.

[0134] Figure 3 and Figure 4 shows the effect of the synthetic images of the present invention, Figure 3 In group A, it is a fracture of lumbar vertebra 1, in group B, it is a fracture of lumbar vertebra 2, in group C, it is a fracture of lumbar vertebra 3, in group D, it is a fracture of lumbar vertebra 4, and in group E, it is a fracture of lumbar vertebra 3; Figure 4 In group A and group B, it is a fracture of lumbar vertebra 1, and in group C, it is a fracture of lumbar vertebra 3.

[0135] The synthetic images show good similarity in both non-fractured vertebral bodies and non-fractured vertebral bodies; among the fractured vertebral bodies, the degree of vertebral body compression, whether the fracture protrudes into the spinal canal, and the degree of protrusion into the spinal canal are well shown; at the same time, subjective and objective evaluation methods are used to evaluate the image quality. The objective evaluation uses structural similarity (SSIM), and the value range of SSIM is 0-1. The closer to 1, the higher the similarity quality of the image; for subjective evaluation, 3 senior spine surgeons are invited to participate in the evaluation, and the fractured vertebral body in the sagittal CT is defined as the region of interest (SROI).

[0136] In the objective evaluation results, the SSIM between the sagittal CT synthesized by PR-GAN and the real CT reached 0.773, and the SSIM between the coronal CT synthesized by PR-GAN and the real CT reached 0.649; in the subjective evaluation results, the sensitivity of diagnosing sagittal fractures using the synthetic CT of PR-GAN reached 96.23%, and the specificity reached 97.94%. The sensitivity of diagnosing coronal fractures reached 95.12%, and the specificity reached 95.2%; the sensitivity of diagnosing fractures protruding into the spinal canal using the CT synthesized by PR-GAN reached 86.67%, and the specificity reached 98.33%.

[0137] The present invention elaborates on the principle and implementation manner of the present invention by applying specific examples. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A lumbar fracture image conversion method based on PR-GAN, characterized by: The following steps are involved: S1. Collect X-ray images and CT images of lumbar fracture of the patient; S2. preprocessing the collected lumbar fracture X-ray images and lumbar fracture CT images to construct a training set; S3. Build PR-GAN model; S4. Input the training set into the PR-GAN model for training; S5. Input the untrained lumbar fracture X-ray image into the trained PR-GAN model, and convert the untrained lumbar fracture X-ray image into the target lumbar fracture CT image through the PR-GAN model.

2. The lumbar fracture image conversion method based on PR-GAN according to claim 1 is characterized in that: In step S2, the preprocessing includes: S21. cropping the acquired lumbar fracture X-ray image and lumbar fracture CT image; S22. Unify the brightness and contrast of X-ray images and CT images of lumbar fractures of different patients, and use the two-dimensional median filter algorithm in open CV to reduce the noise of the images; S23. Pair the unified X-ray image of the patient's lumbar fracture and the CT image of the lumbar fracture to align the structures.

3. The lumbar fracture image conversion method based on PR-GAN according to claim 2 is characterized in that: In step S21, the standard for cropping the image is: cropping the image with the fractured vertebra as the center, retaining one vertebra above and below, and placing the spinal canal in the middle of the image, and the size of the cropped image is unified to 512*512dip.

4. The lumbar fracture image conversion method based on PR-GAN according to claim 1 is characterized in that: Step S3 includes: S31. Construct a residual network and obtain the objective function of the PR-GAN model: S32. Introduce the position attention mechanism into the PR-GAN model to form a position attention mechanism-residual network model; S33. Optimize the PR-GAN model.

5. The lumbar fracture image conversion method based on PR-GAN according to claim 4 is characterized in that: Step S32 is: first, provide a local feature map X, generate three feature maps Q, K, V through the convolution layer, and obtain the corresponding attention score by performing matrix multiplication on Q and K; then apply the softmax activation function to calculate the corresponding spatial attention map and its related weights; finally, apply residual combination learning, multiply the output result by the learnable parameter gamma to change its attention score to 0-1, and then obtain the output by performing dot multiplication with V and summing them, and then combine them with the original input feature map X to obtain the final output feature map.

6. The lumbar fracture image conversion method based on PR-GAN according to claim 4 is characterized in that: In step S33, the position attention mechanism-residual network model is added between the third and fourth downsampling layers of the discriminator, and spectral normalization is added at the convolution layer of the discriminator.

7. The lumbar fracture image conversion method based on PR-GAN according to claim 4 is characterized in that: Step S33 is: first, the X-ray image to be detected is input into the generator. The structure of each downsampling layer of the generator is convolution, IN regularization and Leaky ReLU activation function. Before downsampling, the image enhancement operation is performed first, and the ReflectionPad2d layer is introduced. After the first downsampling, the data becomes 64×256×256, after the second downsampling, the data becomes 128×128×128, and after the third downsampling, the data becomes 256×64×64. At this time, it reaches the bottom. Nine residual modules are introduced between downsampling and upsampling to deepen the network and enhance the data at the same time; then the upsampling operation is performed. The structure of the upsampling is deconvolution, IN normalization, ReLU activation, etc. The activation function is used to restore the size of the image, and the ReflectionPad2d layer is used for data enhancement. The first upsampling receives the data from Resnet_block and converts the data to 256×64×64. The second upsampling converts the data to 128×128×128. After the third upsampling, the data becomes 64×256×256. The discriminator uses the data from the generator as input and converts the data to 64×128×128 through the convolution layer and LeakyReLU. Through three convolutions, IN regularization and Leaky ReLU activation function, the data is converted from 64×128×128 to 128×64×64 and then to 256×32×32. At this time, PositionalAttention Resnet is added to mark and focus on the features of the fracture site in the image, so that it generates an image that meets the CT features, and finally to 512×31×31. Then, a convolution layer is used to convert the data to 1×30×30.

8. The lumbar fracture image conversion method based on PR-GAN according to claim 1 is characterized in that: Step S4 uses an alternating training method to perform training.

9. The lumbar fracture image conversion method based on PR-GAN according to claim 8, characterized in that: Step S4 includes: S41. Generator training: first, the input image is passed through the generator to generate the corresponding output image; then the loss function between the generated image and the target image is calculated, and the error is back-propagated to update the parameters of the generator; finally, during the training process, the parameters of the generator are continuously updated by using stochastic gradient descent; S42. The training of the discriminator first randomly selects a group of real images and a group of generated images; then these images are input into the discriminator, and the probabilities that the real images and the generated images belong to the real data are calculated; the parameters of the discriminator are updated by calculating the loss function corresponding to the real images and the generated images, as well as the overall loss function of the discriminator, and back-propagating the error; finally, during the training process, the parameters of the discriminator are continuously updated by using stochastic gradient descent.

10. A lumbar fracture image conversion system based on PR-GAN, characterized by: include: An image acquisition module, used for acquiring X-ray images and CT images of lumbar fractures of patients; An image preprocessing module is used to preprocess the collected lumbar fracture X-ray images and lumbar fracture CT images to construct a training set; Model building module, used to build PR-GAN model; Model training module, used to input the training set into the PR-GAN model for training; The image simulation module inputs the untrained lumbar fracture X-ray image into the trained PR-GAN model, and converts the untrained lumbar fracture X-ray image into the target lumbar fracture CT image through the PR-GAN model.

Citation Information

Patent Citations

  • Method for reconstructing CT (Computed Tomography) picture by utilizing biplane X-ray picture

    CN115719391A

  • X-ray reconstruction CT method based on depth separable convolution and parallel network architecture

    CN117911553A

  • Method for converting CT image of lumbar fracture into MR image

    CN118297859A

  • Device and method for obtaining reconstructed CT image and diagnosing fracture by using x-ray image

    WO2024080612A1