A lumbar disc herniation image recognition conversion method and system

By using deep learning technology and generative adversarial networks, the ITAD-pix model was constructed to convert CT images of lumbar disc herniation into MRI images. This solved the difficulty of combining imaging methods in lumbar disc herniation, enabling efficient diagnosis and treatment selection, and improving diagnostic accuracy and patient recovery.

CN120182404BActive Publication Date: 2026-01-23YICHANG CENT PEOPLES HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510238548.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2026-01-23
Estimated Expiration
2045-03-03

AI Technical Summary

Technical Problem

Existing imaging methods for diagnosing lumbar disc herniation, CT and MRI each have their own advantages and disadvantages, and cannot be effectively combined, leading to difficulties in the selection of diagnostic and treatment options. Furthermore, image conversion technology is not yet mature in the field of lumbar disc herniation.

Method used

By employing deep learning techniques, particularly generative adversarial networks (GANs) combined with an iterative normalization generation mechanism, the ITAD-pix model is constructed to convert CT images of lumbar disc herniation into MRI images. Through training and conversion using the ITAD-pix model, mutual conversion between CT and MRI images can be achieved, assisting in diagnosis and the selection of treatment plans.

Benefits of technology

It enables high-quality conversion of CT images to MRI images, helping spinal surgeons to quickly and accurately diagnose lumbar disc herniation, providing better treatment options, and improving diagnostic efficiency and postoperative recovery for patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182404B_ABST
    Figure CN120182404B_ABST
Patent Text Reader

Abstract

The application provides a lumbar disc herniation image recognition conversion method and system, first, CT images and MRI images of lumbar disc herniation of a patient are collected; then the collected CT images and MRI images of lumbar disc herniation are preprocessed to construct a training set; then an ITAD-pix model is constructed; the training set is input into the ITAD-pix model for training; finally, the CT images of lumbar disc herniation not participating in the training are input into the trained ITAD-pix model, and the CT images of lumbar disc herniation not participating in the training are converted into target MRI images of lumbar disc herniation through the ITAD-pix model; the CT images of lumbar disc herniation are converted into MRI images, so as to combine the advantages of CT and MRI, which is beneficial to rapid diagnosis of lumbar disc herniation, accurate selection of a treatment scheme, postoperative recovery and long-term prognosis of a patient.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to a lumbar disc herniation image recognition and conversion method and system. BACKGROUND

[0002] Lumbar disc herniation refers to the displacement of the contents of the intervertebral disc (nucleus pulposus) through its outer membrane (annulus fibrosus), which usually occurs in the posterolateral region of the intervertebral disc. Depending on the volume of the herniated material, it can compress and irritate the lumbar nerve roots and dural sac. Lumbar disc herniation is the most common diagnosis in lumbar degenerative abnormalities and is the leading cause of adult spinal surgery. Typical clinical manifestations include initial low back pain, followed by progressive sciatica. For patients with lumbar disc herniation, diagnosis is made through magnetic resonance imaging (MRI), computed tomography (CT) combined with clinical symptoms and physical examination. MRI, as the gold standard for diagnosis, has many limitations; early diagnosis and active treatment have significant implications for improving patient outcomes.

[0003] Intervertebral disc degeneration is the root cause of lumbar disc herniation. With age, the intervertebral disc gradually degenerates, the water content of the annulus fibrosus and nucleus pulposus gradually decreases, the nucleus pulposus loses elasticity, and the annulus fibrosus gradually develops fissures. Under the influence of accumulated strain and external forces, the intervertebral disc ruptures, and the nucleus pulposus, annulus fibrosus, and even the endplate protrude posteriorly, severely compressing the nerves and causing symptoms. Accumulated damage is the main cause of intervertebral disc degeneration. Repeated bending, twisting, and other movements most easily cause intervertebral disc damage, so this disease has a certain relationship with occupation. During pregnancy, the entire ligament system is in a relaxed state, and the lumbosacral region bears more stress than usual, increasing the risk of disc herniation. About 32% of adolescent patients under the age of 20 have a positive family history. Congenital developmental abnormalities of the lumbosacral region, such as lumbarization of the sacrum, sacralization of the lumbar spine, and asymmetry of the articular processes, cause the lower lumbar spine to bear abnormal stress, which increases the damage to the intervertebral disc.

[0004] Lumbar disc herniation is common in patients aged 20-50 years, with a male to female ratio of about (4-6):1. Patients often have a history of bending labor or long-term sitting work, and the first onset often occurs during the process of half-bending heavy lifting or sudden twisting. The following are possible symptoms and characteristics:

[0005] Low back pain: Most patients with lumbar disc herniation have low back pain. Low back pain can occur before leg pain, or at the same time or after leg pain. The cause of low back pain is the stimulation of the dural nerve fibers in the outer annulus fibrosus and posterior longitudinal ligament by the intervertebral disc herniation.

[0006] Sciatica: Since about 95% of lumbar disc herniations occur in the L4-L5 and L5-S1 spaces, most are accompanied by sciatica. Sciatica usually occurs gradually, with pain radiating from the buttocks, posterolateral thigh, lateral leg, to the heel or dorsum of the foot.

[0007] Cauda equina syndrome: Severe lumbar disc herniation can lead to changes in spinal shape, with significant deformity or deformation. Movements in daily activities can exacerbate the deformation, and over time can lead to lumbar lordosis or kyphosis.

[0008] Lumbar disc herniation treatment is generally divided into non-surgical treatment and surgical treatment.

[0009] Non-surgical treatment indications: (1) Patients with initial onset and short course of disease; (2) Symptoms can be relieved after rest; (3) Cannot perform surgery due to systemic disease or local skin disease; (4) Those who do not agree to surgery. It can include conservative treatment, rest and activity restriction and other treatments. Conservative treatment is often used for patients with no obvious nerve function damage, no obvious spinal instability or other serious complications, aiming to reduce pain, improve function and promote patient recovery. Patients may need to strictly limit activity and strictly rest in bed for 3 weeks, and gradually get out of bed with a waist wrap. Other treatments are usually pelvic traction therapy and physiotherapy.

[0010] Surgical treatment indications: (1) Severe lumbar and leg pain, repeated attacks, ineffective after more than six months of non-surgical treatment, and gradually worsening, affecting work and life; (2) Central type protrusion with cauda equina syndrome and sphincter dysfunction, should undergo emergency surgery; (3) Those with obvious nerve involvement. Traditional open surgery includes total laminectomy and nucleus pulposus removal, hemilaminectomy and nucleus pulposus removal, and laminectomy and nucleus pulposus removal. Microsurgical lumbar disc removal uses a microscope to assist surgery and remove the intervertebral disc. Minimally invasive disc removal surgery includes microendoscopic discectomy (MED), percutaneous endoscopic lumbar discectomy (PELD), and unilateral biportal endoscopy (UBE) under intervertebral disc removal. In recent years, percutaneous endoscopic spinal endoscopy technology represented by PELD has developed rapidly, and its application in clinical practice is becoming more and more widespread due to its small damage and rapid recovery. Artificial disc replacement is still controversial, and careful selection is required for this surgery.

[0011] The diagnosis of lumbar disc herniation requires the use of different imaging methods, including X-ray, CT scan and MRI. These methods have their own imaging principles, advantages and disadvantages, and imaging features, which can help understand lumbar disc herniation and its imaging characteristics. In addition, patient signs and physical examination are also used for diagnosis.

[0012] X-ray imaging is a technique that uses X-rays to produce images of the body. X-rays are a type of electromagnetic radiation that can pass through the body and be detected by a special camera. The amount of X-rays absorbed by different tissues in the body varies, which allows for the creation of images. For example, bones absorb more X-rays than soft tissues, resulting in a brighter image on the X-ray film. X-ray imaging is often used as a routine screening tool for various conditions, including lumbar disc herniation. In some cases, the X-ray film may appear completely normal in patients with lumbar disc herniation. However, there are some common imaging findings that may be seen on X-ray films. These include scoliosis of the lumbar spine, reduced or absent lordosis on lateral films, and narrowing of the intervertebral space.

[0013] The advantages of X-ray imaging include its speed (scanning can be completed within a minute), convenience, and low cost. However, X-ray imaging has some limitations. It cannot show soft tissue conditions, assess spinal cord compression, or evaluate the location and size of a herniated disc. DR requires exposure to low doses of radiation, which may not be suitable for certain individuals such as pregnant women.

[0014] X-ray imaging is primarily used in clinical settings to identify signs of degeneration, such as calcification of the annulus fibrosus, osteophytes, hypertrophy of the articular processes, and sclerosis.

[0015] CT scanning is a type of X-ray imaging that uses a rotating X-ray source and detector to produce images. The imaging process works as follows: X-ray source: A rotating X-ray source emits X-rays at different angles to pass through the patient's body. Detector: The detector receives the X-ray signals that have been absorbed and attenuated by the body tissues. Computer reconstruction: The computer uses mathematical algorithms to calculate and reconstruct the data collected to generate cross-sectional (tomographic) images. Multiplanar reconstruction: CT scanning provides information about the severity of lumbar disc herniation on CT images. The herniated disc may appear to have a decreased height, resulting in a narrowed intervertebral space. The water content in the disc may decrease, appearing as a low-density area on the CT image. The disc may take on an oval or curved shape, indicating posterior or lateral disc herniation. If there is a herniated disc tissue, it may appear as a low-density image on the CT scan. Under prolonged pressure, osteophytes (bone spurs) may develop, appearing as jagged protrusions on the edges of the vertebral bodies on the CT image. Bone marrow edema may also occur near the disc or around the vertebral bodies.

[0016] Advantages of CT imaging: CT provides high-resolution three-dimensional images, showing the lumbar spine in more detail, and has a higher sensitivity and specificity for diagnosing lumbar disc herniation. It can better show the details of the spinal bone structure. The manifestations of lumbar disc herniation on CT include posterior edge deformation of the intervertebral disc, compression and deformation of the dural sac, displacement of epidural fat, soft tissue density shadow in the epidural space, and compression and displacement of the nerve root sheath. Stereoscopic anatomy: Compared with DR, CT can perform multi-planar reconstruction, showing the stereoscopic anatomical structure of multiple angles such as transverse, coronal and sagittal planes, which helps to assess the location and size of the protrusion. In addition, CT scanning speed is fast (can be completed within one minute), imaging time is short, and is very useful for rapid diagnosis in emergency situations. Disadvantages: Compared with DR, CT scanning has a relatively high radiation dose, and long-term frequent examination may have some impact on patient health. It cannot show the condition of soft tissue and cannot assess the compression of the spinal cord. High cost: Compared with X-ray examination, CT scanning costs more, which may increase medical expenses

[0017] The main clinical application of CT is: CT scan plays an important role in the diagnosis and treatment of lumbar disc herniation. CT can clearly show the morphological changes of the lumbar intervertebral disc and its relationship with the surrounding tissues (such as nerve roots and spinal cord), thereby helping doctors to confirm whether there is disc herniation and the specific location and degree of the herniation.

[0018] The principle of MRI imaging mainly utilizes the magnetic resonance phenomenon. MRI uses a strong magnetic field and non-loss radio waves to obtain images. Under the action of the magnetic field, atomic nuclei (such as hydrogen nuclei) in the human body will produce resonance signals. The imaging process includes the following steps. Excitation of atomic nucleus resonance: the patient is placed in a strong magnetic field, and the atomic nucleus changes between different energy levels, absorbing and releasing energy. Signal reception: the probe generates radio waves to capture the signals released by the atomic nucleus. Signal processing and image reconstruction: the computer processes the received signals to generate images, showing the detailed structure of the tissue.

[0019] MRI can clearly show the image of human anatomical structure, which is of great help to the diagnosis of lumbar disc herniation. MRI can be used to observe the degeneration of each intervertebral disc comprehensively. Healthy intervertebral discs usually show high signals (T2 weighted images) on MRI, while protruding or degenerative intervertebral discs show low signals, appearing dark, which usually indicates a decrease in water content in the intervertebral disc. It can also understand the degree and location of the nucleus pulposus, clearly show the protrusion of the intervertebral disc, and help to evaluate the type of protrusion (such as central protrusion, lateral protrusion). MRI can assess whether the protruding disc compresses the nerve root or spinal cord. The compressed nerve usually shows signal changes or swelling on the image, showing the changes of the nerve. It can also show inflammation or edema around the protruding disc, which is often accompanied by reactive changes in muscle or membrane. It can also identify other space-occupying lesions in the spinal canal.

[0020] Advantages of MRI imaging: high soft tissue resolution, MRI is very good at displaying soft tissue structures, and can clearly show structures such as paravertebral tissue and spinal cord. Multi-planar imaging: sagittal, coronal and transverse imaging can be performed, providing the possibility of observing the lumbar disc herniation from multiple angles. No radiation: MRI imaging has no radiation, avoiding the impact of X-ray radiation on patients, especially suitable for the examination of special groups such as pregnant women and children. Disadvantages: long examination time: MRI imaging requires a longer examination time (usually more than half an hour), and MRI cannot be performed in emergency situations such as night emergency. Patients need to maintain a relatively fixed posture, which is a great challenge for patients with severe lumbar disc herniation; not suitable for all patients: for some patients, there are metal implants such as cardiac pacemakers, and some special diseases such as claustrophobia, which may not be suitable for MRI examination. MRI is expensive (about 500 for one part), which causes economic burden to the examinee.

[0021] The clinical application of MRI mainly lies in: MRI can not only accurately evaluate the damage of the disc and the compression of the nerve, but also provide detailed images of the surrounding tissues. Therefore, it is very advantageous in repeated evaluation of the patient's condition. It provides an important reference for the development of treatment plans.

[0022] Contrast detection methods such as spinal cord contrast, epidural contrast, and disc contrast can indirectly show whether there is a disc herniation and its degree. Since these methods are invasive and some have complications, and some are technically complex, they are less commonly used in clinical practice. They are only used when the general diagnostic method is not clear.

[0023] The present application provides a lumbar disc herniation image recognition conversion method and system, which converts the CT image of lumbar disc herniation into an MRI image, thereby combining the advantages of CT and MRI, which is beneficial to the rapid diagnosis of lumbar disc herniation by spine surgeons, accurate selection of treatment plans, and long-term prognosis of patients after surgery. SUMMARY

[0024] The purpose of the present application is to provide a lumbar disc herniation image recognition conversion method and system to solve the problems of the prior art.

[0025] To achieve the above-mentioned purpose, the present application provides the following solutions:

[0026] The present application provides a lumbar disc herniation image recognition conversion method, comprising the following steps:

[0027] S1. Collecting CT images and MRI images of lumbar disc herniation of patients;

[0028] S2. Preprocess the collected lumbar disc herniation CT images and lumbar disc herniation MRI images to construct a training set;

[0029] S3. Construct an ITAD-pix model;

[0030] S4. Input the training set into the ITAD-pix model for training;

[0031] S5. Input the lumbar disc herniation CT images not involved in training into the trained ITAD-pix model, and convert the lumbar disc herniation CT images not involved in training into target lumbar disc herniation MRI images through the ITAD-pix model.

[0032] Preferably, in step S2, the preprocessing includes:

[0033] S21. Cut the collected lumbar disc herniation CT images and lumbar disc herniation MRI images;

[0034] S22. Denoise the part of the images and unify the brightness, grayscale and contrast;

[0035] S23. Pair the unified lumbar disc herniation CT images and lumbar disc herniation MRI images to align the structure.

[0036] Preferably, in step S21, the standards for cutting the images include: sagittal image cutting: cutting from the lower edge of the second lumbar vertebrae, the upper edge of the first sacral vertebrae, the front edge of the fourth or fifth lumbar vertebrae, and the middle back of the lumbar spine; transverse image cutting: cutting from the front edge of the lumbar disc, the lamina, and the left and right sides of the lumbar disc.

[0037] Preferably, step S3 includes:

[0038] S31. Construct the model of ITAD-pix based on the residual network structure;

[0039] S32. Generate an AdaIN layer;

[0040] S33. Add the AdaIN layer to each sampling layer in the generator and discriminator to optimize the ITAD-pix model.

[0041] Preferably, step S32 includes: using an encoder-decoder architecture, where the encoder f is fixed to the first few layers of the pre-trained VGG-19, and after encoding the content and style images in the feature space, both feature maps are fed into the AdaIN layer, which aligns the mean and variance of the content feature map with the mean and variance of the style feature map, thereby generating the target feature map t:

[0042] t = AdaIN(f(c), f(s));

[0043] The randomly initialized decoder g is trained to map t back into image space, generating a stylized image T(c, s):

[0044] T(c, s) = g(t).

[0045] Preferably, the step S33 comprises: first inputting the lumbar disc herniation CT image to be detected into the generator, the structure of each layer of the generator is down-sampling of convolution, IN regularization and Leaky ReLU activation function, and a step of image enhancement is performed before down-sampling, a ReflectionPad2d layer is introduced, and the relevant features obtained after the first down-sampling are taken as the input of the next down-sampling, at this time, the data becomes 64x256x256, after the second down-sampling, the data becomes 128x128x128, after the third down-sampling, the data becomes 256x64x64, at this time, the bottom is reached, 9 residual modules are introduced between the down-sampling and the up-sampling to deepen the network and enhance the data, then the up-sampling operation is performed, the structure of the up-sampling is deconvolution, IN normalization, ReLU activation function to restore the size of the image, and also uses the ReflectionPad2d layer to perform data enhancement, the first up-sampling receives the data from the Resnet_block, and the data becomes 256x64x64, at this time, the AdaIN layer is introduced to learn and transform the local context information, the second up-sampling changes the data to 128x128x128, and the third up-sampling changes the data to 64x256x256; the discriminator takes the data from the generator as input, and changes the data to 64x128x128 through the convolution layer and the LeakyReLU, and changes the data from 64x128x128 to 128x64x64, then to 256x32x32, and finally to 512x31x31 through three convolution, IN regularization and Leaky ReLU activation functions, and then changes the data to 1x30x30 through a convolution layer.

[0046] Preferably, iteration corresponding code is added in the last generation step, so that the generated MRI image is repeatedly iteratively generated.

[0047] Preferably, the step S4 adopts an alternating training method for training.

[0048] Preferably, the step S4 comprises:

[0049] S41. The training of the generator, the given input image is generated by the generator to generate the corresponding output image, and the generated image and the target image are sent into the discriminator for training, the training process first generates the corresponding output image by the generator through the input image; then the loss function between the generated image and the target image is calculated, and the error is back propagated to update the parameters of the generator; finally, in the training process, the parameters of the generator are constantly updated in the way of stochastic gradient descent;

[0050] S42. The training of the discriminator, the discriminator discriminates the generated image and the target image to judge whether these images are real, the training process first randomly selects a group of real images and a group of generated images; then the images are input into the discriminator, and the probabilities that the real images and the generated images belong to the real data are calculated; the loss functions corresponding to the real images and the generated images, and the loss function of the discriminator as a whole are calculated, and the parameters of the discriminator are updated by back propagating the error; finally, in the training process, the parameters of the discriminator are constantly updated in the way of stochastic gradient descent.

[0051] The application also provides a lumbar disc herniation image recognition and conversion system, comprising:

[0052] An image acquisition module is configured to acquire lumbar disc herniation CT images and lumbar disc herniation MRI images of a patient;

[0053] An image preprocessing module is configured to preprocess the acquired lumbar disc herniation CT images and lumbar disc herniation MRI images, and construct a training set;

[0054] A model construction module is configured to construct an ITAD-pix model;

[0055] A model training module is configured to input the training set into the ITAD-pix model for training;

[0056] An image recognition and conversion module is configured to input a lumbar disc herniation CT image not participating in the training into the trained ITAD-pix model, and convert the lumbar disc herniation CT image not participating in the training into a target lumbar disc herniation MRI image through the ITAD-pix model.

[0057] The application has the following beneficial technical effects compared with the prior art:

[0058] The application provides a lumbar disc herniation image recognition conversion method and system, which comprises the following steps: firstly, collecting lumbar disc herniation CT images and lumbar disc herniation MRI images of a patient; secondly, preprocessing the collected lumbar disc herniation CT images and lumbar disc herniation MRI images to construct a training set; thirdly, constructing an ITAD-pix model; inputting the training set into the ITAD-pix model for training; and finally, inputting the lumbar disc herniation CT images that do not participate in the training into the trained ITAD-pix model, and converting the lumbar disc herniation CT images that do not participate in the training into target lumbar disc herniation MRI images through the ITAD-pix model; the lumbar disc herniation CT images are converted into MRI images, so as to combine the advantages of CT and MRI, which is beneficial to the rapid diagnosis of lumbar disc herniation by a spine surgeon, the accurate selection of a treatment scheme, and the postoperative recovery and long-term prognosis of a patient. BRIEF DESCRIPTION OF DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings needed in the embodiments. Obviously, the drawings described below are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0060] Figure 1 The image preprocessed in the present application;

[0061] Figure 2 The ITAD-pix model structure diagram in the present application;

[0062] Figure 3 The effect diagram of the recognition and conversion in the present application. DETAILED DESCRIPTION

[0063] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0064] The purpose of the present application is to provide a lumbar disc herniation image recognition conversion method and system to solve the problems in the prior art.

[0065] The present application is a specific application of deep learning in medical images. Deep learning is a machine learning technique that simulates the structure and function of the human brain's neural network, allowing computers to automatically learn and understand data. It uses multi-level neural network models to process complex data, abstract and understand data at a high level, and thus efficiently process and analyze image, voice, text and other data. Common neural network structures in deep learning include convolutional neural networks (CNN), recurrent neural networks (RNN), generative adversarial networks (GAN), long short-term memory networks (LSTM), autoencoders (AE), attention mechanisms (AM), etc.

[0066] Deep learning has the following advantages:

[0067] Automatic feature learning: Deep learning models can automatically learn features from data without manually extracting features, making them more adaptable.

[0068] Processing large-scale data: It can efficiently process large-scale data, and its learning ability improves with the increase of data volume, making it suitable for large amounts of medical image data in the medical field.

[0069] Improved prediction accuracy: In medical image analysis, deep learning models can provide more accurate and precise prediction results.

[0070] Potential versatility: It is suitable for different types of medical image data, such as X-rays, MRIs, CTs, etc., and has a certain degree of versatility.

[0071] Interpretability: Some deep learning models have a certain degree of interpretability, which can help doctors understand the diagnosis process.

[0072] Current applications of deep learning in medical images include:

[0073] Lesion segmentation: In medical images, deep learning can be used to locate and segment lesion areas, such as papers reporting the segmentation of normal or diseased vertebrae in spinal surgery.

[0074] Image reconstruction: Deep learning can be used in medical images to reconstruct three-dimensional images from two-dimensional images, such as papers reporting the reconstruction of lumbar three-dimensional CT from lumbar DR in spinal surgery.

[0075] Image simulation: There are papers reporting the conversion of single or multi-modal imaging images, such as the conversion of carotid and aortic contrast-free CT images to contrast CT images, and the conversion of normal lumbar cross-sectional CT images to MRI images.

[0076] Disease auxiliary diagnosis: deep learning can be used in medical images to diagnose tumors, lung diseases, cardiovascular diseases, etc. For example, mammogram image analysis for breast cancer screening, lung nodule auxiliary detection, and tumor detection and classification for CT or MRI scans.

[0077] The purpose of the present application is to convert lumbar disc herniation CT images into MRI images and perform auxiliary diagnosis. CT and MRI are images of two different imaging principles, and the conversion between them is naturally suitable for applying GAN for learning and training.

[0078] The GAN network is composed of a generator (Generator) and a discriminator (Discriminator). Among them, the generator is responsible for converting a random noise vector into an image or other type of data similar to the training data, while the discriminator is responsible for distinguishing the difference between the image generated by the generator and the real data. The generator and the discriminator gradually optimize their network parameters through repeated games, until they reach a balance - the generator can generate highly realistic data, and the discriminator cannot accurately distinguish whether the generated data is real data. In the training process of GAN, the generator and the discriminator are in opposition to each other, and constantly optimize their performance in the game. Specifically, the generator tries to confuse the discriminator so that it cannot accurately distinguish between generated data and real data, while the discriminator tries to accurately identify generated data and real data as much as possible, so that in the process of backpropagating errors, the generator can adjust its data generation strategy according to the feedback from the discriminator, so as to gradually generate more realistic data. The objective function in the GAN network is shown in the following figure:

[0079]

[0080] where and represent the probability of real data and generated data, and from the above loss function it can be seen that the calculation of the loss function is in D (discriminator). Because the output of D is generally True / Fake judgment, the binary cross-entropy function is used as a whole. The left side contains two parts: minG and maxD. For maxD, for 1-D(G(z)), it is hoped that it can approach 1 as much as possible, that is, the discriminator can accurately distinguish the generated image. For minG, because it wants the generated data to pass the discriminator's judgment, it needs to make D(G(z)) approach 1, that is, the generated data needs to pass the discriminator, so for it should be as small as possible. This embodies the idea of game learning.

[0081] The following is a detailed introduction to the operating principle of GAN:

[0082] Generator:

[0083] The Generator is a neural network that takes random noise (latent space vectors) as input and tries to generate realistic data. It does this by repeatedly learning the characteristics of the data distribution, trying to generate data that can "fool" the Discriminator.

[0084] Discriminator:

[0085] The Discriminator is another neural network, similar to a binary classifier, that takes real data and fake data generated by the Generator and tries to classify them correctly. The goal of the Discriminator is to accurately distinguish between real data and fake data generated by the Generator.

[0086] Adversarial Training:

[0087] The Generator and Discriminator are in a constant adversarial competition. During training, the Generator tries to generate data that is realistic enough to "fool" the Discriminator, while the Discriminator works to learn to accurately distinguish between real and fake data. This adversarial training process iterates until the Generator generates data that is realistic enough to be indistinguishable from real data by the Discriminator.

[0088] The training process of GANs includes:

[0089] 1) Initialization of network parameters:

[0090] The weights of the Generator and Discriminator are usually randomly initialized.

[0091] 2) Adversarial training loop:

[0092] In each training iteration, the Generator and Discriminator are trained alternately. For each training iteration: the Generator receives random noise and generates data samples. The Discriminator receives real data and fake data generated by the Generator, and makes classification judgments respectively. The Discriminator calculates the loss (the accuracy of distinguishing real and fake data), and the Generator calculates the loss (trying to "fool" the Discriminator). Through the loss function, the weights of the Generator and Discriminator are updated to improve the generation ability of the Generator and the discrimination ability of the Discriminator.

[0093] 3) Convergence and generation:

[0094] As training progresses, the Generator gradually learns the characteristics of the data distribution, and the generated data becomes increasingly realistic. When the data generated by the Generator is realistic enough to be accurately distinguished by the Discriminator, the model reaches a state of convergence.

[0095] After convergence, the generator has learned enough distributional features to reach an optimal state. We input the original image into the generator, which will output the target image using the learned features.

[0096] The present application aims to convert lumbar disc herniation CT images into MRI images, which is naturally suitable for GAN training and image generation. However, the learning of GAN is too general, which will learn the features of each vertebral body and paravertebral body widely, which is not conducive to clinical needs, because we need to focus on the location and size of the lumbar disc herniation to assist clinical diagnosis and treatment selection. Therefore, in this application, we innovatively introduce an enhanced GAN model, which introduces an iterative instance normalization generation mechanism to enhance the performance of the lumbar disc herniation.

[0097] Iterative instance normalization generation is a method of optimizing data scale and distribution through multiple iterations. Its purpose is to ensure that the input data remains within a reasonable range (e.g., mean 0, variance 1) in each iteration, thereby accelerating model convergence and improving training stability. By optimizing the distribution of input data through multiple iterations, it becomes more stable during training. Instance normalization uses a novel adaptive instance normalization layer (AdaIN), which matches the mean and variance of the content features with those of the style features, achieving perfect matching in image style transfer.

[0098] The instance normalization residual layer decoder is mainly a mirror of the encoder, with all pooling layers replaced by nearest upsampling to reduce the chessboard effect. We use reflective padding in f and g to avoid boundary artifacts. Another important architectural choice is whether the decoder should use instance layers, batch layers, or no normalization layers.

[0099] In image style transfer, the model can extract features from content and style images through multi-channel fusion, then apply iterative normalization to ensure each feature remains stable in each iteration, resulting in higher quality images. In image restoration, signals from different feature extractors (e.g., color information, texture information, etc.) are fused while iterative normalization is used to maintain stability during image restoration, which is crucial for restoration quality.

[0100] The above is the principle and method based on the present application, in order to make the above-mentioned purpose, characteristics and advantages of the present application more obvious and easy to understand, the following will be combined with the drawings and specific embodiments to further explain the present application in detail.

[0101] Example 1:

[0102] The present application provides a lumbar disc herniation image recognition conversion method, comprising the following steps:

[0103] S1. Collecting CT images and MRI images of lumbar disc herniation of patients;

[0104] S2. Preprocessing the collected CT images and MRI images of lumbar disc herniation to construct a training set;

[0105] In order to exclude the interference of surrounding soft tissues and make the model learn the characteristics of lumbar disc herniation more accurately, the CT and MRI images of lumbar disc herniation are preprocessed, and the preprocessing includes:

[0106] S21. Cutting the collected CT images and MRI images of lumbar disc herniation;

[0107] S22. Denoising part of the images and unifying brightness, grayscale and contrast;

[0108] S23. Pairing the unified CT images and MRI images of lumbar disc herniation to align the structure, Figure 1 The pictures after preprocessing are shown;

[0109] S3. Constructing an ITAD-pix model;

[0110] The ITAD-pix model in this embodiment adopts the Pix2Pix principle, Pix2Pix is a common generative adversarial network for image style conversion, and the present application proposes a new network model ITAD-pix on the basis;

[0111] The objective function of Pix2Pix is based on conditional generative adversarial network (CGAN), and the loss of CGAN is different from that of GAN. It makes certain changes on the basis of GAN, that is, adds the conditional characteristics, and adds an L1 loss to make the images of the source domain and the target domain as close as possible. The objective function of CGAN can be expressed as:

[0112]

[0113] CGAN is similar to GAN, but GAN generates images according to random noise z, while CGAN generates images according to random noise z and input image x; The difference between CGAN and GAN is that, in addition to generating images that can deceive D, CGAN also needs to make the generated images as close as possible to the target domain images y; As for the selection of loss, CGAN selects L1 loss, which is expressed as follows:

[0114]

[0115] L1 loss is selected instead of L2 loss to ensure less blur, L1 loss can better preserve detailed information, L1 loss has higher tolerance to noise and outliers than L2 loss, which makes the generated image clearer and can better preserve the details of the original image, L1 loss can better solve the problem of mode collapse, L2 loss will have the problem of mode collapse in the generated image, that is, the generated image is too smooth and loses some features of the original image, while L1 loss can better avoid this problem, so that the generated image is more diverse, and the final objective function is:

[0116]

[0117] The generator of the Pix2Pix model selects the Unet structure, the Unet structure, as the name implies, its overall network structure looks like a capital English letter U, including a shrinking path and an expanding path, also known as an encoder-decoder structure, the two paths are connected through a jump connection; The generator adopts the "convolution-BatchNorm-ReLu" network module, image translation needs to be from one high-resolution image to another high-resolution image, that is, the input and output are different in surface details, but have exactly the same underlying approximate model structure, so the input and output need to realize rough alignment operation to make the model converge better effect; For many image translation related problems, Pix2Pix proposes a novel method of jump connection, which adds a jump connection between the i-th layer and the n-i-th layer of Unet, where n is the total number of layers, each jump connection only connects all channels of the i-th layer with the channels of the n-i-th layer, which can bypass the information bottleneck.

[0118] First, input an image, which is changed to 64 channels after two convolutions, then pass through a pooling layer without changing the number of channels, only the size of the image is changed to half of the original, then the operation is similar, until it becomes the type of 1x512x32x32 (batchsize=1, channel number 512, picture size 32x32) 1x1024x28x28 is changed again after two convolution layers.

[0119] The discriminator of Pix2Pix adopts PatchGAN, which adopts the same network module as the generator. The difference between it and the discriminator of the general GAN network is that PatchGAN outputs an N*N matrix (N=70 is selected in Pix2Pix), and all elements in the matrix have only two values, True or False. This increases the receptive field of the model to the original image, so that the model generates images more stably. The author indicates that for low-frequency information, selecting L1 loss can achieve good reconstruction; and for high-frequency information, the PatchGAN redesigned by the author is needed, which only judges whether an N*N patch is true or false, so it can better distinguish high-frequency information. Even if N is much smaller than the size of the original image, PatchGAN can still produce good results. It has fewer parameters, runs faster, and can be applied to images of any size. One of the advantages of PatchGAN is that the fixed-size discriminator can be used for images of any size during testing. The generator convolution is applied to a larger image than the trained generator. For example, 256x256 pixel size images are selected during the training process, and finally 512x512 pixel size images are used during testing.

[0120] Step S3 comprises:

[0121] S31. Constructing the model of ITAD-pix based on the residual network structure; the model generator of ITAD-Pix2pix adopts the residual network (ResNet) structure. The residual network proposes the concept of shortcut connection, which greatly improves the depth of the network and effectively solves the problem of gradient disappearance caused by excessive depth; in the traditional convolutional neural network, the output of the network layer is directly transmitted to the next layer after being transformed by the activation function; while in ResNet, the input of each network layer will not only be transmitted to the next layer, but also be directly transmitted to several layers behind through shortcut connection (Shortcut Connection), forming a residual block. The advantage of this is that even if the number of network layers increases, since there is a direct path, the original input information can be more easily propagated to the later layers, avoiding the problem of rapid decay of gradient in deep network and gradient explosion; the residual block in ResNet is composed of two or three convolutional layers, including an identity mapping and a nonlinear activation function (such as ReLU), the identity mapping directly transmits the input to the output, and the nonlinear activation function introduces nonlinear transformation, this design allows the network to learn the identity mapping when needed, i.e. directly transmitting the input to the output, so as to better adapt to different complexity and difficulty tasks, ITAD-pix uses 9 residual blocks (residual blocks), which deepens the network depth and can better capture the features and details of the image;

[0122] S32. Generating AdaIN layer; Adaptive Instance Normalization (AdaIN) is to input content image c and arbitrary style image s, and synthesize the output image that recombines the content of the former and the style of the latter; a simple encoder-decoder architecture is adopted, in which the encoder f is fixed in the first few layers of the pre-trained VGG-19, and after encoding the content and style images in the feature space, both feature maps are fed into the AdaIN layer, which aligns the mean and variance of the content feature map with the mean and variance of the style feature map, thereby generating the target feature map t:

[0123] t = AdaIN(f(c), f(s));

[0124] The randomly initialized decoder g is trained to map t back to the image space to generate the stylized image T(c, s):

[0125] T(c, s) = g(t);

[0126] The decoder is mainly a mirror image of the encoder, all the pooling layers are replaced by nearest up-sampling to reduce the checkerboard effect; reflection padding is used in f and g to avoid boundary artifacts; another important architectural choice is whether the decoder should use instance layers, batch layers, or no normalization layers at all.

[0127] S33. Add AdaIN layers to each sampling layer in the generator and discriminator to optimize the ITAD-pix model; in order to find the best model for optimization, we add AdaIN layers to each sampling layer in the generator and discriminator to improve the conversion performance and imaging effect of the model on different real style images, Figure 2 Specifically shows the structure of ITAD-Pix;

[0128] First, input the CT picture to be detected into the generator, and the structure of each layer of the generator is convolution, IN regularization and Leaky ReLU activation function. Before downsampling, an image enhancement operation is performed, and a ReflectionPad2d layer is introduced to perform symmetry on the edges of the image up and down and left and right, increasing the resolution of the image. After the first downsampling, the relevant features obtained are used as the input for the next downsampling. At this time, the data becomes 64x256x256. After the second downsampling, the data becomes 128x128x128. After the third downsampling, the data becomes 256x64x64. At this time, the bottom is reached. Nine residual modules are introduced between the downsampling and the upsampling to deepen the network while enhancing the data. Then the upsampling operation is performed. The role of upsampling is to enlarge and restore the extracted features. The structure of upsampling is deconvolution, IN normalization, and ReLU activation function to restore the size of the image. ReflectionPad2d layer is also used for data enhancement. The first upsampling receives data from the Resnet_block and changes the data to 256x64x64. At this time, Adaptive Instance Normalization is introduced to learn and transform local context information, which can better generate images that meet the MRI features. The second upsampling changes the data to 128x128x128. After the third upsampling, the data becomes 64x256x256.

[0129] The discriminator takes the data from the generator as input, and through the convolution layer and LeakyReLU, the data becomes 64x128x128. Through three times of convolution, IN regularization and Leaky ReLU activation function, the data is converted from 64x128x128 to 128x64x64, then to 256x32x32, and finally to 512x31x31. Then through a convolution layer, the data becomes 1x30x30.

[0130] In the last generation step, the corresponding code is added to generate the MRI image repeatedly, making up for the shortcomings of the previous generation, so as to generate an MRI image with better effect and performance.

[0131] S4. Input the training set into the ITAD-pix model for training;

[0132] For the two different networks of the discriminator and the generator, an alternating training method is used, that is, the entire training process is divided into two parts: the training of the generator and the training of the discriminator.

[0133] The training of the generator, in the training of the generator, the given input image is generated through the generator to generate the corresponding output image, and the generated image and the target image are sent into the discriminator for training; the training process is first to generate the corresponding output image through the generator; then the loss function between the generated image and the target image is calculated, and the error is back propagated to update the parameters of the generator; finally, in the training process, the parameters of the generator are constantly updated in the way of stochastic gradient descent, so as to gradually improve the performance of the generator;

[0134] The training of the discriminator, in the training of the discriminator, the discriminator needs to distinguish the generated image and the target image to judge whether these images are real; the training process is first to randomly select a group of real images and a group of generated images; then input these images into the discriminator, and calculate the probability that the real image and the generated image belong to the real data (i.e. the output value of the discriminator); by calculating the loss function corresponding to the real image and the generated image, and the loss function of the discriminator as a whole, and updating the parameters of the discriminator by back propagation of error; finally, in the training process, the parameters of the discriminator are constantly updated in the way of stochastic gradient descent, so as to improve its performance;

[0135] In the process of alternating training, the training of the generator and the discriminator is repeated, and their parameters are constantly optimized, until the model converges or reaches the preset number of training times. This alternating training method can effectively balance the performance of the generator and the discriminator in the training process, so that the final model can better complete the image translation task;

[0136] S5. Input the lumbar disc herniation CT image not participating in the training into the trained ITAD-pix model, and the generator will convert the style according to the trained parameters to convert the lumbar disc herniation CT image into an MRI image.

[0137] Figure 3 The effect of the ITAD-pix synthesized picture is shown. The synthesized picture shows good similarity in lumbar disc herniation.

[0138] At the same time, subjective evaluation and objective evaluation are used to evaluate the image quality of the application; objective evaluation uses structural similarity (SSIM), the value range of SSIM is 0-1, and the closer to 1 represents the higher image similarity quality; subjective evaluation invites senior spine surgeons to participate in the evaluation, and defines the lumbar disc herniation part in MRI as the region of interest (ROI).

[0139] In the objective evaluation results, the SSIM of the generated sagittal MRI of ITAD-pix and the real MRI reaches 0.782, the SSIM of the synthesized MRI-ROI of ITAD-pix and the real MRI-ROI reaches 0.878; the SSIM of the generated transverse MRI and the real MRI reaches 0.721, and the SSIM of the synthesized MRI-ROI of ITAD-pix and the real MRI-ROI reaches 0.758; in the subjective evaluation results, the sensitivity of the synthesized MRI in diagnosing lumbar disc herniation reaches 85.28%, and the specificity reaches 95.56%; the accuracy of the treatment scheme selected by the MRI synthesized by ITAD-pix reaches 96.38%.

[0140] The embodiment also provides a lumbar disc herniation image recognition and conversion system, comprising:

[0141] An image acquisition module is configured to acquire lumbar disc herniation CT images and lumbar disc herniation MRI images of a patient;

[0142] An image preprocessing module is configured to preprocess the acquired lumbar disc herniation CT images and lumbar disc herniation MRI images, and construct a training set;

[0143] A model construction module is configured to construct an ITAD-pix model;

[0144] A model training module is configured to input the training set into the ITAD-pix model for training;

[0145] An image recognition and conversion module is configured to input a lumbar disc herniation CT image not participating in the training into the trained ITAD-pix model, and convert the lumbar disc herniation CT image not participating in the training into a target lumbar disc herniation MRI image through the ITAD-pix model.

[0146] The lumbar disc herniation image recognition and conversion method and system provided by the application have the following positive significance:

[0147] MRI, as the gold standard for diagnosing lumbar disc herniation, has various inconveniences: high cost, long time consumption, being in a closed space, and the patient needs to remain still. CT is fast, convenient and efficient, but it cannot provide soft tissue-related information to help spine surgeons observe the size and location of the disc herniation. Therefore, the ITAD-pix model is developed to combine the advantages of CT and MRI examination for clinical use.

[0148] MRI can provide high-resolution images of the lumbar region, including bone, soft tissue, nerves and blood vessels. This provides a clear view of the anatomical structure for doctors; effectively identify changes in the shape of the intervertebral disc, such as disc herniation, disc extrusion, bulging, etc. These changes are manifested as high signal and shape changes in the intervertebral disc in MRI images. Through MRI, doctors can see if the disc herniation is compressing the surrounding spinal nerve roots or spinal cord. This is crucial for assessing the cause of symptoms such as pain, numbness, etc. It can help rule out other conditions that may cause similar symptoms, such as spinal stenosis, tumors, fractures, etc. It shows the location and extent of disc herniation, helping doctors develop more effective treatment plans. It provides important imaging evidence for clinical practice, which can help doctors accurately assess the condition and develop appropriate treatment plans.

[0149] The present application greatly improves the diagnostic efficiency of lumbar disc herniation, which is of great benefit to the imaging department and spine surgery. Not only does it improve the MRI imaging efficiency of the imaging department, but it also assists spine surgeons in diagnosing and treating lumbar disc herniation. It can assist spine surgeons in making treatment decisions and selecting the most suitable treatment plan for different patients.

[0150] The present application applies specific examples to illustrate the principles and implementation methods of the present application. The above examples are only used to help understand the method and core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation method and application range will be changed. In summary, the content of the specification should not be understood as a limitation of the present application.

Claims

1. A method for image recognition and conversion of lumbar disc herniation, characterized in that: Includes the following steps: S1. Acquire CT and MRI images of the patient's lumbar disc herniation; S2. Preprocess the acquired CT and MRI images of lumbar disc herniation to construct a training set; S3. Construct the ITAD-pix model; S4. Input the training set into the ITAD-pix model for training; S5. Input the untrained CT images of lumbar disc herniation into the trained ITAD-pix model, and use the ITAD-pix model to convert the untrained CT images of lumbar disc herniation into the target MRI images of lumbar disc herniation. Step S3 includes: S31. Construct an ITAD-pix model based on the residual network structure; S32. Generate the AdaIN layer; S33. Add the AdaIN layer to each sampling layer in the generator and discriminator to optimize the ITAD-pix model; Step S32 includes: employing an encoder-decoder architecture, where the encoder f is fixed in the first few layers of the pre-trained VGG-19, and after encoding the content and style images in the feature space, both feature maps are fed into the AdaIN layer, which aligns the mean and variance of the content feature map with the mean and variance of the style feature map to produce the target feature map t. t = AdaIN(f(c), f(s)); A randomly initialized decoder g is trained to map t back to the image space, generating a stylized image T(c, s): T(c, s) = g(t); Step S33 includes: First, the CT image of the lumbar disc herniation to be detected is input into the generator. The downsampling structure of each layer of the generator is convolution, IN regularization, and Leaky ReLU activation function. Before downsampling, an image enhancement operation is performed by introducing a ReflectionPad2d layer. After the first downsampling, the relevant features are obtained and used as the input for the next downsampling. At this time, the data becomes 64×256×256. After the second downsampling, the data becomes 128×128×128. After the third downsampling, the data becomes 256×64×64. At this point, the bottom layer is reached. Nine residual modules are introduced between downsampling and upsampling to deepen the network and enhance the data. Then, an upsampling operation is performed. The upsampling structure is deconvolution, IN normalization, and ReLU activation function. The activation function restores the image size and also uses the ReflectionPad2d layer for data augmentation. The first upsampling receives data from the Resnet_block and transforms the data into 256×64×64. At this time, the AdaIN layer is introduced to learn and transform local context information. The second upsampling transforms the data into 128×128×128, and the third upsampling transforms the data into 64×256×256. The discriminator takes the data from the generator as input and passes it through a convolutional layer and LeakyReLU to transform the data into 64×128×128. Through three convolutions, IN regularization, and the Leaky ReLU activation function, the data is transformed from 64×128×128 into 128×64×64, then into 256×32×32, and finally into 512×31×31. Then, through a convolutional layer, the data is transformed into 1×30×30. In the final generation step, iterative code was added to generate the MRI images iteratively.

2. The image recognition and conversion method for lumbar disc herniation according to claim 1, characterized in that: In step S2, the preprocessing includes: S21. Cropping the acquired CT and MRI images of lumbar disc herniation; S22. Perform noise reduction on some parts of the image and unify the brightness, grayscale, and contrast; S23. Pair the unified CT and MRI images of lumbar disc herniation to align the structures.

3. The image recognition and conversion method for lumbar disc herniation according to claim 2, characterized in that: In step S21, the standards for cropping the image include: Sagittal image cropping: cropping is performed with the middle and lower edge of the second lumbar vertebra above, the upper edge of the first sacral vertebra below, the anterior edge of the fourth or fifth lumbar vertebra in front, and the middle and posterior edge of the spinous process of the lumbar vertebra in the back; Transverse image cropping: cropping is performed with the anterior edge of the lumbar intervertebral disc above, the lamina behind, and the left and right edges of the lumbar intervertebral disc respectively.

4. The image recognition and conversion method for lumbar disc herniation according to claim 1, characterized in that: Step S4 uses an alternating training method for training.

5. The image recognition and conversion method for lumbar disc herniation according to claim 1, characterized in that: Step S4 includes: S41. Generator training: The given input image is used to generate the corresponding output image through the generator, and the generated image and the target image are fed into the discriminator for training. The training process first generates the corresponding output image through the generator from the input image; then the loss function between the generated image and the target image is calculated, and the error is backpropagated to update the generator parameters; finally, during the training process, stochastic gradient descent is used to continuously update the generator parameters. S42. Training the discriminator: The discriminator distinguishes between generated and target images to determine whether they are real. The training process first randomly selects a set of real images and a set of generated images; then, these images are input into the discriminator, and the probabilities of real and generated images belonging to real data are calculated; the loss functions corresponding to real and generated images, as well as the overall loss function of the discriminator, are calculated, and the discriminator parameters are updated by backpropagating the error; finally, during the training process, stochastic gradient descent is used to continuously update the discriminator parameters.

6. A lumbar disc herniation image recognition and conversion system employing the lumbar disc herniation image recognition and conversion method according to any one of claims 1-5, characterized in that: include: The image acquisition module is used to acquire CT images and MRI images of lumbar disc herniation in patients. The image preprocessing module is used to preprocess the acquired CT and MRI images of lumbar disc herniation to construct a training set. The model building module is used to build ITAD-pix models; The model training module is used to input the training set into the ITAD-pix model for training; The image recognition and conversion module is used to input untrained lumbar disc herniation CT images into the trained ITAD-pix model, and then use the ITAD-pix model to convert the untrained lumbar disc herniation CT images into target lumbar disc herniation MRI images.

Citation Information

Patent Citations

  • Lumbar vertebra MRI image detection algorithm based on optimized Yolov5

    CN117011254A

  • Method for converting CT image of lumbar fracture into MR image

    CN118297859A