Method and device for generating back x-ray image, storage medium and computer equipment
The generation of back x-ray images through deep learning models solves the radiation exposure and psychological stress caused by the frequent use of traditional X-ray examinations, and achieves efficient and safe monitoring methods and cost reduction.
Patent Information
- Application Number
- CN202510128079.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-05
- Publication Date
- 2025-05-13
AI Technical Summary
When used frequently, traditional X-rays will increase the patient's risk of radiation exposure and psychological stress, and are costly and are not suitable for frequent use.
Reduce dependence on traditional X-ray photography by using deep learning models, especially the CycleGAN architecture. The model uses training data sets, including patient back photos and x-ray images, to generate unpaired data sets, reduce the difficulty of obtaining data sets, and improve the generalization ability of the model.
It significantly reduces the risk of radiation exposure in patients, provides an efficient and safe monitoring method, forms realistic back x-ray images, and reduces medical costs and data acquisition difficulty.
Smart Images

Figure CN119992256A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image processing, and in particular to a method, device, storage medium and computer equipment for generating back X-ray images. Background Art
[0002] Scoliosis is a common spinal deformity, especially among adolescents. The incidence of scoliosis is as high as 20%, and the male to female ratio is about 1:3. Its treatment and rehabilitation require long-term follow-up observation.
[0003] Traditionally, doctors would arrange for patients to undergo X-ray examinations regularly to assess the progression of the disease, but this not only increases the risk of radiation exposure for patients, but may also cause unnecessary psychological stress for patients. In addition, due to the high cost of X-rays and the risk of radiation, they are not suitable for frequent use.
[0004] Therefore, how to understand the scoliosis condition of a patient without causing negative problems to the patient has become a technical problem that needs to be urgently solved by technical personnel in this field. Summary of the invention
[0005] The present invention provides a method, device, storage medium and computer equipment for generating back X-ray images, which are mainly capable of understanding the scoliosis condition of a patient without causing negative problems to the patient.
[0006] According to a first aspect of the present invention, there is provided a method for generating a back x-ray image, comprising:
[0007] Acquire a back photograph of the first patient taken by a photographing device;
[0008] The first patient's back photograph is input into a preset deep learning model to output a first back x-ray image; wherein the preset deep learning model is trained based on a data set, and the data set includes a second patient's back photograph and a second back x-ray image, wherein the second back x-ray image includes a third back x-ray image paired with the second patient's back photograph and / or a fourth back x-ray image unpaired with the second patient's back photograph.
[0009] In some embodiments, the training step of the preset deep learning model includes:
[0010] Construct a target deep learning model of a cyclegan architecture, the target deep learning model comprising a first generator, a second generator, a first discriminator, a second discriminator and a loss function; wherein the first generator is used to convert the second patient's back photo into a fifth back X-ray image, and the first discriminator is used to predict the probability of whether the fifth back X-ray image converted by the first generator is a real back X-ray image; the second generator is used to convert the second back X-ray image into a third patient's back photo, and the second discriminator is used to predict the probability of whether the third patient's back photo converted by the second generator is a real patient's back photo; the loss function comprises an adversarial loss and a cycle consistency loss; the adversarial loss comprises the adversarial loss of the first generator and the first discriminator, and also comprises the adversarial loss of the second generator and the second discriminator; the cycle consistency loss comprises the cycle consistency loss of the first generator and the second generator;
[0011] Using the data set, a target deep learning model is trained to determine a preset deep learning model.
[0012] In some embodiments, the step of using the data set to train the target deep learning model to determine the preset deep learning model includes:
[0013] The target deep learning model is trained using the data set, and during the training process, the overall optimization target is made to meet preset requirements; wherein, the overall optimization target is determined based on the adversarial loss function and the cycle consistency loss function, and the preset requirements include minimizing the overall optimization target so that the first generator and the second generator generate realistic images, and also include maximizing the overall optimization target so that the first discriminator and the second discriminator maximize the ability to distinguish between real images and generated images.
[0014] In some embodiments, the formula for the overall optimization objective is:
[0015] L(G,F,D X ,D Y )=λ gan L GAN (G,D Y ,X,Y)+λ gan L GAN (F,D X ,Y,X)+
[0016] λ cycle L cycle (G,F);
[0017] Among them, L GAN (G,D Y ,X,Y) is the adversarial loss between the first generator and the first discriminator; L GAN (F,DX ,Y,X) is the adversarial loss of the second generator and the second discriminator; λ gan and λ cycle are weight coefficients; L cycle (G, F) is the cycle consistency loss.
[0018] In some embodiments, before determining the preset deep learning model, the method further includes:
[0019] Evaluate the trained target deep learning model;
[0020] If the evaluation result reaches the preset result, the trained target deep learning model is determined to be the preset deep learning model.
[0021] In some embodiments, the step of evaluating the trained target deep learning model includes:
[0022] Determining a structural similarity index, a peak signal-to-noise ratio, and a Cobb angle error using the trained target deep learning model;
[0023] If the structural similarity index is greater than a third preset value, the peak signal-to-noise ratio is greater than a fourth preset value, and the Cobb angle error is less than a fifth preset value, it is determined that the evaluation result reaches a preset result.
[0024] In some embodiments, before using the data set to train a target deep learning model to determine a preset deep learning model, the method further includes:
[0025] The second patient's back photograph and the second back X-ray image in the data set are preprocessed respectively, wherein the preprocessing includes clipping, normalization and / or filtering technology processing.
[0026] According to a second aspect of the present invention, there is provided an apparatus for generating a back x-ray image, comprising:
[0027] An acquisition unit, used for acquiring a back photo of the first patient taken by a photographing device;
[0028] An output unit is used to input the first patient's back photograph into a preset deep learning model to output a first back X-ray image; wherein the preset deep learning model is trained based on a data set, and the data set includes a second patient's back photograph and a second back X-ray image, wherein the second back X-ray image includes a third back X-ray image paired with the second patient's back photograph and / or a fourth back X-ray image not paired with the second patient's back photograph.
[0029] According to a third aspect of the present invention, there is provided a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a method for generating a back X-ray image.
[0030] According to a fourth aspect of the present invention, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above method for generating a back X-ray image when executing the program.
[0031] According to a method, device, storage medium and computer equipment for generating back X-ray images provided by the present invention, compared with the traditional reliance on frequent X-ray shooting, the present invention can significantly reduce the risk of radiation exposure to patients, while providing an efficient and safe monitoring method to form realistic back X-ray images. In addition, the method can use an unpaired data set, which reduces the problem of difficulty in obtaining data sets for training models and improves the generalization ability of the model. The method includes: obtaining a first patient back photo taken by a shooting device; inputting the first patient back photo into a preset deep learning model to output a first back X-ray image; wherein the preset deep learning model is obtained by training according to a data set, and the data set includes a second patient back photo and a second back X-ray image, wherein the second back X-ray image includes a third back X-ray image paired with the second patient back photo and / or a fourth back X-ray image unpaired with the second patient back photo. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0033] Figure 1 A flowchart of a method for generating a back X-ray image according to some embodiments is exemplarily shown;
[0034] Figure 2 exemplarily showing two sets of first patient back photographs and first back X-ray images provided according to some embodiments;
[0035] Figure 3 A schematic diagram of the structure of a device for generating a back X-ray image provided by an embodiment of the present invention is shown;
[0036] Figure 4 A schematic diagram of the physical structure of a computer device provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0037] The present invention will be described in detail below with reference to the accompanying drawings and in combination with embodiments. It should be noted that the embodiments and features in the embodiments of the present application can be combined with each other without conflict.
[0038] Scoliosis is a common spinal deformity, especially in adolescents. The incidence of scoliosis is as high as 20%, and the male-female ratio is about 1:3. Its treatment and rehabilitation require long-term follow-up observation. Traditionally, doctors regularly arrange X-ray examinations for patients to assess the progression of the disease, but this not only increases the risk of radiation exposure for patients, but may also bring unnecessary psychological pressure to patients. In addition, due to the high cost of X-ray shooting and the risk of radiation, it is not suitable for frequent use. Therefore, how to understand the patient's scoliosis without causing negative problems to the patient has become a technical problem that needs to be solved urgently by technicians in this field.
[0039] In order to solve the above technical problems, the embodiment of the present application provides a method for generating back X-ray images. Compared with the traditional method of relying on frequent X-ray shooting, the present invention can significantly reduce the risk of radiation exposure to patients, and at the same time provide an efficient and safe monitoring method to form a realistic back X-ray image. In addition, the method can use unpaired data sets, which reduces the difficulty of obtaining data sets for training models and improves the generalization ability of the model.
[0040] Figure 1 A flowchart of a method for generating a back X-ray image according to some embodiments is exemplarily shown, including S100-S200.
[0041] S100: Acquire a first patient's back photograph taken by a photographing device.
[0042] In the embodiment of the present application, the first patient's back photograph can be a back photograph of the patient when the back is not blocked, and the spine can be placed in the middle of the lens so that the spine image is located in the middle of the first patient's back photograph taken.
[0043] In the embodiment of the present application, the photographing device may be a camera, or may be a device such as a mobile phone with a photographing function.
[0044] S200. Input the first patient's back photograph into a preset deep learning model to output a first back x-ray image; wherein the preset deep learning model is trained based on a data set, and the data set includes a second patient's back photograph and a second back x-ray image, wherein the second back x-ray image includes a third back x-ray image paired with the second patient's back photograph and / or a fourth back x-ray image unpaired with the second patient's back photograph.
[0045] In the embodiment of the present application, the third back X-ray image paired with the second patient's back photo refers to a back X-ray image obtained from the back of the patient to whom the second patient's back photo belongs, taken by X-ray. The fourth back X-ray image not paired with the second patient's back photo refers to a back X-ray image obtained from the back of the patient to whom the second patient's back photo belongs, which is not taken by X-ray. The preset deep learning model in the embodiment of the present application is not limited to using only paired second patient's back photos and third back X-ray images, i.e., paired data sets, during training, and can also use unpaired second patient's back photos and fourth back X-ray images, i.e., unpaired data sets. The embodiment of the present application uses an unpaired data set to train the model, which can reduce the difficulty of obtaining the data set and improve the generalization ability of the model. Scoliosis requires relatively long rehabilitation training. The present invention can use the generated back X-ray images to assist in high-frequency scoliosis condition assessment, which can significantly reduce the patient's radiation exposure risk, while providing an efficient and safe monitoring method to form a realistic back X-ray image.
[0046] In the embodiment of the present application, a large number of unlabeled patient back photos (such as photos taken by a mobile phone) and back X-ray images (such as X-ray images of the patient's back taken by an X-ray machine) can be collected as the second patient back photos and second back X-ray images of the training set. The second patient back photos and the second back X-ray images do not need to be strictly matched one by one, that is, it is not required that each patient back photo has an exactly matched back X-ray image, which greatly simplifies the process of data set collection and reduces the difficulty and cost of data set collection, expands the scope of available data, and increases the diversity and representativeness of the data set.
[0047] Since direct use of X-ray examination may cause radiation problems to the patient, in an embodiment of the present application, a first patient's back photo is input into a preset deep learning model, and the preset deep learning model outputs a corresponding first back X-ray image. In this way, the patient does not need to use X-ray shooting directly to determine the first back X-ray image, and then the patient's scoliosis condition can be understood through the first back X-ray image.
[0048] In an embodiment of the present application, before inputting the first patient's back photo into a preset deep learning model to output a first back X-ray image, it also includes using a training set to train a preset deep learning model. In some embodiments, the training steps of the preset deep learning model include S201-S202.
[0049] S201, constructing a target deep learning model of a CycleGAN architecture. The target deep learning model can learn the mapping relationship between the second patient's back photo and the fourth patient's back X-ray image without a direct paired data set.
[0050] The target deep learning model includes two generators (i.e., the first generator and the second generator) and two discriminators (i.e., the first discriminator and the second discriminator). The goal is to achieve bidirectional unsupervised image conversion between two domains (patient back photos and back X-ray images).
[0051] Specifically, the target deep learning model includes a first generator, a second generator, a first discriminator, a second discriminator and a loss function; wherein the first generator is used to convert the second patient's back photo into a fifth back X-ray image, and the first discriminator is used to predict the probability of whether the fifth back X-ray image converted by the first generator is a real back X-ray image; the second generator is used to convert the second back X-ray image into a third patient's back photo, and the second discriminator is used to predict the probability of whether the third patient's back photo converted by the second generator is a real patient's back photo; the loss function includes adversarial loss and cycle consistency loss; the adversarial loss includes the adversarial loss of the first generator and the first discriminator, and also includes the adversarial loss of the second generator and the second discriminator; the cycle consistency loss includes the cycle consistency loss of the first generator and the second generator.
[0052] In the embodiment of the present application, the real patient's back photo is the second patient's back photo in the data set, and the real back X-ray image is the second back X-ray image in the data set.
[0053] The specific formulas of adversarial loss and cycle consistency loss are as follows:
[0054] Generator G and Discriminator D Y The adversarial loss of the first generator and the first discriminator is as follows:
[0055]
[0056] Among them, L GAN (G,D Y ,X,Y) is the adversarial loss of the first generator and the first discriminator, G(x): the generator G converts the second patient's back photo x into the fifth back X-ray image (fake sample); D Y (G(x)): Discriminator D Y (G(x)) predicts the probability that the fifth back x-ray image is a false sample; x~p data (x): Sample in the target domain (second patient’s back photo); D Y (y): Discriminator D Y (y) predict the probability that the second back X-ray image y is the real back X-ray image (real sample); y~p data (y): Samples in the target domain (real back X-ray images). The goal of the generator G is to minimize L GAN (G,DY ,X,Y), to generate realistic back X-ray images and deceive the discriminator D Y .
[0057] Generator F and Discriminator D X The adversarial loss of the second generator and the second discriminator is as follows:
[0058]
[0059] Where, F(t): Generator F converts the second back X-ray image y into the generated third patient back photo (fake sample); y~p data (y): Sample in the target domain (real back X-ray image); D X (F(y)): Discriminator D X Predict the probability that the third patient's back photo is a false sample; D X (x): Discriminator D X Predict the probability that the second patient's back photo x is a true sample; x~p data (x): Samples in the target domain (second patient back photo). The goal of the generator F is to minimize L GAN (F,D i ,Y,X), to generate realistic patient back photos and deceive the discriminator D X .
[0060] The cycle consistency loss constrains the generators G and F to ensure that the generated samples can be restored to the original input and maintain the information integrity during the image conversion process, so that the generated back X-ray images are not only realistic, but also matched with the input photos. The cycle consistency loss formula is as follows:
[0061]
[0062] Among them, F(G(x)): input x restored after two transformations through generators G and F; G(F(y)): input y restored after two transformations through generators F and G; ∥·∥1: L1 norm, used to measure the pixel difference between the generated sample and the real sample. The goal is to make F(G(x))≈x and G(F(y))≈y, that is, to maintain consistency in bidirectional transformations.
[0063] S202: Using the data set, train a target deep learning model to determine a preset deep learning model.
[0064] In an embodiment of the present application, the target deep learning model of the CycleGAN architecture constructed in the data set training step S201 is used.
[0065] In order to ensure the consistency and compatibility of the data set used to train the target deep learning model, in some embodiments, before using the data set to train the target deep learning model to determine the preset deep learning model, the method further includes:
[0066] The second patient's back photograph and the second back X-ray image in the data set are preprocessed respectively, wherein the preprocessing includes clipping, normalization and / or filtering technology processing.
[0067] In some embodiments, the step of preprocessing the second patient's back photo and the second back X-ray image in the data set includes: cropping the second patient's back photo and the second back X-ray image in the data set into a uniform size; illustratively, they can be cropped into 256*256 pixels.
[0068] The cropped second patient's back photograph and the second back X-ray image are normalized respectively.
[0069] In this embodiment, the pixel value range obtained after normalization is between [0, 1] to eliminate the brightness and contrast differences between different images.
[0070] The normalized second patient's back photograph and the second back X-ray image are processed by filtering technology respectively.
[0071] In this embodiment, filtering technology is used to remove noise in the image to ensure the stability of model training and data quality. The image processed by filtering technology can remove noise.
[0072] The pre-processed second patient back photo and the second back X-ray image are stored in two folders respectively, wherein one folder is used to store the pre-processed second patient back photo and the other folder is used to store the pre-processed second back X-ray image.
[0073] In the embodiment of the present application, the preprocessing of the second patient's back photograph and the second back X-ray image in the data set is not limited to the above-mentioned cropping, normalization and filtering technology processing, and other technologies that can further ensure the consistency and compatibility of the second patient's back photograph and the second back X-ray image may also be included.
[0074] In some embodiments, the step of using the data set to train the target deep learning model to determine the preset deep learning model includes:
[0075] The target deep learning model is trained using the data set, and during the training process, the overall optimization target is made to meet preset requirements; wherein, the overall optimization target is determined based on the adversarial loss function and the cycle consistency loss function, and the preset requirements include minimizing the overall optimization target so that the first generator and the second generator generate realistic images, and also include maximizing the overall optimization target so that the first discriminator and the second discriminator maximize the ability to distinguish between real images and generated images.
[0076] It should be noted that the first generator generates a realistic image, which refers to a realistic fifth back X-ray image, and the second generator generates a realistic image, which refers to a realistic third patient back photo. The real image is the image in the data set, and the generated image is the image generated by the first generator and the second generator.
[0077] Specifically, during the training process, the overall optimization goal is to make the generator G and F deceive the discriminator D as much as possible. X and D Y , while minimizing the cycle consistency loss to ensure that information is not lost during the conversion process. Specifically, the goal of the generators G and F is to minimize the following loss function:
[0078]
[0079] The generator generates realistic images (reducing adversarial loss) while maintaining cycle consistency (reducing cycle consistency loss) by minimizing the loss function L.
[0080] The discriminator D X and D Y The goal is to maximize the ability to distinguish between real data and generated data:
[0081]
[0082] The discriminator improves its ability to distinguish between real samples and generated samples by maximizing the adversarial loss, thereby forcing the generator to generate more realistic images.
[0083] In some embodiments, the formula for the overall optimization objective is:
[0084] L(G,F,D X ,D Y )=λ gan L GAN (G,D Y ,X,Y)+λ gan L GAN (F,D X ,Y,X)+
[0085] λ cycle Lcycle (G,F);
[0086] Among them, L GAN (G,D Y ,X,Y) is the adversarial loss between the first generator and the first discriminator; L GAN (F,D X ,Y,X) is the adversarial loss of the second generator and the second discriminator; λ gan and λ cycle are weight coefficients; L cycle (G, F) is the cycle consistency loss.
[0087] The process of the overall optimization goal constantly approaching the preset requirements adopts the Adam optimizer adaptive learning rate strategy. The process of the adaptive learning rate strategy includes:
[0088] 1. Gradient calculation:
[0089]
[0090] is the gradient at the current time t, representing the loss function f for the parameter θ t The gradient of (θ).
[0091] 2. First-order moment estimation (momentum, mean):
[0092] m t =β1m t-1 +(1-β1)g t
[0093] m t is the moving average of the gradient, similar to momentum, and β1∈[0,1) is the momentum decay coefficient.
[0094] 3. Second-order moment estimation (mean of squared gradient):
[0095]
[0096] v t is the moving average of the squared gradient, and β2∈[0,1) is the mean decay coefficient.
[0097] 4. Bias correction: Adam will correct m in the initial stage t and v t Perform bias correction to avoid the problem of too small momentum value in the initial stage.
[0098] Modified first moment estimate:
[0099] Modified second moment estimate:
[0100] 5. Update parameters:
[0101]
[0102] Among them, θ is the model parameter, corresponding to the parameters of the generator and discriminator; η is the learning rate; ∈ is a small constant to prevent division by zero errors.
[0103] In one example, the initial learning rate can be set to η = 1 × 10 -5 , β1=0.5, β2=0.999.
[0104] In the embodiment of the present application, through the adversarial training mechanism, the generator G learns to generate realistic back X-ray images from the patient's back photos, so that the discriminator D Y It is impossible to distinguish between the generated back X-ray images and the real back X-ray images; the generator F learns to restore the generated back X-ray images to the patient's back photos, so that the discriminator D X It is impossible to distinguish the generated patient back photos from the real patient back photos. At the same time, the cycle consistency loss constraint is used to ensure that the integrity of the input information is maintained as much as possible during the bidirectional conversion process of patient back photo X→back X-ray image Y→patient back photo X.
[0105] In some embodiments, before determining the preset deep learning model, S203-S204 is also included.
[0106] S203, evaluating the trained target deep learning model;
[0107] In some embodiments, after sufficient training, the model performance is evaluated by a series of quantitative and qualitative evaluation indicators. Specifically, the similarity between the generated back X-ray image and the real back X-ray image is evaluated by using indicators such as structural similarity index (SSIM) and peak signal-to-noise ratio (PSNR). Specifically, the step of evaluating the trained target deep learning model includes:
[0108] Determining a structural similarity index, a peak signal-to-noise ratio, and a Cobb angle error using the trained target deep learning model;
[0109] Among them, the structural similarity index (SSIM) measures the similarity of image structures, and its calculation formula is:
[0110]
[0111] Where: μ x and μ y are the means of the generated back X-ray image x (generated image) and the real back X-ray image y (real image); and is the variance; σ xyis the covariance; C1 and C2 are constants used for stability.
[0112] Peak signal-to-noise ratio (PSNR) reflects the clarity and noise level of an image, and its calculation formula is:
[0113]
[0114] Where MAX is the maximum possible pixel value of the image and MSE (mean square error) is defined as:
[0115]
[0116] I and K represent the generated image and the real image respectively, and m and n are the width and height of the image.
[0117] In addition, clinical experts were invited to conduct qualitative evaluations of the generated back X-ray images, including image clarity, accuracy of spinal structure, and measurement accuracy of the Cobb angle, to ensure the effectiveness of the generated images in clinical applications. The measurement accuracy of the Cobb angle can be evaluated using the following error formula:
[0118] Error = |θ 生成 -θ 真实 |
[0119] Specifically, the Cobb angle is calculated as follows:
[0120] Cobb angle = |θ1-θ2|
[0121] Wherein, θ1 and θ2 are the inclination angles of the diseased vertebrae on both sides of the spine, respectively. Through this formula, the severity of scoliosis can be quantified and accurately classified and labeled accordingly. In some embodiments, the labeling is performed according to the size of the Cobb angle, which is mainly divided into the following four categories: less than 10 degrees, 10 to 20 degrees, 21 to 45 degrees, and greater than 45 degrees.
[0122] If the structural similarity index is greater than a third preset value, the peak signal-to-noise ratio is greater than a fourth preset value, and the Cobb angle error is less than a fifth preset value, it is determined that the evaluation result reaches a preset result.
[0123] In the embodiment of the present application, if the structural similarity index is greater than the third preset value, it means that the generated image is very similar to the real image. If the peak signal-to-noise ratio is greater than the fourth preset value, it means that the clarity of the generated image is high and the noise level is low. If the Cobb angle error is less than the fifth preset value, it further means that the generated image is very similar to the real image.
[0124] S204: If the evaluation result reaches the preset result, determine that the trained target deep learning model is the preset deep learning model.
[0125] If the evaluation results do not meet the preset results, continue to fine-tune the model parameters, including learning rate, loss function weight, and network structure, to achieve the best generation effect and clinical application accuracy.
[0126] Deploy the trained preset deep learning model to the actual environment. Users only need to upload an ordinary back photo to quickly obtain a simulated back X-ray image. The generated back X-ray image is displayed in high resolution. The generated effect is as follows Figure 2 As shown, Figure 2 exemplarily showing two sets of first patient back photos and first back X-ray images provided according to some embodiments, Figure 2 The two images on the left in are a set of first patient back photos and a first back X-ray image generated based on the first patient back photos. Figure 2 The two images on the right side of are another set of first patient back photos and a first back X-ray image generated based on the first patient back photos.
[0127] The embodiments of the present application also provide automatic measurement results of the Cobb angle, which is convenient for doctors to quickly assess the degree of scoliosis and provide reference for doctors to analyze the condition. Set a period to collect feedback, analyze data, adjust model parameters, optimize training data, and introduce technical improvements. Through this mechanism, the model is continuously optimized to ensure the accuracy of its clinical application. And share the improved results with doctors in a timely manner. This method not only reduces the patient's radiation exposure, but also improves data acquisition efficiency, enhances privacy protection, and reduces medical costs.
[0128] The method in the embodiment of the present application can reduce radiation exposure: avoid radiation damage caused by frequent X-ray shooting, and ensure the health of patients; improve data acquisition efficiency: the use of unpaired data sets greatly reduces the difficulty of data collection, expands the scope of available data, and allows more patients to participate in research and treatment; enhance privacy protection: ordinary photos (photos of the patient's back) are easier to obtain and less likely to involve sensitive medical information than X-ray images (back X-ray images), which helps to protect patient privacy and reduce the psychological burden of patients; reduce medical costs: reduce dependence on expensive X-ray equipment, indirectly reduce the cost of medical services, and allow more resources to be invested in other key areas; promote the development of telemedicine: the technical solution of the present invention supports remote uploading of photos, and doctors can make preliminary diagnoses without the patient having to go to the hospital in person, which greatly facilitates patients in remote areas or with limited mobility, and promotes the popularization and quality improvement of telemedicine services; promote medical data standardization: provides a set of standardized data processing processes for the diagnosis and monitoring of scoliosis, is conducive to data sharing and comparison between different medical institutions, promotes the standardization of medical data, and facilitates cross-institutional clinical research and cooperation.
[0129] The following describes the actual application process of the above method. In order to implement the back X-ray image generation method, you first need to prepare the experimental environment. In terms of software, Python 3.8 is used as the programming language, relying on the PyTorch 1.10 deep learning framework and related tool libraries such as OpenCV, Pillow, NumPy and Pandas for data processing. In addition, WandB (Weights & Biases) will be used as an experimental tracking platform to record training progress and evaluation results.
[0130] During the data preparation phase, a large number of unlabeled patient back photos and back X-ray images are collected from public databases and cooperative medical institutions. These data do not need to correspond one to one, that is, ordinary photos and X-ray images can come from different patients. All images will be cropped to only include the back area and adjusted to a uniform size (for example, 256x256 pixels). At the same time, they will be normalized so that the pixel value range is between [0,1] to eliminate the brightness and contrast differences between different images. The noise in the image is removed by filtering technology to ensure the stability of model training and data quality. The preprocessed images will be stored in two folders, one for ordinary photos (domain A) and the other for X-ray images (domain B).
[0131] In terms of model configuration, the CycleGAN architecture is selected, which includes two generators (G: ordinary photo → X-ray image, F: X-ray image → ordinary photo) and two discriminators (D X : Determine whether the generated X-ray image is a real X-ray image, D Y : Determine whether the generated ordinary photo is a real ordinary photo).
[0132] The loss function consists of two parts: GAN Loss and Cycle-Consistency Loss. The weight coefficients of each part are λ gan = 0.5, and λ cycle =10.
[0133] L(G,F,D X ,D Y )
[0134] =λ gan L GAN (G,D Y ,X,Y)+λ gan L GAN (F,D X ,Y,X)+λ cycle L cycle(G, F) The optimizer uses the Adam optimizer, and the initial learning rate is η = 1×10 -5 , β1=0.5, β2=0.999.
[0135] In order to reduce the memory usage and speed up training, Xformers memory efficient attention mechanism is enabled (--enable_xformers_memory_efficient_attention). Considering the high resolution of a single image and the complexity of the model, a smaller batch size (--train_batch_size=1) is selected, and the gradient accumulation step (--gradient_accumulation_steps=1) is used to simulate the effect of a larger batch.
[0136] In the training process, the total number of training iterations is set to 25,000 (--max_train_steps=25000), and a validation is performed after every 250 training steps (--validation_steps=250) to monitor model performance and prevent overfitting. The entire training process will be tracked in real time through the WandB platform to record training progress, loss changes, and other important indicators.
[0137] During the model evaluation phase, the quality of the generated images will be measured using criteria such as the structural similarity index (SSIM) and peak signal-to-noise ratio (PSNR), and the similarity between the generated X-ray images and the real X-ray images will be visually checked. In addition, professional doctors will be invited to evaluate the generated images to ensure their clinical applicability. If possible, cross-validation methods will be used to test the stability and generalization ability of the model.
[0138] In the application deployment phase, the model weights with the best performance during the training process will be saved regularly, and a RESTful API service will be built to allow users to upload ordinary photos and obtain the corresponding X-ray images by calling the trained model. In order to simplify the operation process, a friendly graphical user interface (GUI) will be provided to doctors to facilitate them to quickly obtain the imaging data required for diagnosis. It should be noted that the model output is only for auxiliary diagnosis reference, and the final diagnosis result still needs to be made by professional doctors based on actual conditions. With the accumulation of more high-quality data and the advancement of technology, the model will continue to be optimized and updated to improve its accuracy and reliability.
[0139] Finally, in the entire process of data collection, processing and model application, we strictly abide by ethical standards and protect patient privacy, including data desensitization, anonymization, and compliance with HIPAA and other regulations to ensure data security and patient rights. We use algorithms to remove or replace personal information in images to ensure data anonymity, use differential privacy technology to protect patient privacy by adding random noise, and comply with relevant regulations such as HIPAA and GDPR to ensure the security and privacy of medical data.
[0140] In the embodiments of the present application, a method, device, storage medium and computer equipment for generating back X-ray images are provided. Compared with the traditional reliance on frequent X-ray shooting, the present invention can significantly reduce the risk of radiation exposure to patients, and at the same time provide an efficient and safe monitoring method to form realistic back X-ray images. The method can use an unpaired data set, which reduces the problem of difficulty in obtaining data sets for training models and improves the generalization ability of the model. The method includes: obtaining a first patient back photo taken by a shooting device; inputting the first patient back photo into a preset deep learning model to output a first back X-ray image; wherein the preset deep learning model is obtained by training according to a data set, and the data set includes a second patient back photo and a second back X-ray image, wherein the second back X-ray image includes a third back X-ray image paired with the second patient back photo and / or a fourth back X-ray image unpaired with the second patient back photo.
[0141] Further, as Figure 1 In a specific implementation, an embodiment of the present invention provides a device for generating a back X-ray image, such as Figure 3 As shown, the device includes: an acquisition unit 31 and an output unit 32.
[0142] An acquisition unit, used for acquiring a back photo of the first patient taken by a photographing device;
[0143] An output unit is used to input the first patient's back photograph into a preset deep learning model to output a first back X-ray image; wherein the preset deep learning model is trained based on a data set, and the data set includes a second patient's back photograph and a second back X-ray image, wherein the second back X-ray image includes a third back X-ray image paired with the second patient's back photograph and / or a fourth back X-ray image not paired with the second patient's back photograph.
[0144] In a specific application scenario, the device further includes:
[0145] A construction unit, used to construct a target deep learning model of a CycleGAN architecture, wherein the target deep learning model includes a first generator, a second generator, a first discriminator, a second discriminator and a loss function; wherein the first generator is used to convert a second patient's back photo into a fifth back X-ray image, and the first discriminator is used to predict the probability of whether the fifth back X-ray image converted by the first generator is a real back X-ray image; the second generator is used to convert the second back X-ray image into a third patient's back photo, and the second discriminator is used to predict the probability of whether the third patient's back photo converted by the second generator is a real patient's back photo; the loss function includes an adversarial loss and a cycle consistency loss; the adversarial loss includes the adversarial loss of the first generator and the first discriminator, and also includes the adversarial loss of the second generator and the second discriminator; the cycle consistency loss includes the cycle consistency loss of the first generator and the second generator;
[0146] A training unit is used to train a target deep learning model using the data set to determine a preset deep learning model.
[0147] In a specific application scenario, the training unit is specifically used to: use the data set to train the target deep learning model, and make the overall optimization target meet the preset requirements during the training process; wherein, the overall optimization target is determined based on the adversarial loss function and the cycle consistency loss function, and the preset requirements include minimizing the overall optimization target so that the first generator and the second generator generate realistic images, and also include maximizing the overall optimization target so that the first discriminator and the second discriminator maximize the ability to distinguish between real images and generated images.
[0148] In a specific application scenario, the formula for the overall optimization objective is:
[0149] L(G,F,D X ,D Y )=λ gan L GAN (G,D Y ,X,Y)+λ gan L GAN (F,D X ,Y,X)+
[0150] λ cycle L cycle (G,F);
[0151] Among them, L GAN (G,D Y ,X,Y) is the adversarial loss between the first generator and the first discriminator; L GAN (F,D X ,Y,X) is the adversarial loss of the second generator and the second discriminator; λgan and λ cycle are weight coefficients; L cycle (G, F) is the cycle consistency loss.
[0152] In a specific application scenario, the device further includes:
[0153] An evaluation unit, used to evaluate the trained target deep learning model;
[0154] The first determination unit is used to determine that the trained target deep learning model is a preset deep learning model if the evaluation result reaches a preset result.
[0155] In a specific application scenario, the evaluation unit includes:
[0156] A second determination unit is used to determine a structural similarity index, a peak signal-to-noise ratio, and a Cobb angle error using the trained target deep learning model;
[0157] The third determination unit is used to determine that the evaluation result reaches the preset result if the structural similarity index is greater than a third preset value, the peak signal-to-noise ratio is greater than a fourth preset value, and the Cobb angle error is less than a fifth preset value.
[0158] In a specific application scenario, the device further includes: a preprocessing unit for preprocessing the second patient's back photo and the second back X-ray image in the data set, respectively, wherein the preprocessing includes cropping, normalization and / or filtering technology processing.
[0159] It should be noted that for other corresponding descriptions of the functional modules involved in the apparatus for generating a back X-ray image provided by the embodiment of the present invention, reference can be made to Figure 1 The corresponding description of the method shown will not be repeated here.
[0160] Based on the above Figure 1 The method shown, accordingly, an embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the following steps are implemented: obtaining a first patient back photo taken by using a shooting device; inputting the first patient back photo into a preset deep learning model to output a first back X-ray image; wherein the preset deep learning model is trained according to a data set, and the data set includes a second patient back photo and a second back X-ray image, wherein the second back X-ray image includes a third back X-ray image paired with the second patient back photo and / or a fourth back X-ray image not paired with the second patient back photo.
[0161] Based on the above Figure 1 The method shown and Figure 3The embodiment of the device shown in the figure, the embodiment of the present invention also provides a physical structure diagram of a computer device, such as Figure 4 As shown, the computer device includes: a processor 41, a memory 42, and a computer program stored in the memory 42 and executable on the processor, wherein the memory 42 and the processor 41 are both arranged on a bus 43, and the processor 41 implements the following steps when executing the program: obtaining a first patient back photograph taken by a shooting device; inputting the first patient back photograph into a preset deep learning model to output a first back X-ray image; wherein the preset deep learning model is obtained by training based on a data set, and the data set includes a second patient back photograph and a second back X-ray image, wherein the second back X-ray image includes a third back X-ray image paired with the second patient back photograph and / or a fourth back X-ray image not paired with the second patient back photograph.
[0162] Through the technical solution of the present invention, compared with the traditional reliance on frequent X-ray shooting, the present invention can significantly reduce the risk of radiation exposure to patients, while providing an efficient and safe monitoring method to form realistic back X-ray images. In addition, the method can use an unpaired data set, which reduces the problem of difficulty in obtaining a data set for training the model and improves the generalization ability of the model. The method includes: obtaining a first patient back photo taken by a shooting device; inputting the first patient back photo into a preset deep learning model to output a first back X-ray image; wherein the preset deep learning model is trained based on a data set, the data set includes a second patient back photo and a second back X-ray image, wherein the second back X-ray image includes a third back X-ray image paired with the second patient back photo and / or a fourth back X-ray image unpaired with the second patient back photo.
[0163] Obviously, those skilled in the art should understand that the above modules or steps of the present invention can be implemented by a general computing device, they can be concentrated on a single computing device, or distributed on a network composed of multiple computing devices, and optionally, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, and in some cases, the steps shown or described can be executed in a different order than here, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.
[0164] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A method for generating a back X-ray image, characterized in that: include: Acquire a back photograph of the first patient taken by a photographing device; The first patient's back photograph is input into a preset deep learning model to output a first back x-ray image; wherein the preset deep learning model is trained based on a data set, and the data set includes a second patient's back photograph and a second back x-ray image, wherein the second back x-ray image includes a third back x-ray image paired with the second patient's back photograph and / or a fourth back x-ray image unpaired with the second patient's back photograph.
2. The method according to claim 1, characterized in that The training steps of the preset deep learning model include: Construct a target deep learning model of a cyclegan architecture, the target deep learning model comprising a first generator, a second generator, a first discriminator, a second discriminator and a loss function; wherein the first generator is used to convert the second patient's back photo into a fifth back X-ray image, and the first discriminator is used to predict the probability of whether the fifth back X-ray image converted by the first generator is a real back X-ray image; the second generator is used to convert the second back X-ray image into a third patient's back photo, and the second discriminator is used to predict the probability of whether the third patient's back photo converted by the second generator is a real patient's back photo; the loss function comprises an adversarial loss and a cycle consistency loss; the adversarial loss comprises the adversarial loss of the first generator and the first discriminator, and also comprises the adversarial loss of the second generator and the second discriminator; the cycle consistency loss comprises the cycle consistency loss of the first generator and the second generator; Using the data set, a target deep learning model is trained to determine a preset deep learning model.
3. The method according to claim 2, characterized in that The step of using the data set to train the target deep learning model to determine the preset deep learning model includes: The target deep learning model is trained using the data set, and during the training process, the overall optimization target is made to meet preset requirements; wherein, the overall optimization target is determined based on the adversarial loss function and the cycle consistency loss function, and the preset requirements include minimizing the overall optimization target so that the first generator and the second generator generate realistic images, and also include maximizing the overall optimization target so that the first discriminator and the second discriminator maximize the ability to distinguish between real images and generated images.
4. The method according to claim 3, characterized in that The formula for the overall optimization objective is: L(G,F,D X ,D Y )=λ gan L GAN (G,D Y ,X,Y)+λ gan L GAN (F,D X ,Y,X)+ l cycle L cycle (G,F); Among them, L GAN (G,D Y ,X,Y) is the adversarial loss between the first generator and the first discriminator; L GAN (F,D X ,Y,X) is the adversarial loss of the second generator and the second discriminator; λ gan and λ cycle are weight coefficients; L cycle (G,F) is the cycle consistency loss.
5. The method according to claim 4, characterized in that Before determining the preset deep learning model, the method further includes: Evaluate the trained target deep learning model; If the evaluation result reaches the preset result, the trained target deep learning model is determined to be the preset deep learning model.
6. The method according to claim 5, characterized in that The step of evaluating the trained target deep learning model includes: Determining a structural similarity index, a peak signal-to-noise ratio, and a Cobb angle error using the trained target deep learning model; If the structural similarity index is greater than a third preset value, the peak signal-to-noise ratio is greater than a fourth preset value, and the Cobb angle error is less than a fifth preset value, it is determined that the evaluation result reaches a preset result.
7. The method according to claim 2, characterized in that Before using the data set to train the target deep learning model to determine the preset deep learning model, the method further includes: The second patient's back photograph and the second back X-ray image in the data set are preprocessed respectively, wherein the preprocessing includes clipping, normalization and / or filtering technology processing.
8. A device for generating a back x-ray image, characterized in that: include: An acquisition unit, used for acquiring a back photo of the first patient taken by a photographing device; An output unit is used to input the first patient's back photograph into a preset deep learning model to output a first back X-ray image; wherein the preset deep learning model is trained based on a data set, and the data set includes a second patient's back photograph and a second back X-ray image, wherein the second back X-ray image includes a third back X-ray image paired with the second patient's back photograph and / or a fourth back X-ray image not paired with the second patient's back photograph.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.