Generation model training method, training device, image generation method, image generation device, and program

By training a generative model with specific data and prompts, the method enhances image contrast, addressing accuracy issues and enabling effective diagnosis with reduced contrast agent use.

WO2025249085A1PCT designated stage Publication Date: 2025-12-04TEIKYO UNIVERSITY
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2025/016611
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-05-27
Filing Date
2025-05-02
Publication Date
2025-12-04

AI Technical Summary

Technical Problem

Existing methods for generating contrast-enhanced images without using contrast agents face accuracy issues and struggle to specify conditions effectively in generative models, leading to reduced image contrast that hinders accurate diagnosis.

Method used

A method for training a generative model using specific training data and prompts to enhance image contrast, including information about the contrast agent concentration, inspection method, energy, and artifacts, allowing the model to generate images with higher contrast similar to those obtained with normal concentration contrast agents.

Benefits of technology

The method improves image contrast, enabling accurate diagnosis by generating images with sufficient contrast using low-concentration contrast agents, thereby reducing side effects while maintaining diagnostic quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2025016611_04122025_PF_FP_ABST
    Figure JP2025016611_04122025_PF_FP_ABST
Patent Text Reader

Abstract

This generation model training method uses, as training data: first data D1 including a first training image which is obtained by using a contrast agent and includes a first region representing a target tissue, and a first prompt pertaining to the first training image; and second data D2 including a second training image that includes a second region, representing the target tissue, and that has a lower contrast than the contrast of the first training image, and a second prompt pertaining to the second training image. The generation model training method uses said training data to train a generation model M so as to execute a generation process for, in accordance with a condition prompt Px and an original image X including a third region representing the target tissue, generating a converted image Y which includes a fourth region corresponding to the third region and which has a contrast higher than the contrast of the original image X. The first prompt includes information indicating that the first training image was obtained using a contrast agent having a normal concentration, the second prompt includes information indicating that the second training image was obtained using a contrast agent having a concentration lower than a normal concentration, and the condition prompt represents a condition of the generation process and includes information indicating that a contrast agent having a normal concentration is used.
Need to check novelty before this filing date? Find Prior Art

Description

Method for training generative model, training device, image generation method, image generation device, and program

[0001] The present disclosure relates to techniques for improving the contrast of images obtained using contrast agents.

[0002] In various diagnostic imaging techniques, contrast agents are used to enhance the contrast between a target tissue and other tissues, such as angiography, X-ray fluoroscopy, computed tomography (CT), ultrasound imaging, and magnetic resonance imaging (MRI).

[0003] However, depending on the type of contrast agent, it may cause side effects in the human body. For example, iodine-based contrast agents administered intravascularly are known to cause acute contrast-induced nephropathy as a major side effect. Therefore, in consideration of safety for the human body, it is desirable to reduce the amount of contrast agent used. However, reducing the amount of contrast agent used results in a problem of reduced contrast. From the perspective of making an accurate diagnosis, reduced contrast is undesirable.

[0004] Therefore, a technique has been proposed for obtaining a contrast-enhanced image (i.e., an image similar to that obtained using a contrast agent of a normal concentration) without using a contrast agent (or with a low concentration contrast agent).

[0005] For example, Patent Document 1 discloses a technology aimed at obtaining contrast-enhanced MRI images (magnetic resonance images) without using a contrast agent. Specifically, the technology of Patent Document 1 establishes a generative model using MRI images taken without using a contrast agent and MRI images taken with a contrast agent as training data. When the MRI image taken without using a contrast agent is input to the generative model, a contrast-enhanced MRI image is output. However, the technology of Patent Document 1 has accuracy problems because the contrast-enhanced MRI image is generated from an MRI image not using a contrast agent.

[0006] On the other hand, Patent Document 2 discloses a technology for establishing a generative model using, as training data, a full-contrast-dose image captured using a normal dose of contrast agent, a low-contrast-dose image captured using a contrast agent at a dose less than the full dose, and a zero-contrast-dose image captured without administering contrast agent. Then, by inputting the low-contrast-dose image and the zero-contrast-dose image into the generative model, a full-contrast-dose image (i.e., an image with enhanced contrast) is output. Patent Document 2 anticipates that a full-contrast-dose image can be generated with high accuracy by taking into account not only the zero-contrast-dose image but also the low-contrast-dose image.

[0007] Special table No. 2022-547047 Publication No. 2020-536638

[0008] However, the technology of Patent Document 2 was unable to generate a desired full contrast agent dose image after easily specifying various conditions for the generation model. In consideration of the above circumstances, the present invention aims to improve the contrast of an image obtained using a contrast agent by easily specifying conditions in a generation model.

[0009] [1] A method for training a generative model implemented by a computer, using as training data: a first training image, the first training image being an image obtained using a contrast agent, the first training image including a first region representing a target tissue; and a first prompt related to the first training image; a second training image including a second region representing the target tissue and having a lower contrast than the first training image; and second data including the second prompt related to the second training image; the method training a generative model to perform a generation process to generate an original image including a third region representing the target tissue and a transformed image including a fourth region corresponding to the third region and having a higher contrast than the original image in response to a conditional prompt; the first prompt including information indicating that the first training image was obtained using a contrast agent of a normal concentration; the second prompt including information indicating that the second training image was obtained using a contrast agent of a lower concentration than the normal concentration; and the conditional prompt including information indicating that a contrast agent of a normal concentration is used as a condition for the generation process.

[0010] [2] A method for training a generative model according to [1], wherein the first prompt includes information about the inspection method used to capture the first training image.

[0011] [3] A method for training a generative model according to [1] or [2], wherein the first prompt includes information about the energy used to capture the first training image.

[0012] [4] A method for training a generative model according to any one of [1] to [3], wherein the first prompt includes information representing the type of artifact.

[0013] [5] A method for training a generative model according to any one of [1] to [4], wherein the first prompt includes information about noise.

[0014] [6] A training device comprising: a training unit that trains a generative model to perform a generation process using as training data: a first training image, the first training image including a first region representing a target tissue, and a first prompt related to the first training image; a second training image including a second region representing the target tissue and having a lower contrast than the first training image; and second data including the second prompt related to the second training image; and to generate an original image including a third region representing the target tissue and, in response to a conditional prompt, a transformed image including a fourth region corresponding to the third region and having a higher contrast than the original image; wherein the first prompt includes information indicating that the first training image was obtained using a contrast agent of a normal concentration; the second prompt includes information indicating that the second training image was obtained using a contrast agent of a lower concentration than the normal concentration; and the conditional prompt includes information indicating a condition for the generation process, indicating that a contrast agent of a normal concentration should be used.

[0015] [7] A program that causes a computer to function as a training unit that trains a generative model to perform a generation process using as training data: a first training image, the first training image including a first region representing a target tissue, and a first prompt related to the first training image; a second training image including a second region representing the target tissue and having a lower contrast than the first training image; and second data including the second prompt related to the second training image; and generates an original image including a third region representing the target tissue and, in response to a conditional prompt, a converted image including a fourth region corresponding to the third region and having a higher contrast than the original image, wherein the first prompt includes information indicating that the first training image was obtained using a contrast agent of a normal concentration; the second prompt includes information indicating that the second training image was obtained using a contrast agent of a lower concentration than the normal concentration; and the conditional prompt includes information indicating a condition for the generation process, indicating that a contrast agent of a normal concentration should be used.

[0016] [8] An image generation method implemented by a computer that executes a generation process to generate the transformed image by inputting the original image and the conditional prompt into a generative model trained by the training method of [1] to [5].

[0017] [9] The image generation method of [7], wherein the target tissue is a tubular tissue, and in the generation process, the line structure is extracted from the original image, and the converted image is generated based on the extracted line structure and the condition prompt.

[0018]

[10] An image generation device comprising a generation unit that executes a generation process to generate the transformed image by inputting the original image and the conditional prompt into a generation model trained by the training method of [1] to [5].

[0019]

[11] A program that causes a computer system to function as a generation unit that executes a generation process to generate the transformed image by inputting the original image and the condition prompt into a generation model trained by the training method of [1] to [5].

[0020] According to the present invention, the contrast of images obtained using a contrast agent can be improved by simply specifying conditions in a generative model.

[0021] 1 is a block diagram illustrating a configuration of an image processing system according to an embodiment; FIG. 2 is a schematic diagram illustrating an original image and a transformed image according to an embodiment; FIG. 3 is a block diagram illustrating a functional configuration of an image processing system according to an embodiment; FIG. 4 is a schematic diagram illustrating a plurality of first data and a plurality of second data according to an embodiment; FIG. 5 is a flowchart illustrating an example of a training process according to an embodiment; FIG. 6 is a flowchart illustrating an example of a generation process according to an embodiment; 7 is an actual original image and a transformed image; 8 is an actual original image and a transformed image; 9 is an actual original image and a transformed image; 10 is an actual original image and a transformed image; 11 is an actual original image and a transformed image; 12 is a graph illustrating Michelson contrast for the actual original image and the transformed image; 13 is a graph illustrating RMS contrast for the actual original image and the transformed image; 14 is a graph illustrating the entropy relationship between the actual original image and the transformed image; 15 is a graph illustrating SNR for the actual original image and the transformed image.

[0022] FIG. 1 is a block diagram illustrating the configuration of an image processing system 100 according to an embodiment. The image processing system 100 is a computer system for generating a new image (hereinafter referred to as a "converted image Y") by converting an existing image (hereinafter referred to as an "original image X"). That is, the image processing system 100 is a system that performs image processing to generate the converted image Y from the original image X. The original image X and the converted image Y are images of tissue (hereinafter referred to as "target tissue") to be observed in a subject (typically a human), and are composed of a plurality of pixels arranged in a matrix. The target tissue is, for example, a tubular tissue. In the following description, a case where the target tissue is a blood vessel will be described as an example.

[0023] FIG. 2 shows a schematic diagram of an original image X and a converted image Y. The original image X is an image obtained by administering a contrast agent to a subject. Specifically, the original image X is an image including a region representing a target tissue V (hereinafter referred to as a "target region Rx1") and other regions (hereinafter referred to as a "non-target region Rx2"). The target region Rx1 is a portion that is emphasized by the contrast agent, and the non-target region Rx2 is a region other than the region representing the target tissue V (i.e., a region including the background and tissues other than the target tissue V). The target region Rx1 is an example of a "third region."

[0024] In this embodiment, the original image X is an image obtained by digital subtraction angiography (DSA). Digital subtraction angiography is an examination method in which X-ray images of a subject are obtained before and after the administration of a contrast agent, and a DSA image showing blood vessels is obtained by subtracting the X-ray image before the administration of a contrast agent (mask image) from the X-ray image after the administration of a contrast agent (original image). The original image X is typically a grayscale image.

[0025] In this embodiment, an image obtained by administering a contrast agent of a lower concentration (hereinafter referred to as "low concentration") than a contrast agent of a normal concentration (hereinafter referred to as "normal concentration") is exemplified as original image X. Original image X is a DSA image obtained by subtracting an X-ray image taken before administering the contrast agent to the subject from an X-ray image taken after administering the low-concentration contrast agent to the subject. The contrast agent is, for example, an iodine-based contrast agent.

[0026] The converted image Y is an image corresponding to the original image X, and includes a region representing the target tissue V (hereinafter referred to as the "target region Ry1") and other regions (hereinafter referred to as the "non-target region Ry2"). The target region Ry1 is a region corresponding to the target region Rx1 of the original image X, and the non-target region Ry2 is a region corresponding to the non-target region Rx2 of the original image X. The target region Ry1 is an example of the "fourth region."

[0027] 2, the contrast of the converted image Y is higher than the contrast of the original image X. In other words, the converted image Y is an image in which the contrast of the original image X has been improved. In other words, the converted image Y is an image with a contrast similar to that obtained by capturing an image using a contrast agent of a normal concentration.

[0028] In an image obtained using a contrast agent, the difference between the density (pixel value) of the target region representing the target tissue V and the density (pixel value) of the other non-target regions becomes larger (i.e., the contrast becomes larger). The higher the concentration of the contrast agent, the higher the contrast of the image, making it easier to grasp the target tissue V. On the other hand, the lower the concentration of the contrast agent, the lower the contrast of the image, making it more difficult to grasp the target tissue V.

[0029] The contrast in an image is an index that represents the degree of spread of the distribution of pixel values ​​in the image. In this embodiment, the difference between the average pixel value in the target region and the average pixel value in the non-target region can be said to be the contrast of the image. In this embodiment, the contrast is, for example, Michelson contrast. Michelson contrast is the image contrast calculated by (Lmax - Lmin) / (Lmax + Lmin). Lmax is the maximum luminance value in the image, and Lmin is the minimum luminance value in the image. Alternatively, the RMS (Root Mean Square) contrast may be used as the contrast. As described above, since the contrast agent increases the difference between the pixel values ​​in the target region and the pixel values ​​in the non-target region, typically, the Michelson contrast and RMS contrast also increase in an image using a contrast agent.

[0030] As can be understood from the above explanation, the contrast of the original image X corresponds to low density, and the contrast of the converted image Y corresponds to normal density.

[0031] The normal concentration is a general concentration determined, for example, according to the weight of the subject (test subject), and is changed appropriately depending on the type of target tissue and the subject's weight. For example, when the target tissue is a blood vessel and the subject is a typical adult male, the normal concentration (iodine content) of the contrast agent is approximately 150 to 400 mgI / mL. In contrast, the low concentration is set as follows, from the perspective of sufficiently reducing the burden on the subject and obtaining a converted image Y with high accuracy. For the same subject, the lower limit of the low concentration is, for example, 0.1 times or more the normal concentration, preferably 0.25 times or more the normal concentration, and more preferably 0.5 times or more the normal concentration. The upper limit of the low concentration is, for example, 0.9 times or less the normal concentration, preferably 0.8 times or less the normal concentration, and more preferably 0.7 times or less the normal concentration.

[0032] However, the normal concentration is not uniquely defined, but refers to the concentration of the contrast agent at which an image with high enough contrast to enable a medical professional to clearly identify the target tissue can be captured. Furthermore, the low concentration is not uniquely defined, but refers to the concentration of the contrast agent at which an image with low enough contrast to enable a medical professional to clearly identify the target tissue can be captured. In other words, in this embodiment, regardless of the concentration of the contrast agent actually used, an image with sufficient contrast for accurate diagnosis is defined as an image using a normal-concentration contrast agent, and an image with insufficient contrast for accurate diagnosis is defined as an image using a low-concentration contrast agent.

[0033] The image processing system 100 is realized by an information device such as a smartphone, a tablet terminal, or a personal computer. As illustrated in Fig. 1, the image processing system 100 includes a control device 11, a storage device 12, a display device 13, and an input device 14. The image processing system 100 may be realized as a single device, or may be realized as a plurality of devices configured separately from each other.

[0034] The control device 11 is composed of one or more processors that control each element of the image processing system 100. For example, the control device 11 is composed of one or more types of processors, such as a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a field programmable gate array (FPGA), or an application specific integrated circuit (ASIC).

[0035] The storage device 12 is one or more memories that store programs executed by the control device 11 and various data used by the control device 11. The storage device 12 is configured with a known storage medium such as a magnetic storage medium or a semiconductor storage medium. The storage device 12 may also be configured with a combination of multiple types of storage media. Furthermore, the storage device 12 may be a portable storage medium that is detachable from the image processing system 100, or a storage medium to which the control device 11 can write or read data via a communication network (e.g., cloud storage).

[0036] The storage device 12 stores, for example, an original image X. The original image X is captured by any known imaging device (imaging diagnostic device). In this embodiment, the original image X is obtained by an imaging device (e.g., a blood vessel imaging device) that can capture the original image X by irradiating the human body with X-rays. That is, the X-rays are the energy used to capture the original image X (DSA image). The energy used to capture the original image X differs depending on the type of imaging device.

[0037] The display device 13 displays an image under the control of the control device 11. For example, the original image X and the converted image Y are displayed on the display device 13. The display device 13 is configured with a display panel such as a liquid crystal panel or an organic EL (Electroluminescence) panel.

[0038] The input device 14 is an input device that accepts instructions from a user. The input device 14 is, for example, an operator operated by the user or a touch panel that detects contact by the user. Note that the display device 13 or the input device 14, which are separate from the image processing system 100, may be connected to the image processing system 100 via a wired or wireless connection. Instructions from the user may also be input by voice. In this embodiment, for example, prompts (Px, P1, P2) described below are input by the input device 14.

[0039] 3 is a block diagram illustrating an example of the functional configuration of the image processing system 100. The control device 11 executes a program stored in the storage device 12 to realize multiple functions (a training unit 112, a generation unit 114, and a generative model M) for processing a transformed image Y from an original image X.

[0040] A generative model M is used to generate a transformed image Y using an original image X. The generative model M is a statistical model that generates a transformed image Y in response to the original image X and a prompt (hereinafter referred to as a "condition prompt Px") that represents a condition for specifying the generation of the transformed image Y.

[0041] Specifically, the generative model M is a model that is configured, for example, based on a neural network and is capable of generating an image based on the condition prompt Px. In this embodiment, a diffusion model (DDPM: Denoising Diffusion Probabilistic Models) is used that generates a new image through a de-diffusion process that iteratively removes noise from an image. An example of such a diffusion model is Stable Diffusion. For example, a model that is an extension of Stable Diffusion (e.g., InstructPix2Pix or ControlNet) is preferably used as the generative model M. However, the generative model M is not limited to Stable Diffusion as long as it is a model that can generate a converted image Y from an original image X based on the condition prompt Px. Details of the condition prompt Px will be described later.

[0042] The generative model M is realized by a combination of a program (e.g., a program module constituting artificial intelligence software) that causes the control device 11 to execute a calculation to generate the converted image Y, and multiple parameters that are applied to the calculation.

[0043] The multiple parameters that define the generative model M are optimized by prior machine learning and stored in the storage device 12. Specifically, the multiple parameters of the generative model M are set so that an image close to the initial image is generated when a diffusion process that iteratively adds normally distributed noise to an image proceeds in the reverse direction (de-diffusion process). The multiple parameters are established by training using training data.

[0044] The training unit 112 uses training data to train the generative model M. The training data includes a plurality of first data D and a plurality of second data D. Fig. 4 is a schematic diagram showing the plurality of first data D and the plurality of second data D.

[0045] The first data D1 includes a first training image T1 and a first prompt P1 representing the first training image T1. The first training image T1 is an image obtained using a normal concentration contrast agent and corresponds to the converted image Y. In this embodiment, as described above, regardless of the concentration of the contrast agent actually used, if the contrast is sufficient to perform an accurate diagnosis, the image obtained using a normal concentration contrast agent is designated as the first training image T1. Specifically, the first training image T1 includes a first region R1 (an example of a "first region") representing the target tissue V and a second region R2 other than the first region R1.

[0046] The second data D2 includes a second training image T2 and a second prompt P2 representing the second training image T2. In this embodiment, the second training image T2 is an image obtained using a low-concentration contrast agent and corresponds to the original image X. That is, the contrast of the second training image T2 is lower than that of the first training image T1. As described above, in this embodiment, regardless of the concentration of the contrast agent actually used, if the contrast is insufficient for accurate diagnosis, the image obtained using a low-concentration contrast agent is designated as the second training image T2. Specifically, the second training image T2 includes a third region R3 (an example of a "second region") representing the target tissue V and a fourth region R4 other than the third region R3. The second region R2 and the fourth region R4 may include the background, tissue other than the target tissue, or a medical instrument.

[0047] The first prompt P1 and the second prompt P2 will be described below. When it is not necessary to distinguish between the first training image T1 and the second training image T2, they will be simply referred to as "training images." The training images are the same type of images as the original image X, and are images captured using the same examination method (digital subtraction angiography) as the original image X.

[0048] The first prompt P1 is information about the first training image T1. The second prompt P2 is information about the second training image T2. In this embodiment, the first prompt P1 and the second prompt P2 are illustrated as character strings in a natural language (so-called character tags). For example, the first prompt P1 and the second prompt P2 may each include character strings indicating the following information (1) to (8):

[0049] <Information (1): Information Representing the Name of Target Tissue Included in Training Image> Information (1) is, for example, a character string (for example, "blood_vessel") representing a blood vessel that is the target tissue.

[0050] <Information (2): Information Representing the Concentration of the Contrast Agent in the Training Image> In the first prompt P1, the information (2) is a character string representing that the contrast agent is of normal concentration, such as a character string (e.g., "normal_contrast") representing the contrast of the first training image T1 (i.e., the contrast when a normal concentration of contrast agent is used). As described above, the contrast when a normal concentration of contrast agent is used refers to contrast that is sufficient for accurate diagnosis. In the second prompt P2, the information (2) is a character string representing that the contrast agent is of low concentration, such as a character string (e.g., "low_contrast") representing the contrast of the second training image T2 (i.e., the contrast when a low concentration of contrast agent is used). As described above, the contrast when a low concentration of contrast agent is used refers to contrast that is not sufficient for accurate diagnosis.

[0051] <Information (3): Information Regarding the Examination Method Used to Capture the Training Images> As described above, the examination method in this embodiment is digital subtraction angiography (DSA). Therefore, information (3) is, for example, a character string representing digital subtraction angiography (e.g., "dsa"). Alternatively, information (3) may be a character string representing the imaging device (imaging diagnostic device) used in the examination method.

[0052] <Information (4): Information Regarding the Energy Used to Capture the Training Images> In this embodiment, the energy is X-rays, as described above. Therefore, information (4) is a character string representing X-rays (e.g., "x-ray"). The energy used to capture the training images varies depending on the inspection method and inspection device, and may include ultrasound waves or electromagnetic waves.

[0053] <Information (5): Information Representing the Type of Medical Instrument> For example, if a training image includes a medical instrument (e.g., a catheter), a character string representing the type of this medical instrument (e.g., "catheter") is included as information (5). Furthermore, when performing IVR (Interventional Radiology), a coil (a rolled wire) may be placed in a blood vessel for treatment purposes. When such a coil is used, an image of the coil may be included in the training image, or artifacts (unwanted components or disturbances) may occur. Therefore, it is preferable that information (5) include a character string representing the coil remaining in the blood vessel (e.g., "coil"). Note that, for training images that include a medical instrument other than a catheter, whether inside or outside the body, a common character string (e.g., "coil") may be included as information (5) regardless of the type of medical instrument, in order to conveniently specify that the medical instrument is a medical instrument other than a catheter.

[0054] <Information (6): Information Representing the Type of Artifact> For example, the training images may contain artifacts (unwanted components or disturbances) related to the subject's movement. The movement of the subject (typically a human) includes, for example, not only the subject's bodily movement but also movement caused by the subject's breathing, heartbeat, pulsation, or intestinal peristalsis. If the training images contain such artifacts related to the subject's movement (so-called motion artifacts), it is preferable to include a string representing the subject's movement (e.g., "motion_artifact") as information (6). The type of artifact is not limited to motion artifacts, and may also include, for example, artifacts derived from the imaging device or blood flow.

[0055] <Information (7): Noise-Related Information> When capturing training images using X-rays, a portion of the X-ray beam (up and down and / or left and right) may be blocked near the X-ray irradiation port to avoid unnecessary radiation exposure. In this case, noise may occur at a position in the training image corresponding to the blocked portion (typically, at the edge of the training image) due to the influence of X-rays scattered by the subject. Therefore, in this embodiment, it is preferable to include a character string (e.g., "noisy_frame") representing noise that occurs at a specific position in the training image when a portion of the X-ray beam is blocked. However, the noise-related information is not limited to the above examples. For example, various types of information such as information representing the type or name of noise, information representing the location where the noise occurs, or information representing the cause of the noise are exemplified as noise-related information (7).

[0056] <Information (8): Other Information> The first prompt P1 and the second prompt P2 may include various other information as appropriate. Information (8) is, for example, a character string indicating the type of training image (e.g., grayscale / monochrome / color) (e.g., "greyscale" for grayscale).

[0057] The training unit 112 trains the generative model M by inputting a first training image T1 to which a first prompt P1 including information (1) to (8) is assigned and a second training image T2 to which a second prompt P2 including information (1) to (8) is assigned. The generative model M is trained by adjusting multiple parameters to generate an optimal output for the input images and prompts. The assignment of the first prompt P1 to the first training image T1 and the second prompt P2 to the second training image T2 is performed, for example, by an administrator of the image processing system 100. However, it is not necessary to assign all of the above-described information (1) to (8) as the first prompt P1 and the second prompt P2; it is sufficient that at least information (2) is assigned. A transformed image Y is generated from an original image X using the generative model M established as described above.

[0058] A plurality of first data D and a plurality of second data D are used to train the generative model M. The plurality of first training images T are, for example, images of target tissues V (blood vessels) of different subjects captured using a normal concentration of contrast agent, and an appropriate first prompt P is assigned to each of the plurality of first training images T. The assigned first prompt P may vary depending on the content of the first training image T. Note that the concentration of the contrast agent used in each first training image T may differ for each subject.

[0059] Similarly, the second training images T2 are images of target tissues V (blood vessels) of different subjects captured using a low concentration of contrast agent, and an appropriate second prompt P2 is assigned to each of the second training images T2. The second prompt P2 assigned may vary depending on the content of the second training image T2. Note that the concentration of the contrast agent used in each second training image T2 may vary from subject to subject.

[0060] FIG. 5 is a flowchart illustrating an example of a process (training process) for training the generative model M executed by the control device 11. For example, the training process of FIG. 5 is initiated in response to an instruction from an administrator of the image processing system 100 via the input device 14. When the process of FIG. 5 is initiated, the training unit 112 receives training data (Sa1). Specifically, the training unit 112 receives a plurality of first data D1 including a first training image T1 and a first prompt P1, and a plurality of second data D2 including a second training image T2 and a second prompt P2. The first data D1 and the second data D2 are input by a user operating the input device 14. The plurality of first data D1 and the plurality of second data D2 may be received individually in order or collectively. The training unit 112 inputs the received training data into the generative model M (Sa2) and trains the generative model M (Sa3).

[0061] The generative model M may be trained (established) by performing additional learning (for example, fine tuning or transfer learning) on ​​an existing general-purpose model using training data.

[0062] The generator 114 executes a generation process for generating a transformed image Y by inputting the original image X and the condition prompt Px into the generation model M. Therefore, a transformed image Y (an image having sufficient contrast to allow for an accurate diagnosis) similar to that obtained using a normal-concentration contrast agent is generated from the original image X (an image having insufficient contrast to allow for an accurate diagnosis) obtained using a low-concentration contrast agent. In other words, the contrast of the original image X obtained using the contrast agent can be improved.

[0063] The conditional prompt Px is information (in this embodiment, a character string in a natural language) that represents the conditions for the generation process. The conditional prompt Px in this embodiment includes information indicating elements and features that should be included in the converted image Y (so-called positive prompts) and information indicating elements and features that should be suppressed in the converted image Y (so-called negative prompts). Specifically, the positive prompts allow the user to specify elements and features that should be maintained in the original image X and elements and features that should be changed or added to the original image X. The negative prompts allow the user to specify elements and features that should be excluded (reduced) from the original image X. Examples of information included in the conditional prompt Px include, for example, the following information (I) to (VII), which is typically similar to the information (information (1) to (8)) used in the first prompt P1 and the second prompt P2. Like the first prompt P1 and the second prompt P2, the conditional prompt Px in this embodiment is a character string in a natural language.

[0064] <Information (I): Information Representing the Name of Target Tissue> Information (I) is information representing the name of the target tissue desired in the transformed image Y (to be included in the transformed image Y), and like information (1), is a character string representing a blood vessel that is the target tissue (e.g., "blood_vessel"). In other words, information (I) is a character string representing the name of the target tissue included in the original image. Information (I) is specified as a positive prompt.

[0065] <Information (II): Information Representing Contrast> Information (II) is the desired contrast in the converted image Y, and is information representing the contrast obtained when a normal concentration contrast agent is used. Therefore, information (II) is the same as the character string (e.g., "normal_contrast") as information (2) in the first prompt P1. Information (II) is specified as a positive prompt.

[0066] <Information (III): Information Regarding Examination Method> Information (III) is information indicating the examination method by which the converted image Y is to be captured, and, like information (3), is a character string representing digital subtraction angiography (e.g., "dsa"). In other words, information (III) is a character string regarding the examination method used to capture the original image. Information (III) is specified as a positive prompt.

[0067] <Information (IV): Information Regarding Energy> Information (IV) is information indicating what energy the converted image Y will appear to have been captured with, and like information (4), is, for example, a character string representing X-rays (e.g., "x-ray"). In other words, information (IV) is a character string regarding the energy used to capture the original image. Information (IV) is specified as a positive prompt.

[0068] <Information (V): Information Representing the Type of Medical Instrument> Information (V) includes, for example, the following information (V-1) and information (V-2).

[0069] Information (V-1) is information indicating the type of medical instrument included in the converted image Y. In this embodiment, the original image X (DSA image) typically includes a catheter. Therefore, it is preferable that information (V-1) include a character string ("catheter") representing the catheter. Information (V-1) is specified as a positive prompt. Note that in this embodiment, it is common for a catheter to be visible in the DSA image creation process, and suppressing the catheter creates a sense of incongruity in the converted image Y. Also, considering that the catheter is located in a position in the original image X that does not overlap with the observation position of the target tissue V, it is preferable that the character string representing the catheter be a positive prompt.

[0070] On the other hand, information (V-2) is information indicating the type of medical instrument to be suppressed in the converted image Y. In this embodiment, a character string (e.g., "coil") indicating the type of medical instrument other than a catheter is exemplified as information (V-2). In this embodiment, it is preferable to specify medical instruments other than catheters as negative prompts because they interfere with observation.

[0071] <Information (VI): Information Representing the Type of Artifact> Information (VI) is information representing the type of artifact to be suppressed in the converted image Y, and like information (6), is, for example, a character string representing the subject's movement (e.g., "motion_artifact"). Information (VI) is specified as a negative prompt.

[0072] <Information (VII): Information about noise> Information (VII) is information about noise to be suppressed in the converted image Y. Like information (7), information (VII) is, for example, a character string (e.g., "noisy_frame") representing noise that occurs at a specific position in the training image when part of the X-ray beam is blocked. Information (VII) is specified as a negative prompt.

[0073] The information included in the conditional prompt Px is not limited to information (I) to (VII). For example, a character string (e.g., "greyscale") representing the desired type of image in the converted image Y may be included as a positive prompt. Furthermore, information (I) to (VII) may emphasize or suppress each piece of information. For example, it is conceivable to emphasize information (II) in comparison with other pieces of information. It is sufficient for the conditional prompt Px to include at least information (II). In other words, in the positive prompt, information (I), (III) to (V-1) are elements or features that are desired to be maintained from the original image X, and information (II) is an element or feature that is desired to be changed from the original image X.

[0074] In practice, the prompts (first prompt P1, second prompt P2, and conditional prompt Px) are input to the generative model M as feature vectors that alternatively represent the prompts received from the user.

[0075] In this embodiment, as described above, the target tissue is a blood vessel (tubular tissue). It is assumed that the blood vessel is linear in the original image X. Therefore, in the generation process of this embodiment, it is preferable to extract a line structure (i.e., a region corresponding to the blood vessel) and generate a converted image Y based on the extracted line structure and the condition prompt Px. Any known image processing technique (e.g., ControlNet) is used to extract the line structure in the converted image Y. Because the converted image Y is generated based on the line structure extracted from the original image X, a converted image Y can be obtained that further emphasizes the blood vessels in the original image X. Furthermore, various techniques for removing noise from the original image X can also be used as appropriate.

[0076] FIG. 6 is a flowchart illustrating an example of the generation process executed by the control device 11. For example, the generation process of FIG. 6 is initiated in response to an instruction from a user via the input device 14. When the process of FIG. 6 is initiated, the generation unit 114 receives an original image X and a condition prompt Px assigned to the original image X (Sb1). The original image X and the condition prompt Px are input by a user operating the input device 14. The generation unit 114 inputs the received original image X and condition prompt Px into the generative model M (Sb2) and outputs a converted image Y (Sa3). The output converted image Y is displayed on the display device 13. In other words, the generation process is a process of generating a converted image Y with sufficient contrast for accurate diagnosis from an original image X with insufficient contrast for accurate diagnosis.

[0077] As can be understood from the above explanation, by inputting the original image X and the condition prompt Px including information indicating that a normal concentration contrast agent is to be used into the generation model M, it is possible to generate a converted image Y having higher contrast than the original image X. In other words, by simply specifying conditions in the generation model M using the condition prompt Px, it is possible to generate a converted image Y with improved contrast.

[0078] Since the first prompt P1 and the second prompt P2 contain information about the inspection method used to capture the training image, by specifying information about the inspection method as the condition prompt Px, a converted image Y can be generated that takes into account the characteristics of the inspection method.

[0079] Since the first prompt P1 and the second prompt P2 contain information about the energy used to capture the training image, by specifying information about the energy as the condition prompt Px, a converted image Y can be generated that takes into account the characteristics of that energy.

[0080] Since the first prompt P1 and the second prompt P2 contain information representing the type of artifact contained in the training image, by specifying information representing the type of artifact as the conditional prompt Px (negative prompt), it is possible to generate a converted image Y in which the artifact is suppressed.

[0081] 7 to 11 illustrate, for reference, an actual original image X (Input) and a transformed image Y (Output) obtained using the generative model M according to this embodiment. The original image X is a DSA image obtained by administering a low-concentration contrast agent to a subject (including images captured using a low-concentration contrast agent due to poor contrast for some reason even when a normal-concentration contrast agent is used).

[0082] The 1328 original images X and the 1328 converted images Y generated from each original image X were evaluated as follows (FIGS. 12 to 15).

[0083] 12 and 13 show the results (box plots) of calculating the Michelson contrast and RMS contrast for the original image X and the converted image Y. As can be seen from Fig. 12 and Fig. 13, it was confirmed that the converted image Y tends to have higher values ​​than the original image X for both the Michelson contrast and the RMS contrast.

[0084] Furthermore, a quantitative evaluation using entropy was performed on the converted image Y. Entropy is an index that indicates the disorder of an image, and tends to take a larger value as the contrast increases.

[0085] Specifically, the image was histogrammed and normalized, and then the entropy was calculated. The entropy is calculated as -ΣP i log2P i (P i is the probability of occurrence of each pixel value). The results are shown in Figure 14. As can be seen from Figure 14, the entropy of the converted image Y also tends to be higher than that of the original image X. From the above, it can be said that the contrast of the converted image Y has been improved.

[0086] Figure 15 shows the results of calculating the SNR (signal-noise ratio) for the original image X and the converted image Y. As can be seen from Figure 15, it was confirmed that the SNR also dropped significantly for the converted image Y. Note that the results of Figures 12 to 15 can be further improved by further increasing the amount of training data.

[0087] <Modifications> Specific modifications that can be added to the above-described embodiments are exemplified below. Multiple embodiments arbitrarily selected from the following examples may be combined as appropriate within the scope of not contradicting each other.

[0088] (1) In the above-described embodiment, the second training image T2 of the second data D2 used as training data may be an image obtained by actually administering a low-concentration contrast agent to a subject, or may be an image generated by image processing so that the contrast between the third region R3 and the fourth region R4 is the same as that obtained by using a low-concentration contrast agent (hereinafter referred to as a “pseudo-low-concentration image”).

[0089] For example, the pseudo-low-density image is generated by image processing a first training image T1 captured using a normal-density contrast agent to generate a second training image T2. Note that the first training image T1 used to generate the pseudo-low-density image may or may not be the same as the first training image T1 used to train the generative model M.

[0090] Specifically, the image processing is a process for generating a second training image T2 from a first training image T1 in which the contrast between the third region R3 and the fourth region R4 (the density difference between the third region R3 and the fourth region R4) is lower than the contrast between the first region R1 and the second region R2 in the first training image T1 (the density difference between the first region R1 and the second region R2). In other words, a second training image T2 corresponding to a low concentration of contrast agent is generated. As preprocessing, the resolution of the first training image T1 is set to 1024 x 1024 and converted to 256 gradations (8 bits). However, the preprocessing of the first training image T1 is not limited to the above example.

[0091] The image processing generates a second training image T2 that includes a third region R3 obtained by changing the pixel values ​​(density) of the first region R1 according to the density value of the low density, and a fourth region R4 obtained by maintaining the pixel values ​​(density) of the second region R2. Through the above image processing, the second training image T2 is generated in which the contrast between the third region R3 and the fourth region R4 is lower than the contrast between the first region R1 and the second region R2.

[0092] The third region R3 is obtained by changing the pixel values ​​of each pixel in the first region R1 to approach the pixel values ​​of the second region R2 (i.e., to reduce the contrast between the first region R1 and the second region R2). The pixel value of the second region R2 here is, for example, the mode of the second region R2. It is assumed that the area of ​​the second region R2 in the first training image T1 is significantly larger than that of the first region R1. Therefore, typically, the mode of the first training image T1 becomes the mode of the second region R2. However, for example, a user may visually select the pixel value of any pixel in the second region R2 as the pixel value of the second region R2.

[0093] Assuming a first training image T1 in which the first region R1 is close to black (0) and the second region R2 is close to white (255), the third region R3 of the second training image T2 will be lighter (i.e., have a larger pixel value) than the first region R1 of the first training image T1.

[0094] For example, the second training image T2 is generated by converting the pixel value of each pixel in the first training image T1 using the following equation (1): G1 is the pixel value of the pixel in the first training image T1, and G2 is the converted pixel value (i.e., the pixel value of the pixel in the second training image T2). α is a value corresponding to the low density value. BG is a value obtained by subtracting a value (e.g., 20) corresponding to noise (e.g., noise caused by a living body or an imaging device) from the mode in the first training image T1 (i.e., the mode in the second region R2).

[0095] G2=BG-α(BG-G1)...(1)

[0096] Of the pixels in the first training image T1, those in which G1 is smaller than BG (G1<BG) (i.e., the first region R1) are subjected to conversion using equation (1). For the first training image T1, pixels in which G1 is smaller than BG are typically pixels corresponding to the first region R1 (i.e., pixels representing the target tissue V). The pixel values ​​of each pixel in the first region R1 are a superposition of pixel values ​​attributable to the background and pixel values ​​attributable to the target tissue V. The term "α(BG-G1)" in equation (1) is calculated by extracting pixel values ​​attributable to the target tissue V for the pixels in the first region R1 and then multiplying them by α.

[0097] For example, if G1 is 50, BG is 200, and α is 0.2 (low density is 0.2 times the normal density), G2 will be 170. Note that the formula for converting the pixel value of each pixel in the first training image T1 (first region R1) is not limited to the above formula (1) as long as it is possible to convert pixel values ​​according to the density value of the low density.

[0098] In other words, the contrast between the third region R3 and the fourth region R4 in the second training image T2 corresponds to the contrast between the region representing the target tissue and the other regions in an image obtained using a low concentration contrast agent.

[0099] A plurality of second training images T2 may be generated from a single first training image T1. That is, a plurality of second training images T2 having different contrasts between the third region R3 and the fourth region R4 are generated from a single first training image T1. Specifically, a plurality of second training images T2 are generated from a single first training image T1 by varying the density value of the low density. For example, a plurality of second training images T2 are generated by setting the low density at predetermined intervals (e.g., 0.02 times) within a range of 0.25 to 0.9 times the normal density.

[0100] By using a pseudo low-density image generated from the first training image T1 as the second training image T2, a large number of second training images T2 can be generated without administering a low-concentration contrast agent to the subject.

[0101] If it is possible to generate a second training image T2 in which the contrast between the third region R3 and the fourth region R4 is lower than the contrast between the first region R1 and the second region R2, the second training image T2 may be generated by changing both the pixel values ​​of the first region R1 and the second region R2 (for example, by changing the range of available pixel values). However, if the second training image T2 includes the fourth region R4 obtained by maintaining the pixel values ​​in the second region R2, the second training image T2 can more closely resemble an image obtained in an actual clinical setting using a low-concentration contrast agent or an image using a normal-concentration contrast agent but with reduced contrast.

[0102] (2) In the above-described embodiment, the original image X includes, for example, an image that looks like it was taken using a low-concentration contrast agent (an image with insufficient contrast for accurate diagnosis) due to poor contrast caused by some reason (e.g., insufficient contrast due to the patient's body size or insufficient flow rate or velocity of the contrast agent) even though a normal-concentration contrast agent was used. Similarly, the second training image T2 includes an image that looks like it was taken using a low-concentration contrast agent (an image with insufficient contrast for accurate diagnosis).

[0103] Furthermore, in the present disclosure, the term "low concentration contrast agent" includes not only a case where the concentration of the contrast agent to be administered is reduced by diluting the prepared contrast agent with physiological saline, but also a case where a contrast agent of a normal concentration is used and the flow rate or flow velocity of the contrast agent injection is reduced, thereby reducing the administered amount and the effective concentration in the blood vessel.

[0104] (3) In the above-described embodiment, the training image and the original image X are images captured using digital subtraction angiography (DSA images). However, the first training image T1 and the original image X are not limited to DSA images. For example, images captured using a contrast agent in various imaging techniques, such as angiography other than digital subtraction angiography, X-ray fluoroscopy, computed tomography (CT), ultrasound imaging, and magnetic resonance imaging (MRI), can be used as the first training image T1 and the original image X. The type of contrast agent administered to the subject can be changed as appropriate depending on the type of imaging technique. Furthermore, the target tissue V is not limited to blood vessels, but can be any tissue that can be enhanced by a contrast agent, and may be tubular tissue other than blood vessels (e.g., vascular organs such as lymphatic vessels, bile ducts, and pancreatic ducts).

[0105] (4) The subject to which the present invention is applicable is not limited to humans, but may be other animals.

[0106] (5) In the above-described embodiment, for example, training images or original images X obtained by using a contrast agent of a normal concentration on simulated tissue (for example, simulated blood vessels or simulated organs) may be used.

[0107] (6) The image processing system 100 is an example of a training device and an image generating device. However, the training device and the image generating device may each be configured as separate, independent devices.

[0108] (7) The present invention can also be conceived as a program that causes a computer to function as a training unit that trains a generative model to perform a generation process using as training data: a first training image, the first training image including a first region representing a target tissue, and a first prompt representing the first training image; a second training image including a second region representing the target tissue and having a lower contrast than the first training image; and second data including the second prompt representing the second training image; and generates an original image including a third region representing the target tissue and, in response to a conditional prompt, a converted image including a fourth region corresponding to the third region and having a higher contrast than the original image, wherein the first prompt includes information indicating that the first training image was obtained using a contrast agent of a normal concentration; the second prompt includes information indicating that the second training image was obtained using a contrast agent of a lower concentration than the normal concentration; and the conditional prompt includes information indicating a condition for the generation process, indicating that a contrast agent of a normal concentration should be used.

[0109] The present invention can also be conceived as a program that causes a computer system to function as a generation unit that executes a generation process to generate the transformed image by inputting an original image and a condition prompt into the generation model.

[0110] The program and generative model according to the present invention can be provided, for example, in a form stored on a computer-readable recording medium, or in a form distributed from a distribution device via a communication network and installed on a computer.

[0111] 11: Control device 12: Storage device 13: Display device 14: Input device 100: Image processing system 112: Training unit 114: Generation unit D1: First data D2: Second data M: Generative model P1: First prompt P2: Second prompt R1: First region R2: Second region R3: Third region R4: Fourth region Rx1, Ry1: Target region Rx2, Ry2: Non-target region T1: First training image T2: Second training image V: Target tissue X: Original image Y: Transformed image

Claims

1. A computer-implemented method for training a generative model using as training data: first data including first training images, the first training images being images obtained using a contrast agent and including a first region representing a target tissue, and a first prompt related to the first training images; second data including second training images, the second training images including a second region representing the target tissue and having a lower contrast than the first training images, and a second prompt related to the second training images; and training a generative model to perform a generation process to generate, in response to a conditional prompt, an original image including a third region representing the target tissue and a transformed image including a fourth region corresponding to the third region and having a higher contrast than the original image, wherein the first prompt includes information indicating that the first training images were obtained using a normal concentration of contrast agent; the second prompt includes information indicating that the second training images were obtained using a lower-than-normal concentration of contrast agent; and the conditional prompt includes information indicating a condition for the generation process, including information indicating that a normal concentration of contrast agent is used.

2. The method for training a generative model according to claim 1, wherein the first prompt includes information about an inspection method used to capture the first training image.

3. The method for training a generative model according to claim 1, wherein the first prompt includes information about the energy used to capture the first training image.

4. The method for training a generative model according to claim 1, wherein the first prompt includes information representing a type of artifact.

5. The method for training a generative model of claim 1, wherein the first prompt includes information about noise.

6. A training device comprising a training unit that trains a generative model to execute a generation process using as training data: a first training image, the first training image being an image obtained using a contrast agent and including a first region representing a target tissue, and a first prompt related to the first training image; a second training image, the second training image being including a second region representing the target tissue and having a lower contrast than the first training image, and second data including a second prompt related to the second training image; and to generate, in response to a conditional prompt, an original image including a third region representing the target tissue and a converted image having a higher contrast than the original image, the converted image including a fourth region corresponding to the third region, the converted image having a higher contrast than the original image; wherein the first prompt includes information indicating that the first training image was obtained using a normal concentration of contrast agent; the second prompt includes information indicating that the second training image was obtained using a contrast agent at a concentration lower than the normal concentration; and the conditional prompt includes information indicating a condition for the generation process, indicating that a normal concentration of contrast agent should be used.

7. A program that causes a computer to function as a training unit that trains a generative model using as training data: a first training image, the first training image being an image obtained using a contrast agent and including a first region representing a target tissue, and a first prompt related to the first training image; a second training image, the second training image being including a second region representing the target tissue and having a lower contrast than the first training image, and second data being including a second prompt related to the second training image; and executes a generation process to generate, in response to a conditional prompt, an original image including a third region representing the target tissue and a converted image having a higher contrast than the original image, the converted image including a fourth region corresponding to the third region, the converted image having a higher contrast than the original image, the converted image including a fourth region corresponding to the third region, the converted image having a higher contrast than the original image, the converted image being obtained using a contrast agent of a normal concentration; the first prompt including information indicating that the first training image was obtained using a contrast agent of a normal concentration; the second prompt including information indicating that the second training image was obtained using a contrast agent of a lower concentration than the normal concentration; and the conditional prompt including information indicating a condition for the generation process, the converted image being obtained using a contrast agent of a normal concentration.

8. An image generation method implemented by a computer, which executes a generation process to generate the transformed image by inputting the original image and the conditional prompt to a generative model trained by the training method of claim 1.

9. The image generation method according to claim 8, wherein the target tissue is a tubular tissue, and the generation process extracts the line structures in the original image, and generates the converted image based on the extracted line structures and the condition prompt.

10. An image generation device comprising a generation unit that executes a generation process to generate the transformed image by inputting the original image and the conditional prompt to a generation model trained by the training method of claim 1.

11. A program that causes a computer system to function as a generation unit that executes a generation process to generate the transformed image by inputting the original image and the condition prompt into a generation model trained by the training method of claim 1.

Citation Information

Patent Citations

  • Deep virtual contrast

    EP3739522A1

  • Deep Spectral Bolus Tracking

    JP2022509016A

  • Machine learning contrast enhancement

    JP2024507766A