Training method of generation model, training device, image generation method, image generation device, and program

By training a generative model with specific prompts and data types, the method enhances image contrast for accurate diagnosis while reducing contrast agent use, addressing the challenge of side effects and accuracy in existing techniques.

JP2025178956APending Publication Date: 2025-12-09TEIKYO UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024085848
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-05-27
Publication Date
2025-12-09

AI Technical Summary

Technical Problem

Existing image enhancement techniques using contrast agents face challenges in achieving high contrast without causing side effects, and existing generative models struggle to accurately generate images with full contrast agent dose without specifying conditions.

Method used

A method for training a generative model using specific prompts and data types to enhance contrast, including information about contrast agent concentration, inspection method, energy, and artifacts, allowing the model to generate images with improved contrast.

Benefits of technology

The method effectively improves image contrast by specifying conditions, generating images with sufficient contrast for accurate diagnosis while reducing the need for high contrast agent doses, thus minimizing side effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025178956000001_ABST
    Figure 2025178956000001_ABST
Patent Text Reader

Abstract

To improve the contrast of an image obtained using a contrast medium by simply designating conditions in a generation model.SOLUTION: There is provided a generation model training method for training a generation model M for generating, according to an original image X including a third region of a target tissue and a condition prompt Px representing a condition of a generation process, a converted image Y including a fourth region corresponding to the third region and having a contrast higher than that of the original image X, by using, as training data, a first data D1 including a first training image including a first region of the target tissue obtained by using a contrast agent of normal concentration and a first prompt representing that the first region is obtained by using the contrast agent of the normal concentration, a second training image including a second region of the target tissue and having a contrast lower than that of the first training image, and a second data D2 including a second prompt representing that the second region is obtained by using a contrast agent having a low concentration.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to techniques for improving the contrast of images obtained using contrast agents. [Background technology]

[0002] In various imaging diagnostics, contrast agents are used to enhance the contrast of tissues to be observed (hereinafter referred to as "target tissues"), such as angiography, X-ray fluoroscopy, computed tomography (CT), ultrasound imaging, and magnetic resonance imaging (MRI).

[0003] However, depending on the type of contrast agent, it may cause side effects in the human body. For example, iodine-based contrast agents administered intravascularly are known to cause acute contrast-induced nephropathy as a major side effect. Therefore, in consideration of safety for the human body, it is desirable to reduce the amount of contrast agent used. However, reducing the amount of contrast agent used results in a problem of reduced contrast. From the perspective of making an accurate diagnosis, reduced contrast is undesirable.

[0004] Therefore, a technique has been proposed for obtaining a contrast-enhanced image (that is, an image similar to that obtained using a contrast agent of a normal concentration) without using a contrast agent (or using a contrast agent of a low concentration).

[0005] For example, Patent Document 1 discloses a technology for establishing a generative model using MRI images taken without a contrast agent and MRI images taken with a contrast agent as training data, with the aim of obtaining contrast-enhanced MRI images (magnetic resonance images) without the use of a contrast agent. When the MRI images taken without a contrast agent are input into the generative model, a contrast-enhanced MRI image is output. However, the technology in Patent Document 1 has accuracy problems because the contrast-enhanced MRI image is generated from an MRI image taken without a contrast agent.

[0006] On the other hand, Patent Document 2 discloses a technology for establishing a generation model using, as training data, a full contrast agent dose image captured using a normal dose of contrast agent, a low contrast agent dose image captured using a contrast agent at a dose less than the full contrast agent dose, and a zero contrast agent dose image captured without administering contrast agent. Then, by inputting the low contrast agent dose image and the zero contrast agent dose image into the generation model, a full contrast agent dose image (i.e., an image with enhanced contrast) is output. Patent Document 2 anticipates that a full contrast agent dose image can be generated with high accuracy by taking into account not only the zero contrast agent dose image but also the low contrast agent dose image. [Prior art documents] [Patent documents]

[0007] [Patent Document 1] Special Publication No. 2022-547047 [Patent Document 2] Special Publication No. 2020-536638 Summary of the Invention [Problem to be solved by the invention]

[0008] However, the technology of Patent Document 2 was unable to generate a full contrast agent dose image after easily specifying various conditions. In consideration of the above circumstances, the present invention aims to improve the contrast of an image obtained using a contrast agent by easily specifying conditions in a generation model. [Means for solving the problem]

[0009] [1] A method for training a generative model implemented by a computer, using as training data: a first training image, the first training image being an image obtained using a contrast agent, the first training image including a first region representing a target tissue; and a first prompt related to the first training image; a second training image including a second region representing the target tissue and having a lower contrast than the first training image; and second data including a second prompt related to the second training image; the method training a generative model to perform a generation process to generate an original image including a third region representing the target tissue and a converted image, in response to a conditional prompt, the converted image including a fourth region corresponding to the third region and having a higher contrast than the original image; the first prompt including information indicating that the first training image was obtained using a contrast agent of a normal concentration; the second prompt including information indicating that the second training image was obtained using a contrast agent of a lower concentration than the normal concentration; and the conditional prompt including information indicating that a contrast agent of a normal concentration is used as a condition for the generation process.

[0010] [2] The method for training a generative model according to [1], wherein the first prompt includes information about the inspection method used to capture the first training image.

[0011] [3] The method for training a generative model according to [1] or [2], wherein the first prompt includes information about the energy used to capture the first training image.

[0012] [4] A method for training a generative model according to any one of [1] to [3], wherein the first prompt includes information representing the type of artifact.

[0013] [5] The method for training a generative model according to any one of [1] to [4], wherein the first prompt includes information about noise.

[0014] [6] A training device comprising: a training unit that trains a generative model to perform a generation process using as training data: a first training image, the first training image including a first region representing a target tissue, and a first prompt related to the first training image; a second training image including a second region representing the target tissue and having a lower contrast than the first training image; and second data including the second prompt related to the second training image; and to generate, in response to a conditional prompt, an original image including a third region representing the target tissue and a converted image including a fourth region corresponding to the third region and having a higher contrast than the original image; wherein the first prompt includes information indicating that the first training image was obtained using a contrast agent of a normal concentration; the second prompt includes information indicating that the second training image was obtained using a contrast agent of a lower concentration than the normal concentration; and the conditional prompt includes information indicating a condition for the generation process, indicating that a contrast agent of a normal concentration should be used.

[0015] [7] A program that causes a computer to function as a training unit that trains a generative model to perform a generation process using as training data: a first training image, the first training image including a first region representing a target tissue, and a first prompt related to the first training image; a second training image, the second training image including a second region representing the target tissue and having a lower contrast than the first training image; and second data including the second prompt related to the second training image; and generates an original image including a third region representing the target tissue and, in response to a conditional prompt, a converted image including a fourth region corresponding to the third region and having a higher contrast than the original image, wherein the first prompt includes information indicating that the first training image was obtained using a contrast agent of a normal concentration; the second prompt includes information indicating that the second training image was obtained using a contrast agent of a lower concentration than the normal concentration; and the conditional prompt includes information indicating a condition for the generation process, indicating that a contrast agent of a normal concentration should be used.

[0016] [8] An image generation method implemented by a computer that executes a generation process to generate the transformed image by inputting the original image and the conditional prompt into a generative model trained by the training method of [1] to [5].

[0017] [9] The image generation method of [7], wherein the target tissue is a tubular tissue, and in the generation process, the line structure is extracted from the original image, and the converted image is generated based on the extracted line structure and the condition prompt.

[0018]

[10] An image generation device comprising a generation unit that executes a generation process to generate the transformed image by inputting the original image and the condition prompt into a generative model trained by the training method of [1] to [5].

[0019]

[11] A program that causes a computer system to function as a generation unit that executes a generation process to generate the transformed image by inputting the original image and the condition prompt into a generative model trained by the training method of [1] to [5]. [Effects of the Invention]

[0020] According to the present invention, the contrast of images obtained using a contrast agent can be improved by simply specifying conditions in a generative model. [Brief explanation of the drawings]

[0021] [Figure 1] FIG. 1 is a block diagram illustrating a configuration of an image processing system according to an embodiment. [Figure 2] 2A and 2B are schematic diagrams illustrating an original image and a converted image according to the embodiment. [Figure 3] FIG. 1 is a block diagram illustrating a functional configuration of an image processing system according to an embodiment. [Figure 4] FIG. 2 is a schematic diagram showing a plurality of first data and a plurality of second data according to the embodiment. [Figure 5]10 is a flowchart illustrating an example of a training process according to the embodiment. [Figure 6] 10 is a flowchart illustrating an example of a generation process according to the embodiment. [Figure 7] The actual original image and the transformed image. [Figure 8] The actual original image and the transformed image. [Figure 9] The actual original image and the transformed image. [Figure 10] The actual original image and the transformed image. [Figure 11] The actual original image and the transformed image. [Figure 12] 1 is a graph showing Michelson contrast for actual original and transformed images; [Figure 13] 10 is a graph showing the RMS contrast of an actual original image and a transformed image. [Figure 14] 10 is a graph showing the relationship between the entropy of an actual original image and a transformed image. [Figure 15] 10 is a graph showing the SNR of an actual original image and a transformed image. DETAILED DESCRIPTION OF THE INVENTION

[0022] FIG. 1 is a block diagram illustrating the configuration of an image processing system 100 according to an embodiment. The image processing system 100 is a computer system for generating a new image (hereinafter referred to as a "converted image Y") by converting an existing image (hereinafter referred to as an "original image X"). That is, the image processing system 100 is a system that executes image processing for generating the converted image Y from the original image X. The original image X and the converted image Y are images captured of tissue (hereinafter referred to as a "target tissue") to be observed in a subject (typically a human), and are composed of a plurality of pixels arranged in a matrix. The target tissue is, for example, a tubular tissue. In the following explanation, a case where the target tissue is a blood vessel will be exemplified.

[0023] FIG. 2 shows a schematic diagram of an original image X and a converted image Y. The original image X is an image obtained by using (administering) a contrast agent to a subject. Specifically, the original image X is an image including a region representing a target tissue V (hereinafter referred to as a "target region Rx1") and other regions (hereinafter referred to as a "non-target region Rx2"). The target region Rx1 is a portion that is emphasized by the contrast agent, and the non-target region Rx2 is a region other than the region representing the target tissue V (i.e., a region including the background and tissues other than the target tissue V). The target region Rx1 is an example of a "third region."

[0024] In this embodiment, the original image X is an image obtained by digital subtraction angiography (DSA). Digital subtraction angiography is an examination method in which X-ray images of a subject are obtained before and after the administration of a contrast agent, and a DSA image showing blood vessels is obtained by subtracting the X-ray image before the administration of a contrast agent (mask image) from the X-ray image after the administration of a contrast agent (original image). The original image X is typically a grayscale image.

[0025] In this embodiment, an image obtained by administering a contrast agent of a lower concentration (hereinafter referred to as "low concentration") than a contrast agent of a normal concentration (hereinafter referred to as "normal concentration") is exemplified as original image X. Original image X is a DSA image obtained by subtracting an X-ray image taken before administering the contrast agent to the subject from an X-ray image taken after administering the low-concentration contrast agent to the subject. The contrast agent is, for example, an iodine-based contrast agent.

[0026] The converted image Y is an image corresponding to the original image X, and includes a region representing the target tissue V (hereinafter referred to as the "target region Ry1") and other regions (hereinafter referred to as the "non-target region Ry2"). The target region Ry1 is a region corresponding to the target region Rx1 of the original image X, and the non-target region Ry2 is a region corresponding to the non-target region Rx2 of the original image X. The target region Ry1 is an example of the "fourth region."

[0027] 2, the contrast of the converted image Y is higher than the contrast of the original image X. In other words, the converted image Y is an image in which the contrast of the original image X has been improved. In other words, the converted image Y is an image with a contrast similar to that obtained by capturing an image using a contrast agent of a normal concentration.

[0028] In an image obtained using a contrast agent, the difference in density (pixel value) between the target region representing the target tissue V and the density (pixel value) of other non-target regions becomes larger (i.e., contrast is created). The higher the concentration of the contrast agent, the higher the contrast of the image, making it easier to identify the target tissue V. On the other hand, the lower the concentration of the contrast agent, the lower the contrast of the image, making it more difficult to identify the target tissue V.

[0029] The contrast in an image is an index that indicates the degree of spread of the distribution of pixel values ​​of the image. The contrast in this embodiment is, for example, Michelson contrast. The Michelson contrast is the contrast of an image calculated by (Lmax-Lmin) / (Lmax+Lmin). Lmax is the maximum luminance value in the image, and Lmin is the minimum luminance value in the image. Alternatively, the RMS (Root Mean Square) contrast may be used as the contrast.

[0030] As can be understood from the above explanation, the contrast of the original image X corresponds to low density, and the contrast of the converted image Y corresponds to normal density.

[0031] The normal concentration is, for example, a general concentration determined according to the weight of the subject (test subject), and is changed appropriately according to the type of target tissue and the weight of the subject. In contrast, the low concentration is set as follows, from the viewpoint of sufficiently reducing the burden on the subject and obtaining a converted image Y with high accuracy. For the same subject, the lower limit of the low concentration is, for example, 0.1 times or more the normal concentration, preferably 0.25 times or more the normal concentration, and more preferably 0.5 times or more the normal concentration. The upper limit of the low concentration is, for example, 0.9 times or less the normal concentration, preferably 0.8 times or less the normal concentration, and more preferably 0.7 times or less the normal concentration.

[0032] However, the normal concentration is not uniquely defined, but refers to the concentration of the contrast agent at which an image with high enough contrast to enable a medical professional to clearly identify the target tissue can be captured. Furthermore, the low concentration is not uniquely defined, but refers to the concentration of the contrast agent at which an image with low enough contrast to enable a medical professional to clearly identify the target tissue can be captured. In other words, in this embodiment, regardless of the concentration of the contrast agent actually used, an image with sufficient contrast for accurate diagnosis is defined as an image using a normal-concentration contrast agent, and an image with insufficient contrast for accurate diagnosis is defined as an image using a low-concentration contrast agent.

[0033] The image processing system 100 is realized by an information device such as a smartphone, a tablet terminal, or a personal computer. As illustrated in Fig. 1, the image processing system 100 includes a control device 11, a storage device 12, a display device 13, and an input device 14. The image processing system 100 may be realized as a single device, or may be realized as a plurality of devices configured separately from each other.

[0034] The control device 11 is composed of one or more processors that control each element of the image processing system 100. For example, the control device 11 is composed of one or more types of processors such as a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), or an ASIC (Application Specific Integrated Circuit).

[0035] The storage device 12 is one or more memories that store programs executed by the control device 11 and various data used by the control device 11. The storage device 12 is configured with a known storage medium such as a magnetic storage medium or a semiconductor storage medium. The storage device 12 may be configured with a combination of multiple types of storage media. Furthermore, the storage device 12 may be a portable storage medium that is detachable from the image processing system 100, or a storage medium (e.g., cloud storage) that the control device 11 can write to or read from via a communication network.

[0036] The storage device 12 stores, for example, an original image X. The original image X is captured by any known imaging device (imaging diagnostic device). In this embodiment, the original image X is obtained by an imaging device (e.g., a blood vessel imaging device) that can capture the original image X by irradiating the human body with X-rays. That is, the X-rays are the energy used to capture the original image X (DSA image). The energy used to capture the original image X differs depending on the type of imaging device.

[0037] Display device 13 displays an image under the control of control device 11. For example, original image X and converted image Y are displayed on display device 13. Display device 13 is configured with a display panel such as a liquid crystal panel or an organic EL (Electroluminescence) panel.

[0038] The input device 14 is an input device that accepts instructions from a user. The input device 14 is, for example, an operator operated by a user or a touch panel that detects contact by a user. Note that the display device 13 or the input device 14, which are separate from the image processing system 100, may be connected to the image processing system 100 by wire or wirelessly. Instructions from a user may also be input by voice. In this embodiment, for example, prompts (Px, P1, P2) described below are input by the input device 14.

[0039] 3 is a block diagram illustrating an example of the functional configuration of the image processing system 100. The control device 11 executes a program stored in the storage device 12 to realize multiple functions (a training unit 112, a generation unit 114, and a generative model M) for processing a converted image Y from an original image X.

[0040] A generative model M is used to generate a transformed image Y using an original image X. The generative model M is a statistical model that generates the transformed image Y in response to the original image X and a prompt (hereinafter referred to as a "condition prompt Px") that represents a condition for specifying the generation of the transformed image Y.

[0041] Specifically, the generative model M is configured, for example, based on a neural network and is capable of generating an image based on the condition prompt Px. In this embodiment, a diffusion model (DDPM: Denoising Diffusion Probabilistic Models) is used, which generates a new image through a de-diffusion process that iteratively removes noise from an image. An example of such a diffusion model is Stable Diffusion. Details of the condition prompt Px will be described later.

[0042] The generative model M is realized by a combination of a program (for example, a program module constituting artificial intelligence software) that causes the control device 11 to execute a calculation to generate the converted image Y, and a plurality of parameters that are applied to the calculation.

[0043] A plurality of parameters defining the generative model M are optimized by prior machine learning and stored in the storage device 12. Specifically, the plurality of parameters of the generative model M are set so that an image close to the initial image is generated when a diffusion process in which normally distributed noise is repeatedly added to an image proceeds in the reverse direction (de-diffusion process). The plurality of parameters are established by training using training data.

[0044] The training unit 112 uses training data to train the generative model M. The training data includes a plurality of first data D1 and a plurality of second data D2. Fig. 4 is a schematic diagram showing the plurality of first data D1 and the plurality of second data D2.

[0045] The first data D1 includes a first training image T1 and a first prompt P1 representing the first training image T1. The first training image T1 is an image obtained using a contrast agent of a normal concentration and corresponds to the converted image Y. In this embodiment, as described above, regardless of the concentration of the contrast agent actually used, if the contrast is sufficient for accurate diagnosis, the image obtained using a contrast agent of a normal concentration is set as the first training image T1. Specifically, the first training image T1 includes a first region R1 (an example of a "first region") representing the target tissue V and a second region R2 other than the first region R1.

[0046] The second data D2 includes a second training image T2 and a second prompt P2 representing the second training image T2. In this embodiment, the second training image T2 is an image obtained using a low-concentration contrast agent and corresponds to the original image X. That is, the contrast of the second training image T2 is lower than that of the first training image T1. As described above, in this embodiment, regardless of the concentration of the contrast agent actually used, if the contrast is insufficient for accurate diagnosis, the second training image T2 is an image obtained using a low-concentration contrast agent. Specifically, the second training image T2 includes a third region R3 (an example of a "second region") representing the target tissue V and a fourth region R4 other than the third region R3. The second region R2 and the fourth region R4 may include the background, tissue other than the target tissue, or a medical instrument.

[0047] The first prompt P1 and the second prompt P2 will be described below. When there is no need to distinguish between the first training image T1 and the second training image T2, they will be simply referred to as "training images." The training images are the same type of images as the original image X, and are images captured using the same examination method (digital subtraction angiography) as the original image X.

[0048] The first prompt P1 is information about the first training image T1. The second prompt P2 is information about the second training image T2. In this embodiment, the first prompt P1 and the second prompt P2 are illustrated as character strings in a natural language. For example, each of the first prompt P1 and the second prompt P2 may include character strings indicating the following information (1) to (8):

[0049] <Information (1): Information indicating the name of the target tissue contained in the training image> The information (1) is, for example, a character string (for example, "blood_vessel") that represents a blood vessel, which is the target tissue.

[0050] <Information (2): Information representing the concentration of the contrast agent in the training image> In the first prompt P1, the information (2) is a character string indicating that the contrast agent is of normal concentration, for example, a character string (e.g., "normal_contrast") indicating the contrast of the first training image T1 (i.e., the contrast when a contrast agent of normal concentration is used). As described above, the contrast when a contrast agent of normal concentration is used means the contrast sufficient for making an accurate diagnosis. In the second prompt P2, the information (2) is a character string indicating that the contrast agent is low in concentration, for example, a character string (e.g., "low_contrast") indicating the contrast of the second training image T2 (i.e., the contrast when a low concentration of contrast agent is used). As described above, the contrast when a low concentration of contrast agent is used means the contrast is insufficient to make an accurate diagnosis.

[0051] <Information (3): Information about the inspection method used to capture the training images> As described above, the examination method of this embodiment is digital subtraction angiography (DSA). Therefore, information (3) is, for example, a character string representing digital subtraction angiography (e.g., "dsa"). Alternatively, information (3) may be a character string representing the imaging device (imaging diagnostic device) used in the examination method.

[0052] <Information (4): Information about the energy used to capture the training images> In this embodiment, the energy is X-rays, as described above. Therefore, the information (4) is a character string representing X-rays (e.g., "x-ray"). The energy used to capture the training images varies depending on the inspection method and inspection device, and may also be ultrasound or electromagnetic waves.

[0053] <Information (5): Information indicating the type of medical device> For example, if a training image includes a medical device (e.g., a catheter), a character string representing the type of medical device (e.g., "catheter") is included as information (5). Furthermore, when performing IVR (Interventional Radiology), a coil (a rolled wire) may be placed in a blood vessel for treatment purposes. When such a coil is used, an image of the coil may be included in the training image, or artifacts (unwanted components or disturbances) may occur. Therefore, it is preferable that information (5) include a character string representing the coil remaining in the blood vessel (e.g., "coil"). Note that, for training images that include a medical device other than a catheter, whether inside or outside the body, a common character string (e.g., "coil") may be included as information (5) for convenience, regardless of the type of medical device, in order to specify that the medical device is a medical device other than a catheter.

[0054] <Information (6): Information indicating the type of artifact> For example, the training images may contain artifacts (unwanted components or disturbances) related to the subject's movement. The movement of the subject (typically a human) includes, for example, not only the subject's bodily movement but also movement caused by the subject's breathing, heartbeat, pulsation, or intestinal peristalsis. When the training images contain such artifacts related to the subject's movement (so-called motion artifacts), it is preferable to include a string representing the subject's movement (e.g., "motion_artifact") as information (6). The type of artifact is not limited to motion artifacts, and artifacts derived from the imaging device and blood flow, for example, are also possible.

[0055] <Information (7): Information about noise> When capturing training images using X-rays, a portion of the X-ray beam (top, bottom, and / or left and right) may be blocked near the X-ray irradiation port to avoid unnecessary exposure. In this case, noise may occur at a position in the training image corresponding to the blocked portion (typically, at the edge of the training image) due to the influence of X-rays scattered by the subject. Therefore, in this embodiment, it is preferable to include a character string (e.g., "noisy_frame") representing noise that occurs at a specific position in the training image when a portion of the X-ray beam is blocked. However, the information about noise is not limited to the above examples. For example, various types of information such as information representing the type or name of noise, information representing the position where the noise occurs, or information representing the cause of the noise are exemplified as the noise-related information (7).

[0056] <Information (8): Other information> The first prompt P1 and the second prompt P2 may include various other information as appropriate. The information (8) is, for example, a character string (e.g., “greyscale” for grayscale) indicating the type of training image (e.g., grayscale / monochrome / color).

[0057] The training unit 112 trains the generative model M by inputting a first training image T1 to which a first prompt P1 including information (1) to (8) is assigned and a second training image T2 to which a second prompt P2 including information (1) to (8) is assigned. The generative model M is trained by adjusting multiple parameters to generate an optimal output for the input image and prompt. Note that the assignment of the first prompt P1 to the first training image T1 and the second prompt P2 to the second training image T2 are performed, for example, by an administrator of the image processing system 100. However, it is not necessary to assign all of the above-mentioned information (1) to (8) as the first prompt P1 and the second prompt P2; it is sufficient that at least information (2) is assigned. A converted image Y is generated from an original image X using the generative model M established as described above.

[0058] A plurality of first data D1 and a plurality of second data D2 are used to train the generative model M. The plurality of first training images T1 are, for example, images of target tissues V (blood vessels) of different subjects captured using a normal concentration of contrast agent, and an appropriate first prompt P1 is assigned to each of the plurality of first training images T1. The assigned first prompt P1 may vary depending on the content of the first training image T1. Note that the concentration of the contrast agent used in each first training image T1 may vary for each subject.

[0059] Similarly, the second training images T2 are images of target tissues V (blood vessels) of different subjects captured using a low concentration of contrast agent, and an appropriate second prompt P2 is assigned to each of the second training images T2. The assigned second prompt P2 may vary depending on the content of the second training image T2. Note that the concentration of the contrast agent used in each second training image T2 may vary from subject to subject.

[0060] FIG. 5 is a flowchart showing an example of a process (training process) for training the generative model M executed by the control device 11. For example, the training process of FIG. 5 is started in response to an instruction from an administrator of the image processing system 100 via the input device 14. When the process of FIG. 5 is started, the training unit 112 receives training data (Sa1). Specifically, the training unit 112 receives a plurality of first data D1 including a first training image T1 and a first prompt P1, and a plurality of second data D2 including a second training image T2 and a second prompt P2. The first data D1 and the second data D2 are input by a user operating the input device 14. The plurality of first data D1 and the plurality of second data D2 may be received individually in order or collectively. The training unit 112 inputs the received training data to the generative model M (Sa2) and trains the generative model M (Sa3).

[0061] The generative model M may be trained (established) by performing additional learning (for example, fine tuning or transfer learning) on ​​an existing general-purpose model using training data.

[0062] The generation unit 114 executes a generation process for generating a transformed image Y by inputting the original image X and the condition prompt Px into the generation model M. Therefore, a transformed image Y (an image having sufficient contrast to make an accurate diagnosis) similar to that obtained using a normal-concentration contrast agent is generated from the original image X (an image having contrast that is not sufficient to make an accurate diagnosis) obtained using a low-concentration contrast agent. In other words, the contrast of the original image X obtained using the contrast agent can be improved.

[0063] The condition prompt Px is information (a character string in a natural language in this embodiment) that represents a condition for the generation process. The condition prompt Px in this embodiment includes information indicating elements and features that should be included in the converted image Y (a so-called positive prompt) and information indicating elements and features that should be suppressed in the converted image Y (a so-called negative prompt). Specifically, the positive prompt allows the user to specify elements and features that should be maintained in the original image X and elements and features that should be changed or added to the original image X. The negative prompt allows the user to specify elements and features that should be excluded (reduced) from the original image X. Examples of information included in the condition prompt Px include, for example, the following information (I) to (VII), and are typically similar to the information (information (1) to (8)) used in the first prompt P1 and the second prompt P2. The condition prompt Px in this embodiment is a character string in a natural language, similar to the first prompt P1 and the second prompt P2.

[0064] <Information (I): Information indicating the name of the target organization> Information (I) is information representing the name of the target tissue desired in the transformed image Y (to be included in the transformed image Y), and like information (1), is a string representing the blood vessel that is the target tissue (e.g., "blood_vessel"). In other words, information (I) is a string representing the name of the target tissue included in the original image. Information (I) is specified as a positive prompt.

[0065] <Information (II): Information that expresses contrast> Information (II) is the desired contrast in the converted image Y, and represents the contrast obtained when a normal concentration contrast agent is used. Therefore, information (II) is the same as the character string (e.g., "normal_contrast") as information (2) in the first prompt P1. Information (II) is specified as a positive prompt.

[0066] <Information (III): Information on testing methods> Information (III) is information indicating the examination method used to capture the converted image Y, and like information (3), is a character string (e.g., "dsa") representing digital subtraction angiography. Information (III) can also be described as a character string relating to the examination method used to capture the original image. Information (III) is specified as a positive prompt.

[0067] <Information (IV): Energy Information> The information (IV) indicates the energy used to capture the converted image Y, and like the information (4), is, for example, a character string representing X-rays (e.g., "x-ray"). In other words, the information (IV) is a character string relating to the energy used to capture the original image. The information (IV) is specified as a positive prompt.

[0068] <Information (V): Information indicating the type of medical device> The information (V) includes, for example, the following information (V-1) and information (V-2).

[0069] Information (V-1) is information indicating the type of medical instrument included in converted image Y. In this embodiment, original image X (DSA image) typically includes a catheter. Therefore, it is preferable that information (V-1) include a character string ("catheter") representing the catheter. Information (V-1) is specified as a positive prompt. Note that in this embodiment, it is common for a catheter to be visible in the DSA image creation process, and suppressing the catheter creates a sense of incongruity in converted image Y. Also, considering that the catheter is located in a position in original image X that does not overlap with the observation position of target tissue V, it is preferable that the character string representing the catheter be a positive prompt.

[0070] On the other hand, information (V-2) is information indicating the type of medical instrument to be suppressed in the converted image Y. In this embodiment, a character string (e.g., "coil") indicating the type of medical instrument other than a catheter is exemplified as information (V-2). In this embodiment, it is preferable to specify medical instruments other than catheters as negative prompts because they interfere with observation.

[0071] <Information (VI): Information indicating the type of artifact> Information (VI) is information that indicates the type of artifact to be suppressed in the converted image Y, and like information (6), is, for example, a character string (e.g., "motion_artifact") that indicates the subject's motion. Information (VI) is specified as a negative prompt.

[0072] <Information (VII): Information about noise> Information (VII) is information about noise to be suppressed in the converted image Y. Like information (7), it is, for example, a string (e.g., "noisy_frame") representing noise that occurs at a specific position in the training image when part of the X-ray beam is blocked. Information (VII) is specified as a negative prompt.

[0073] The information included in the conditional prompt Px is not limited to information (I) to (VII). For example, a character string (e.g., "greyscale") representing the desired type of image in the converted image Y may be included as a positive prompt. Furthermore, information (I) to (VII) may emphasize or suppress each piece of information. For example, it is assumed that information (II) is emphasized in comparison with other information. Note that the conditional prompt Px is sufficient to include at least information (II). In other words, in the positive prompt, information (I), (III) to (V-1) are elements or features that are desired to be maintained from the original image X, and information (II) is an element or feature that is desired to be changed from the original image X.

[0074] In practice, the prompts (first prompt P1, second prompt P2, and conditional prompt Px) are input to the generative model M as feature vectors that alternatively represent the prompts received from the user.

[0075] In this embodiment, as described above, the target tissue is a blood vessel (tubular tissue). It is assumed that the blood vessel is linear in the original image X. Therefore, in the generation process of this embodiment, it is preferable to extract a line structure (i.e., a region corresponding to the blood vessel) and generate a converted image Y based on the extracted line structure and the condition prompt Px. Any known image processing technique (e.g., ControlNet) is used to extract the line structure in the converted image Y. Since the converted image Y is generated based on the line structure extracted from the original image X, it is possible to obtain a converted image Y in which the blood vessels in the original image X are more emphasized. In addition, various techniques for removing noise from the original image X may also be used as appropriate.

[0076] FIG. 6 is a flowchart showing an example of the generation process executed by the control device 11. For example, the generation process of FIG. 6 is started in response to an instruction from the user via the input device 14. When the process of FIG. 6 is started, the generation unit 114 receives an original image X and a condition prompt Px assigned to the original image X (Sb1). The original image X and the condition prompt Px are input by the user operating the input device 14. The generation unit 114 inputs the received original image X and condition prompt Px into the generative model M (Sb2) and outputs a converted image Y (Sa3). The output converted image Y is displayed on the display device 13. In other words, the generation process is a process of generating a converted image Y having sufficient contrast for accurate diagnosis from an original image X having insufficient contrast for accurate diagnosis.

[0077] As can be understood from the above explanation, by inputting the original image X and the condition prompt Px including information indicating that a normal concentration contrast agent is to be used into the generative model M, it is possible to generate a converted image Y with higher contrast than the original image X. In other words, by simply specifying the conditions in the generative model M using the condition prompt Px, it is possible to generate a converted image Y with improved contrast.

[0078] Since the first prompt P1 and the second prompt P2 contain information about the inspection method used to capture the training image, by specifying information about the inspection method as the condition prompt Px, a converted image Y can be generated that takes into account the characteristics of the inspection method.

[0079] Since the first prompt P1 and the second prompt P2 contain information about the energy used to capture the training image, by specifying information about the energy as the condition prompt Px, a converted image Y can be generated that takes into account the characteristics of that energy.

[0080] Since the first prompt P1 and the second prompt P2 contain information representing the type of artifact contained in the training image, by specifying information representing the type of artifact as a conditional prompt Px (negative prompt), a converted image Y can be generated in which the artifact is suppressed.

[0081] 7 to 11 illustrate, for reference, an actual original image X (Input) and a transformed image Y (Output) obtained using the generative model M according to this embodiment. The original image X is a DSA image obtained by administering a low-concentration contrast agent to a subject (including images captured using a low-concentration contrast agent due to poor contrast for some reason even when a normal-concentration contrast agent is used).

[0082] The 1328 original images X and the 1328 converted images Y generated from each original image X were evaluated as follows (FIGS. 12 to 15).

[0083] 12 and 13 show the results (box plots) of calculating the Michelson contrast and RMS contrast for the original image X and the converted image Y. As can be seen from Fig. 12 and Fig. 13, it was confirmed that the converted image Y tends to have higher values ​​than the original image X for both the Michelson contrast and the RMS contrast.

[0084] We also performed a quantitative evaluation using entropy for the transformed image Y. Entropy is an index that represents the disorder of an image, and tends to take a larger value as the contrast increases.

[0085] Specifically, the image was histogrammed and normalized, and then the entropy was calculated. i log2P i (P i is the probability of occurrence of each pixel value). The results are shown in Figure 14. As can be seen from Figure 14, the entropy of the converted image Y also tends to be higher than that of the original image X. From the above, it can be said that the contrast of the converted image Y has improved.

[0086] Fig. 15 shows the results of calculating the SNR (signal-noise ratio) for the original image X and the converted image Y. As can be seen from Fig. 15, it was confirmed that the SNR also dropped significantly in the converted image Y. Note that the results of Figs. 12 to 15 can be further improved by further increasing the number of training data.

[0087] <Modification> Specific modified embodiments that can be added to the embodiments exemplified above are exemplified below. Multiple embodiments arbitrarily selected from the following examples may be combined as appropriate within the scope of not mutually contradicting each other.

[0088] (1) In the above-described embodiment, the second training image T2 of the second data D2 used as training data may be an image obtained by actually administering a low-concentration contrast agent to a subject, or may be an image generated by image processing so that the contrast between the third region R3 and the fourth region R4 is similar to that obtained by using a low-concentration contrast agent (hereinafter referred to as a "pseudo-low-concentration image").

[0089] For example, the pseudo low-density image is generated by image processing a first training image T1 captured using a normal-concentration contrast agent to generate a second training image T2. Note that the first training image T1 used to generate the pseudo low-density image may or may not be the same as the first training image T1 used to train the generative model M.

[0090] Specifically, the image processing is a process for generating, from a first training image T1, a second training image T2 in which the contrast between the third region R3 and the fourth region R4 (the density difference between the third region R3 and the fourth region R4) is lower than the contrast between the first region R1 and the second region R2 in the first training image T1 (the density difference between the first region R1 and the second region R2). In other words, a second training image T2 corresponding to a low concentration of contrast agent is generated. As preprocessing, the resolution of the first training image T1 is set to 1024 x 1024 and converted to 256 gradations (8 bits). However, the preprocessing of the first training image T1 is not limited to the above example.

[0091] The image processing generates a second training image T2 including a third region R3 obtained by changing the pixel values ​​(density) of the first region R1 according to the density value of the low density, and a fourth region R4 obtained by maintaining the pixel values ​​(density) of the second region R2. The above image processing generates a second training image T2 in which the contrast between the third region R3 and the fourth region R4 is lower than the contrast between the first region R1 and the second region R2.

[0092] The third region R3 is obtained by changing the pixel values ​​of each pixel in the first region R1 to approach the pixel values ​​of the second region R2 (i.e., to reduce the contrast between the first region R1 and the second region R2). The pixel value of the second region R2 here is, for example, the most frequent value of the second region R2. It is assumed that the area of ​​the second region R2 in the first training image T1 is significantly larger than that of the first region R1. Therefore, typically, the most frequent value in the first training image T1 becomes the most frequent value of the second region R2. However, for example, a user may visually select the pixel value of any pixel in the second region R2 as the pixel value of the second region R2.

[0093] Assuming a first training image T1 in which the first region R1 is close to black (0) and the second region R2 is close to white (255), the third region R3 of the second training image T2 will be lighter (i.e., have a larger pixel value) than the first region R1 of the first training image T1.

[0094] For example, the second training image T2 is generated by converting the pixel value of each pixel in the first training image T1 using the following formula (1). G1 is the pixel value of the pixel in the first training image T1, and G2 is the converted pixel value (i.e., the pixel value of the pixel in the second training image T2). α is a value corresponding to the low density density value. BG is a value obtained by subtracting a value (e.g., 20) corresponding to noise (e.g., noise caused by a living body or an imaging device) from the mode in the first training image T1 (i.e., the mode in the second region R2).

[0095] G2 = BG - α (BG - G1) (1)

[0096] For pixels in the first training image T1 where G1 is smaller than BG (G1 < BG) (i.e., the first region R1), perform the transformation according to Equation (1). For the first training image T1, pixels where G1 is smaller than BG typically correspond to pixels in the first region R1 (i.e., pixels representing the target tissue V). The pixel values of each pixel in the first region R1 are superimposed with pixel values caused by the background and pixel values caused by the target tissue V. The term "α(BG - G1)" in Equation (1) multiplies α after extracting the pixel values caused by the target tissue V for the pixels in the first region R1.

[0097] For example, when G1 is 50, BG is 200, and α is 0.2 (the concentration value where the low concentration is 0.2 times the normal concentration), G2 becomes 170. Note that the equation for converting the pixel values of each pixel in the first training image T1 (the first region R1) is not limited to the above Equation (1) as long as it is possible to convert the pixel values according to the concentration values of the low concentration.

[0098] The contrast between the third region R3 and the fourth region R4 in the second training image T2 can also be paraphrased as corresponding to the contrast between the region representing the target tissue and other regions in the image obtained using a low concentration contrast agent.

[0099] A plurality of second training images T2 may be generated from one first training image T1. That is, a plurality of second training images T2 with different contrasts between the third region R3 and the fourth region R4 are generated from one first training image T1. Specifically, a plurality of second training images T2 are generated from one first training image T1 by varying the concentration values of the low concentration. For example, the low concentration is set at predetermined intervals (e.g., 0.02 times) within the range of 0.25 to 0.9 times the normal concentration to generate a plurality of second training images T2.

[0100] According to the configuration that uses the pseudo low concentration image generated from the first training image T1 as the second training image T2, a large number of second training images T2 can be generated without administering a low concentration contrast agent to the subject.

[0101] If it is possible to generate second training images T2 in which the contrast between the third region R3 and the fourth region R4 is lower than the contrast between the first region R1 and the second region R2, the second training images T2 may be generated by changing both the pixel values ​​of the first region R1 and the second region R2 (for example, by changing the range of available pixel values). However, second training images T2 that include the fourth region R4 obtained by maintaining the pixel values ​​in the second region R2 can more closely resemble images obtained in actual clinical settings using a low-concentration contrast agent or images using a normal-concentration contrast agent but with reduced contrast.

[0102] (2) In the above-mentioned embodiment, the original image X includes, for example, an image that looks like it was taken using a low-concentration contrast agent (an image with insufficient contrast for accurate diagnosis) due to poor contrast caused by some reason (such as the patient's body size or insufficient flow rate of the contrast agent) even though a normal-concentration contrast agent was used. Similarly, the second training image T2 includes an image that looks like it was taken using a low-concentration contrast agent (an image with insufficient contrast for accurate diagnosis).

[0103] Furthermore, in the present disclosure, the term "low concentration contrast agent" includes not only a case where the concentration of the contrast agent to be administered is reduced by diluting the prepared contrast agent with physiological saline, but also a case where a contrast agent of a normal concentration is used and the flow rate or flow velocity of the contrast agent injection is reduced, thereby reducing the administered amount and the effective concentration in the blood vessel.

[0104] (3) In the above-described embodiment, the training image and the original image X are images captured by digital subtraction angiography (DSA images). However, the first training image T1 and the original image X are not limited to DSA images. For example, images captured using a contrast agent in various imaging techniques, such as angiography other than digital subtraction angiography, X-ray fluoroscopy, computed tomography (CT), ultrasound imaging, and magnetic resonance imaging (MRI), may be used as the first training image T1 and the original image X. The type of contrast agent administered to the subject may be changed as appropriate depending on the type of imaging technique. Furthermore, the target tissue V is not limited to blood vessels, but may be any tissue that can be enhanced by a contrast agent, and may be tubular tissue other than blood vessels (e.g., vascular organs such as lymphatic vessels, bile ducts, and pancreatic ducts).

[0105] (4) The subject to which the present invention is applicable is not limited to humans, but may be other animals.

[0106] (5) In the above-described embodiment, for example, training images or original images X obtained by using a contrast agent of a normal concentration on simulated tissue (for example, simulated blood vessels or simulated organs) may be used.

[0107] (6) The image processing system 100 is an example of a training device and an image generating device. However, the training device and the image generating device may each be configured as separate, independent devices.

[0108] (7) The present invention can also be thought of as a program that causes a computer to function as a training unit that trains a generative model to perform a generation process using as training data: a first training image, the first training image including a first region representing a target tissue, and a first prompt representing the first training image; a second training image, the second training image including a second region representing the target tissue and having a lower contrast than the first training image, and second data including the second prompt representing the second training image; and a generated original image including a third region representing the target tissue and a converted image, in accordance with a conditional prompt, including a fourth region corresponding to the third region and having a higher contrast than the original image; wherein the first prompt includes information indicating that the first training image was obtained using a normal concentration of contrast agent; the second prompt includes information indicating that the second training image was obtained using a contrast agent at a concentration lower than the normal concentration; and the conditional prompt includes information indicating a condition for the generation process, indicating that a normal concentration of contrast agent is to be used.

[0109] The present invention can also be conceived as a program that causes a computer system to function as a generation unit that executes a generation process to generate the transformed image by inputting an original image and a condition prompt into the generation model.

[0110] The program and generative model according to the present invention can be provided, for example, in a form stored on a computer-readable recording medium, or in a form distributed from a distribution device via a communication network and installed on a computer. [Explanation of symbols]

[0111] 11: Control device 12:Storage device 13:Display device 14: Input device 100: Image processing system 112: Training Department 114: Generation part D1: First data D2: Second data M: Generative model P1: First prompt P2: Second prompt R1: 1st area R2: 2nd area R3: 3rd area R4: 4th area Rx1, Ry1: target area Rx2, Ry2: Non-target area T1: First training image T2: Second training image V: Target organization X: Original image Y: Transformed image

Claims

1. first data including a first training image obtained using a contrast agent, the first training image including a first region representing a target tissue, and a first prompt related to the first training image; using as training data second data including a second training image that includes a second region representing a target tissue and has a contrast lower than that of the first training image, and second data including a second prompt related to the second training image; training a generative model to perform a generation process to generate, in response to a conditional prompt, an original image including a third region representing a target tissue, a transformed image including a fourth region corresponding to the third region, the transformed image having a higher contrast than the original image; the first prompt includes information indicating that the first training image was acquired using a normal concentration of contrast agent; the second prompt includes information indicating that the second training image was acquired using a lower than normal concentration of contrast agent; The condition prompt indicates a condition for the generation process and includes information indicating that a normal concentration of contrast agent is to be used. A method for training a computer-implemented generative model.

2. The first prompt includes information regarding the examination method used to capture the first training image.

2. The method of training a generative model of claim 1.

3. The first prompt includes information regarding the energy used to capture the first training image.

2. The method of training a generative model of claim 1.

4. The first prompt includes information representing the type of artifact.

2. The method of training a generative model of claim 1.

5. The first prompt includes information about the noise.

2. The method of training a generative model of claim 1.

6. first data including a first training image obtained using a contrast agent, the first training image including a first region representing a target tissue, and a first prompt related to the first training image; using as training data second data including a second training image that includes a second region representing a target tissue and has a contrast lower than that of the first training image, and second data including a second prompt related to the second training image; a training unit configured to train a generative model to perform a generation process to generate a transformed image having a higher contrast than the original image, the transformed image including a fourth region corresponding to the third region and a fourth region corresponding to the third region, in response to a conditional prompt, the transformed image including a fourth region corresponding to the third region and a fourth region corresponding to the third region, the fourth region being in response to a conditional prompt, the transformed image having a higher contrast than the original image; the first prompt includes information indicating that the first training image was acquired using a normal concentration of contrast agent; the second prompt includes information indicating that the second training image was acquired using a lower than normal concentration of contrast agent; The condition prompt indicates a condition for the generation process and includes information indicating that a normal concentration of contrast agent is to be used. training equipment.

7. first data including a first training image obtained using a contrast agent, the first training image including a first region representing a target tissue, and a first prompt related to the first training image; using as training data second data including a second training image that includes a second region representing a target tissue and has a contrast lower than that of the first training image, and second data including a second prompt related to the second training image; causing the computer to function as a training unit that trains the generative model to perform a generation process that generates, in response to an original image including a third region representing a target tissue and a conditional prompt, a transformed image that includes a fourth region corresponding to the third region and has a higher contrast than the original image; the first prompt includes information indicating that the first training image was acquired using a normal concentration of contrast agent; the second prompt includes information indicating that the second training image was acquired using a lower than normal concentration of contrast agent; The condition prompt indicates a condition for the generation process and includes information indicating that a normal concentration of contrast agent is to be used. program.

8. A generation process for generating the transformed image is executed by inputting the original image and the condition prompt to a generation model trained by the training method of claim 1. A computer-implemented image generation method.

9. the target tissue is a tubular tissue; In the generation process, the line structure is extracted from the original image, and the converted image is generated based on the extracted line structure and the condition prompt. The image generating method of claim 8.

10. a generation unit that executes a generation process to generate the transformed image by inputting the original image and the condition prompt to a generative model trained by the training method of claim 1; Image generating device.

11. a generation unit that executes a generation process for generating the transformed image by inputting the original image and the condition prompt to a generation model trained by the training method of claim 1; A program that makes a computer system function as a

Citation Information

Patent Citations

  • Deep learning-based contrast agent dose reduction in medical imaging

    JP2020536638A

  • Medical Image Enhancement

    JP2022547047A