Image processing device, method, and program
The iterative offset-based method in the image processing device enhances landmark detection accuracy by refining target points from candidate points, addressing the limitations of existing methods in deriving vertebral body corners.
Patent Information
- Application Number
- JP2024011914
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-30
- Publication Date
- 2025-08-12
AI Technical Summary
Existing methods, such as those described in Non-Patent Document 1, struggle to accurately derive landmarks like the corners of vertebral bodies from the centers using offset regression in medical images.
An image processing device and method that iteratively derive target points by acquiring offsets from candidate points, using them as new reference points, until a predetermined condition is satisfied, enhancing accuracy.
This iterative approach allows for the precise detection of landmarks, particularly at corners of vertebral bodies, improving accuracy even in deformed structures.
Smart Images

Figure 2025117188000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to an image processing device, a method, and a program. [Background technology]
[0002] Various methods have been proposed for detecting points representing predefined characteristic structures, i.e., landmarks, from an image (see, for example, Patent Document 1). Landmarks have also been detected using learning models trained through deep learning. Methods such as heat map regression, coordinate point regression, and offset regression are often used for landmark detection using deep learning. In heat map regression, the likelihood of a landmark is given as a correct answer, and deep learning is performed to obtain a detection result that indicates the likelihood of a landmark on the image using a color change or the like. In coordinate point regression, coordinate points representing landmarks in the image are given as correct answers, and deep learning is performed to detect the positions of the landmarks. In offset regression, vectors (offsets) from each pixel included in the image to the landmark are given as correct answers, and deep learning is performed to detect the offsets from each pixel included in the image to the landmark.
[0003] Meanwhile, a method for detecting landmarks using both heat map regression and offset regression has been proposed. For example, Non-Patent Document 1 proposes a method for detecting the centers and four corners of vertebrae that make up the spine as landmarks, in which the centers of the vertebrae are derived by heat map regression and the four corners of the vertebrae are derived by offset regression from the centers of the vertebrae. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Special Publication No. 2022-532039 [Non-patent literature]
[0005] [Non-Patent Document 1] VERTEBRA-FOCUSED LANDMARK DETECTION FOR SCOLIOSIS ASSESSMENT, Yi et al., 9 Jan 2020 Summary of the Invention [Problem to be solved by the invention]
[0006] However, with the method described in Non-Patent Document 1, for example, when a vertebral body is included in the image, it may not be possible to accurately derive landmarks such as the four corners of the vertebral body by offset regression from the center of the vertebral body.
[0007] The present disclosure has been made in consideration of the above circumstances, and aims to enable landmarks of structures included in images to be derived with greater accuracy. [Means for solving the problem]
[0008] An image processing device according to the present disclosure includes at least one processor, The processor performs a first process of acquiring an offset from a reference point of a structure included in the image to a candidate point of a target point related to the reference point; The target point is derived by repeating the second process of obtaining a new offset to a candidate point for a new target point, using the candidate point for the target point derived based on the offset as a new reference point, N times (N≧1) until a predetermined condition is satisfied.
[0009] The image processing method according to the present disclosure includes a first process in which a computer acquires an offset from a reference point of a structure included in an image to a candidate point of a target point related to the reference point; The target point is derived by repeating the second process of obtaining a new offset to a candidate point for a new target point, using the candidate point for the target point derived based on the offset as a new reference point, N times (N≧1) until a predetermined condition is satisfied.
[0010] The image processing program according to the present disclosure includes a procedure for performing a first process for acquiring an offset from a reference point of a structure included in an image to a candidate point of a target point related to the reference point; The computer is caused to execute a second process of obtaining a new offset to a candidate point for a new target point, with the candidate point for the target point derived based on the offset as a new reference point, and repeating this process N times (N≧1) until a predetermined condition is satisfied, thereby deriving the target point. [Effects of the Invention]
[0011] According to the present disclosure, landmarks of structures included in an image can be derived with high accuracy. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 1 is a perspective view showing an overview of a medical information system to which an image processing device according to an embodiment of the present disclosure is applied; [Figure 2] FIG. 1 is a diagram illustrating a hardware configuration of an image processing device according to an embodiment of the present disclosure. [Figure 3] FIG. 1 is a diagram illustrating a functional configuration of an image processing apparatus according to an embodiment of the present disclosure. [Figure 4] Diagram for explaining reference points [Figure 5] Diagram to explain the learning of the derived model [Figure 6] Diagram to explain weights during learning [Figure 7] Diagram showing derived target points [Figure 8] Figure showing further target point derivation [Figure 9] FIG. 1 shows the results of deriving reference points and target points in a medical image. [Figure 10] Diagram showing the imaging plane of the intervertebral disc [Figure 11] A flowchart showing the processing performed in this embodiment [Figure 12] A medical image showing a vertebral body with a compression fracture [Figure 13]Diagram to explain Cobb angle [Figure 14] Diagram showing cross-sectional images of the breast DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, an embodiment of the present disclosure will be described with reference to the drawings. First, the configuration of a medical information system to which an image processing device according to this embodiment is applied will be described. FIG. 1 is a diagram showing a schematic configuration of the medical information system. In the medical information system shown in FIG. 1, a computer 1 incorporating an image processing device according to this embodiment, an imaging device 2, and an image storage server 3 are connected in a communicable state via a network 4.
[0014] The computer 1 includes an image processing device according to this embodiment and has the image processing program of this embodiment installed. The computer 1 may be a workstation or personal computer operated directly by a doctor making a diagnosis, or may be a server computer connected to either of these via a network. The image processing program is stored in a storage device of the server computer connected to the network or in network storage in an externally accessible state, and is downloaded and installed into the computer 1 used by the doctor upon request. Alternatively, the program may be recorded on a recording medium such as a DVD (Digital Versatile Disc) or CD-ROM (Compact Disc Read Only Memory) and distributed, and then installed into the computer 1 from the recording medium.
[0015] The imaging device 2 is a device that captures an image of a diagnostic target region of a subject to generate a two-dimensional or three-dimensional image representing the region, and specifically includes a radiation imaging device, a CT (Computed Tomography) device, an MRI (Magnetic Resonance Imaging) device, a PET (Positron Emission Tomography) device, etc. The image of the subject generated by the imaging device 2 is transmitted to and stored in the image storage server 3. Note that a three-dimensional image includes multiple tomographic images, or an image composed of three-dimensional coordinates generated from multiple tomographic images.
[0016] The image storage server 3 is a computer that stores and manages various data, and is equipped with a large-capacity external storage device and database management software. The image storage server 3 communicates with other devices via a wired or wireless network 4, sending and receiving image data and the like. Specifically, it acquires various data, including image data of images generated by the imaging device 2, via the network, and stores and manages the data on a recording medium such as a large-capacity external storage device. The storage format of the image data and communication between devices via the network 4 are based on protocols such as DICOM (Digital Imaging and Communication in Medicine).
[0017] Next, an image processing device according to this embodiment will be described. Fig. 2 is a diagram showing the hardware configuration of the image processing device according to this embodiment. As shown in Fig. 2, the image processing device 20 includes a CPU (Central Processing Unit) 11, a display 14, an input device 15, a memory 16, and a network I / F (Interface) 17 connected to a network 4. The CPU 11, the display 14, the input device 15, the memory 16, and the network I / F 17 are connected to a bus 19. The CPU 11 is an example of a processor in the present disclosure.
[0018] The memory 16 includes the storage unit 13 and a RAM (Random Access Memory) 18. The RAM 18 is a memory for primary storage, and is, for example, a RAM such as an SRAM (Static Random Access Memory) or a DRAM (Dynamic Random Access Memory).
[0019] The storage unit 13 is a non-volatile memory, and is realized by, for example, at least one of a hard disk drive (HDD), a solid state drive (SSD), an electrically erasable and programmable read-only memory (EEPROM), and a flash memory. The storage unit 13, which serves as a storage medium, stores an image processing program 12 according to this embodiment. The CPU 11 reads the image processing program 12 from the storage unit 13, loads it into the RAM 18, and executes the loaded image processing program 12. The storage unit 13 also stores a center derivation model 22A and a derivation model 23A, which will be described later.
[0020] The display 14 is a device for displaying various screens, such as a liquid crystal display or an EL (Electro Luminescence) display. The input device 15 is a device for a user to input, such as at least one of a keyboard, a mouse, a microphone for voice input, a touchpad for proximity input including contact, and a camera for gesture input. The network I / F 17 is an interface for connecting to the network 4.
[0021] Next, the functional configuration of the image processing device according to this embodiment will be described. Fig. 3 is a diagram showing the functional configuration of the image processing device according to this embodiment. As shown in Fig. 3, the image processing device 20 includes an image acquisition unit 21, a reference point derivation unit 22, a first processing unit 23, a second processing unit 24, and a display control unit 25. When the CPU 11 executes the image processing program 12, the CPU 11 functions as the image acquisition unit 21, the reference point derivation unit 22, the first processing unit 23, the second processing unit 24, and the display control unit 25.
[0022] The image acquisition unit 21 acquires a medical image G0 to be processed from the image storage server 3 in response to an instruction from the operator via the input device 15. In this embodiment, the medical image G0 is an MRI image of a sagittal section including the spinal column of a human body. The MRI image may be a three-dimensional image or a two-dimensional image. The spinal column is made up of multiple vertebrae. The vertebrae are made up of vertebral bodies, spinous processes, etc.
[0023] The reference point derivation unit 22 analyzes the medical image G0 to derive the centers of the vertebrae contained in the medical image G0 as reference points. The centers of the vertebrae may be, for example, the centers of gravity of the vertebrae. For this purpose, the reference point derivation unit 22 uses a center derivation model 22A that has been trained to derive the centers of the vertebrae. The center derivation model 22A is stored in the storage unit 13. A vertebra is an example of a structure disclosed herein.
[0024] The center derivation model 22A is constructed by deep learning a neural network using image data representing an image of a vertebral body whose center has been identified, so that when an image including the vertebral body is input, the center of the vertebral body is identified. The deep learning may use coordinate point regression, which uses the coordinate position of the vertebral body center as ground truth data and learns to output the coordinate position of the vertebral body center in response to input image data. Alternatively, heat map regression may be used, which uses the likelihood of the vertebral body center position as ground truth data and learns to output the likelihood of the position where the vertebral body center is located as a heat map in response to input image data. Alternatively, offset regression may be used, which uses a vector (offset) from each pixel included in the image toward the center of the vertebral body as ground truth data and learns to output the offset from each pixel of the image represented by the image data to the center of the vertebral body in response to input image data.
[0025] As a result, the centers of the vertebrae are derived as reference points B1 to B7 for each of the multiple (seven in this case) vertebrae included in the medical image G0, as shown in Fig. 4. Note that Fig. 4 shows an image of one sagittal cross section included in the medical image G0.
[0026] The first processing unit 23 performs a first process to acquire offsets from each of the reference points B1 to B7 derived by the reference point derivation unit 22 to candidate points for the target points associated with each of the reference points B1 to B7. In this embodiment, both the reference point and the target point are located within the vertebral body. As described above, the reference point is the center of the vertebral body, and the target point is a point on the boundary of a corner of the vertebral body. In this embodiment, the target point associated with the reference point refers to a target point located in the same vertebral body as the reference point. The target point is not limited to a point on the boundary of a corner. For example, the target point may be a point located near the boundary within the vertebral body, approximately a few pixels away from a point on the boundary of the corner of the vertebral body. Note that the medical image G0 is a three-dimensional image, and the vertebral body is roughly a rectangular parallelepiped, so there are eight corners in the vertebral body.
[0027] To obtain the offset, the first processing unit 23 uses a derived model 23A that has been trained to derive an offset to the target point. The derived model 23A is also stored in the storage unit 13.
[0028] The derived model 23A is constructed by deep learning a neural network so that when an image of a vertebral body with identified reference points is input, the neural network outputs a target point. Deep learning uses vectors (offsets) from each pixel in the image to the target point, which is the corner of the vertebral body, as ground truth data, and outputs vectors from each pixel of the image represented by the image data to the target point, i.e., offsets, from the neural network in response to input image data. Note that the number of output channels of the neural network can be, for example, the product of the number of dimensions and the number of target points. For example, if the image data is a three-dimensional image, there are eight corners of the vertebral body, so the number of output channels is 3 × 8 = 24. If the image data is a two-dimensional image, there are four corners of the vertebral body, so the number of output channels is 2 × 4 = 8.
[0029] FIG. 5 is a diagram for explaining the learning of the derived model 23A. Note that although the medical image G0 is a three-dimensional image, FIG. 5 illustrates the explanation using a two-dimensional image. Also, FIG. 5 illustrates the learning of the offset using the point (vertex) xG at the upper right corner of the vertebral body 30 as the target point. During learning, the target point predicted based on the offset at a certain pixel in the image does not match the correct target point xG. Therefore, the target point predicted from a certain pixel xi derived during learning is designated xi′. Note that the coordinates of pixel xi, target point xG, and predicted target point xi′ are three-dimensional in the case of a three-dimensional image, and two-dimensional in the case of a two-dimensional image. Also, pixel xj shown in FIG. 5 is a pixel that is farther away from target point xG than pixel xi.
[0030] In this case, the offset loss Ltotal is derived by the following equations (1) to (3). pred_offseti is the predicted offset from pixel xi to target point xi', gt_offseti is the offset (correct offset) of the correct data from pixel xi to vertex xG shown in Figure 5, and wi and wj are weights used when deriving the loss Loffset for points xi and xj, respectively. In equations (2) to (4), an L1 norm or the like may be used instead of the L2 norm.
number
[0031] When the derived model 23A is trained, a deviation occurs between the target point predicted from each pixel using the offset derived during training and the target point given as the correct answer. Therefore, when training the derived model 23A, the deviation thus derived may be used as a loss. This loss is the loss Ldst shown in the above equation (3). This can further improve the accuracy when deriving the target point.
[0032] Note that only Loffset may be used as the loss when training the derived model 23A. In this case, Ltotal=Loffset.
[0033] In this embodiment, the derived model 23A is constructed by repeating learning until a predetermined condition is satisfied. The predetermined condition may be, but is not limited to, that the loss Ltotal is equal to or less than a predetermined threshold value or that learning has been completed a predetermined number of times.
[0034] Note that during learning, the closer a pixel is to the target point, the greater the weighting may be assigned when deriving the loss. For example, with respect to the target point xG, as shown in FIG. 6, it is preferable to set concentric regions A100, A101, and A102 centered around the target point xG, and assign increasing weightings to the regions A102, A101, and A100 in that order when deriving the loss. In this case, the weights wi and wj are set as shown in the following equation (4). In this way, by performing offset regression with a greater weighting on the loss for pixels closer to the target point, the accuracy of deriving the offset to the target point for pixels near the target point can be further improved.
number
[0035] In addition, when training the derived model 23A, the loss may be calculated only in a predetermined range around the reference point and the target point, rather than the entire image used for training. In other words, the loss may not be calculated outside the predetermined ranges around the reference point and the target point. For example, the loss may not be calculated in the range of the target point area A102 shown in FIG. 6. Similarly, for the reference point, the loss may not be calculated in an area larger than a certain range based on the reference point. This reduces the amount of calculation during training.
[0036] The first processing unit 23 uses the derived model 23A to derive an offset V1 from the reference point to a target point related to the reference point. Meanwhile, in a vertebral body, there is a relatively long distance between its center and a corner. Therefore, when the offset V1 from the reference point to the target point is derived, as shown in FIG. 7, the target point CK11 derived based on the offset V1 may not coincide with the actual target point C11 (i.e., a point at the corner of the vertebral body). Therefore, in this embodiment, the target point derived based on the offset V1 is referred to as a candidate point CK11 for the target point.
[0037] In this embodiment, the second processing unit 24 performs a second process using the derived model 23A to acquire a new offset V2 to the target point C11, using the candidate point CK11 of the target point derived based on the offset V1 as a new reference point.
[0038] Here, candidate point CK11 of the target point derived based on offset V1 is closer to target point C11 than reference point B1. Therefore, when a new offset V2 to the target point is derived using candidate point CK11 as a new reference point, candidate point CK12 of the target point derived based on the new offset V2 will be closer to target point C11 or will coincide with the target point, as shown in FIG.
[0039] In this embodiment, the second processing unit 24 repeats the second process N times (N≧1) until a predetermined condition is satisfied. That is, because the offset is a vector, the second processing unit 24 repeats the second process until the absolute value of the offset from the target point candidate point CK12 to the target point C11, which is derived based on the new offset V2, becomes less than a predetermined threshold value Th1. Note that if the absolute value of the offset from the candidate point CK12 to the target point C11 becomes less than the threshold value Th1, the second processing unit 24 ends the process after only one iteration of the second process.
[0040] On the other hand, there may be cases where the offset from the target point candidate point CK12 derived based on the new offset V2 to the target point C11 is equal to or greater than a predetermined threshold value Th1. In this case, the second processing unit 24 repeats the second process using the target point candidate point CK12 derived based on the new offset V2 as the reference point until the offset from the target point candidate point CK12 derived based on the new offset V2 to the target point becomes less than the threshold value Th1. This allows the second processing unit 24 to derive the corner of the vertebral body as the target point. Figure 9 is a diagram showing the derivation results of the reference point and the target point in the medical image G0. In Figure 9, the reference point is indicated by a white circle and the target point is indicated by a black circle. Furthermore, since the image shown in Figure 9 is a two-dimensional image, the number of target points is four.
[0041] The reference points and target points derived in this way are used in the following processes. For example, the reference points are used to detect the center line of the spine. For example, if the medical image G0 is a scout image used for positioning when imaging using a CT device or an MRI device, the target points are used to draw a line (or a plane in the case of a three-dimensional image) between opposing target points on adjacent vertebrae as shown in Figure 10 and to derive an imaging plane for imaging the intervertebral disc during imaging.
[0042] The condition for the second processing unit 24 to repeat the second process is not limited to the absolute value of the offset from the candidate point CK12 of the target point derived based on the new offset V2 to the target point being less than the threshold value Th1, but may be that N reaches a predetermined number of times.
[0043] Furthermore, when the second processing unit 24 performs the second processing, learning becomes difficult if the target point to be derived is located far from the reference point. Therefore, when the second processing unit 24 performs the second processing, the threshold value Th1 may be increased as the target point to be derived is farther from the reference point. That is, different threshold values Th1 may be set for multiple anatomical structures. The threshold value may also be derived based on the distance between the reference value associated with the anatomical structure and the target value. Specifically, suppose there is a distance L31 between a reference point B31 for a certain vertebra and a derived target point CK31, and a distance L32 between a reference point B32 for another vertebra and a derived target point CK32. If the distance L31 is relatively shorter than the distance L32, the threshold value Th1 used to derive the target point CK31 may be set relatively smaller than the threshold value Th2 used to derive the target point CK32.
[0044] Next, the processing performed in this embodiment will be described. Fig. 11 is a flowchart showing the processing performed in this embodiment. First, the image acquisition unit 21 acquires a medical image G0 for deriving a reference point and a target point from the image storage server 3 (step ST1).
[0045] Next, the reference point derivation unit 22 derives a reference point from the medical image G0 (step ST2). That is, the center of each of the multiple vertebrae included in the medical image G0 is derived as a reference point. Next, the first processing unit 23 performs a first process using the derived model 23A to acquire an offset from the reference point to a target point related to the reference point (step ST3).
[0046] Next, the second processing unit 24 performs a second process using the derived model 23A to acquire a new offset to the target point using the candidate point for the target point derived based on the offset as a new reference point (step ST4).
[0047] Next, the second processing unit 24 determines whether or not a predetermined condition is satisfied (step ST5). If step ST5 is negative, the second processing unit 24 returns to step ST4 and repeats the processes of steps ST4 and ST5. If step ST5 is positive, the medical image G0 in which the derived reference points and target points are drawn is displayed on the display 14 (step ST6), and the process ends. Note that instead of displaying the medical image G0, the medical image G0 may be processed using the reference points and target points as described above.
[0048] In this embodiment, a first process is performed to acquire an offset from a reference point to a target point, and a second process is performed to acquire a new offset to the target point using a candidate point for the target point derived based on the offset as a new reference point. This process is repeated N times (N≧1) until a predetermined condition is satisfied. By repeating the second process in this manner, the new reference point approaches the target point, thereby improving the accuracy of deriving the target point. Therefore, as described in Non-Patent Document 1, the target point can be derived more accurately than a target point derived solely by using the offset acquired by the first process.
[0049] In particular, points at the corners of vertebral bodies are difficult to detect because their characteristics are unclear compared to the center of the vertebral body. Furthermore, points at the corners of vertebral bodies are located far from the center of the vertebral body, which serves as the reference point. According to this embodiment, the second process is repeatedly performed, so that target points whose characteristics are relatively unclear compared to the reference point or that are located far from the reference point can be accurately detected.
[0050] Furthermore, the shape of the vertebral body may be deformed due to a disease such as a compression fracture, for example, as shown in Fig. 12. Even if the shape of the vertebral body is deformed from a healthy state in this way, in this embodiment, the second process is repeated, so that the corner of the deformed vertebral body can be accurately derived as the target point.
[0051] Furthermore, as described in Aubert et al.'s paper (Automatic spine and pelvis detection in frontal X-rays using deep neural networks for patch displacement learning, Benjamin Aubert et al., 2016 IEEE 13th International Symposium on Biomedical Imaging (ISBI), April 13, 2016), a method for detecting landmarks has been proposed, in which patches are set on an image, offsets are derived from the set patches, and the process of moving the patches based on the derived offsets is repeated. In this embodiment, offsets are not derived each time a patch is moved, but offsets for all pixels are derived through a single inference process. Therefore, compared to the method described in Aubert et al.'s paper, the number of inference processes is reduced, resulting in a smaller amount of calculation, allowing the position of the target point to be derived quickly.
[0052] In the above embodiment, the medical image G0 is an MRI image of a sagittal plane including the spine, but this is not limited to this. An MRI image of a coronal plane including the spine may also be used as the medical image G0. In this case, as shown in FIG. 13, reference points and target points are acquired on the coronal plane, i.e., on each vertebra constituting the spine as viewed from the front of the subject. Note that FIG. 13 shows a coronal image of a patient with scoliosis. The reference points derived on the coronal plane are used to detect the centerline of the spine, etc. The target points are used to derive the Cobb angle used to diagnose scoliosis.
[0053] In this case, a line can be drawn between opposing target points on adjacent vertebral bodies, and the angle α formed by the intersection of the lines can be derived as the Cobb angle, as shown in Figure 13. Furthermore, if the coronal section medical image G0 is a scout image used for positioning when imaging using a CT device or an MRI device, a line can be drawn between opposing target points on adjacent vertebral bodies and used to derive the imaging plane for imaging the intervertebral disc during imaging.
[0054] Furthermore, in the above embodiment, the medical image G0 including the spine is the processing target, but this is not limited thereto. Medical images G0 including anatomical structures such as the breast, brain, and heart can also be processed. For example, if the medical image G0 is a tomographic image representing one of the cross sections of an MRI image of the breast, the center of the breast may be derived as the reference point B40, and the nipple, upper end, and lower end of the breast may be derived as target points C41, C42, and C43, respectively, as shown in FIG. 14. Furthermore, if the medical image G0 is a CT image of the brain, the center of the brain may be derived as the reference point, and the center of the orbit and the center of the external auditory canal may be derived as target points for deriving the orbitomeatal line (OM line). Furthermore, if the medical image is an MRI or CT image of the heart, the center of the heart may be derived as the reference point, and the connection points between the heart and large blood vessels, such as the aorta and vena cava, may be derived as target points.
[0055] In the above embodiment, the center derivation model 22A derives the reference point, and the derivation model 23A derives the offset, but this is not limiting. It is also possible to use only one derivation model that derives both the reference point and the offset. In this case, the derivation model performs a common process for deriving the reference point and the offset in the first stage, and then branches off in the second stage depending on the task of deriving the reference point and the task of deriving the offset.
[0056] In the above embodiment, the image processing device according to this embodiment is provided with the reference point derivation unit 22, but this is not limited to this. If the medical image G0 from which the reference points have been derived in advance is stored in the image storage server 3, it is possible to configure the image processing device according to this embodiment without providing the reference point derivation unit 22.
[0057] Furthermore, in the above embodiment, image processing is performed on medical images, but this is not limited to this, and the techniques of the present disclosure can also be applied when deriving reference points and target points for general photographic images other than medical images.
[0058] Furthermore, in the above embodiment, the following various processors can be used as the hardware structure of processing units that perform various processes, such as the image acquisition unit 21, the reference point derivation unit 22, the first processing unit 23, the second processing unit 24, and the display control unit 25. As described above, the various processors include a CPU, which is a general-purpose processor that executes software (programs) and functions as various processing units, as well as dedicated electrical circuits that are processors having a circuit configuration specifically designed to perform specific processes, such as a GPU (Graphics Processing Unit), a Programmable Logic Device (PLD), a processor whose circuit configuration can be changed after manufacture, such as an FPGA (Field Programmable Gate Array), and an ASIC (Application Specific Integrated Circuit).
[0059] A single processing unit may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs or a combination of a CPU and an FPGA). Also, multiple processing units may be configured with a single processor.
[0060] Examples of configuring multiple processing units with a single processor include, first, a form in which one processor is configured with a combination of one or more CPUs and software, and this processor functions as multiple processing units, as typified by computers such as client and server. Second, a form in which a processor is used to realize the functions of an entire system including multiple processing units with a single IC (Integrated Circuit) chip, as typified by systems on chips (SoCs). In this way, various processing units are configured using one or more of the above-mentioned various processors as a hardware structure.
[0061] Furthermore, more specifically, the hardware structure of these various processors can be an electric circuit that combines circuit elements such as semiconductor elements.
[0062] The following are appendices to the present disclosure. (Additional note 1) at least one processor; The processor: performing a first process of acquiring an offset from a reference point of a structure included in an image to a candidate point of a target point related to the reference point; An image processing device that derives the target point by repeating a second process of obtaining a new offset to a new candidate point for the target point, with the candidate point for the target point derived based on the offset as the new reference point, N times (N≧1) until a predetermined condition is satisfied. (Additional note 2) 2. The image processing device according to claim 1, wherein the processor derives the reference points by analyzing the image. (Additional note 3) 3. The image processing device according to claim 2, wherein the processor derives the reference points using a derivation model that has been trained to derive the reference points from the image. (Additional note 4) 4. The image processing device according to claim 3, wherein the learning is learning by heat map regression, coordinate point regression, or offset regression. (Additional note 5) 5. The image processing device according to any one of claims 1 to 4, wherein the predetermined condition is that N has reached a predetermined number of times or that the absolute value of the new offset has become less than a predetermined threshold value. (Additional note 6) 6. The image processing device according to any one of claims 1 to 5, wherein the reference point and the target point associated with the reference point are located within the same structure. (Additional note 7) 7. The image processing device according to claim 6, wherein the target point has relatively vaguer features than the reference point. (Additional note 8) 8. The image processing device according to claim 6, wherein the reference point is inside the structure and the target point is on the boundary of the structure. (Additional note 9) 9. The image processing device according to any one of claims 1 to 8, wherein the processor performs the first process and the second process using a derived model learned by offset regression. (Additional note 10) 10. The image processing device according to claim 9, wherein the learning is learning in which the weight for offset loss increases as the object is closer to the target point. (Additional note 11) 11. The image processing device according to claim 9, wherein the learning is learning that derives offset loss only in a predetermined range around the reference point and a predetermined range around the target point. (Additional note 12) The image processing device according to any one of appendix 9 to 11, wherein the learning is learning using a deviation between the position of the candidate point of the target point repeatedly derived during learning and the correct position of the target point as a further loss. (Additional note 13) a computer performs a first process of acquiring an offset from a reference point of a structure included in an image to a candidate point of a target point related to the reference point; An image processing method for deriving the target point by repeating a second process, in which a candidate point for the target point derived based on the offset is used as the new reference point, to obtain a new offset to the new candidate point for the target point N times (N≧1) until a predetermined condition is satisfied. (Additional note 14) a step of performing a first process of acquiring an offset from a reference point of a structure included in an image to a candidate point of a target point related to the reference point; and a procedure for deriving the target point by repeating a second process of acquiring a new offset to a new candidate point for the target point, with the candidate point for the target point derived based on the offset as the new reference point, N times (N≧1) until a predetermined condition is satisfied. [Explanation of symbols]
[0063] 1. Computer 2. Imaging equipment 3. Image storage server 4 Network 11 CPU 12 Image Processing Programs 13 Storage section 14 Display 15 Input Devices 16 memory 17 Network I / F 18 RAM 19 Bus 20 Image processing device 21 Image acquisition unit 22 Reference point derivation part 22A Center Derived Model 23 First Processing Section 23A Derived Model 24 Second Processing Section 30 vertebral bodies A100~A102 area B1~B7, B40 reference point C11, C41~C42 target point CK11, CK12 candidate points G0 Medical Imaging V1,V2 offset α Cobb angle
Claims
1. at least one processor; The processor: performing a first process of acquiring an offset from a reference point of a structure included in an image to a candidate point of a target point related to the reference point; An image processing device that derives the target point by repeating a second process, in which a candidate point for the target point derived based on the offset is used as the new reference point, to obtain a new offset to the new candidate point for the target point N times (N≧1) until a predetermined condition is satisfied.
2. The image processing device according to claim 1 , wherein the processor derives the reference points by analyzing the image.
3. The image processing device according to claim 2 , wherein the processor derives the reference points using a derivation model that has been trained to derive the reference points from the image.
4. The image processing device according to claim 3 , wherein the learning is learning by heat map regression, coordinate point regression, or offset regression.
5. 2. The image processing device according to claim 1, wherein the predetermined condition is that N has reached a predetermined number of times or that the absolute value of the new offset has become less than a predetermined threshold value.
6. The image processing device according to claim 1 , wherein the reference point and the target point associated with the reference point are located within the same structure.
7. The image processing device according to claim 6 , wherein the target points have relatively vaguer features than the reference points.
8. 8. The image processing device according to claim 6, wherein the reference point is located inside the structure, and the target point is located on the boundary of the structure.
9. The image processing device according to claim 1 , wherein the processor performs the first processing and the second processing using a derived model trained by offset regression.
10. The image processing device according to claim 9 , wherein the learning is performed by increasing the weighting of the offset loss as the offset approaches the target point.
11. The image processing device according to claim 9 or 10, wherein the learning is learning that derives offset losses only in a predetermined range around the reference point and a predetermined range around the target point.
12. 11. The image processing device according to claim 9, wherein the learning is performed using a deviation between the position of the candidate point of the target point repeatedly derived during learning and the correct position of the target point as a further loss.
13. a computer performs a first process of acquiring an offset from a reference point of a structure included in an image to a candidate point of a target point related to the reference point; An image processing method for deriving the target point by repeating a second process, in which a candidate point for the target point derived based on the offset is used as the new reference point, to obtain a new offset to the new candidate point for the target point N times (N≧1) until a predetermined condition is satisfied.
14. a step of performing a first process of acquiring an offset from a reference point of a structure included in an image to a candidate point of a target point related to the reference point; and a procedure for deriving the target point by repeating a second process of obtaining a new offset to a new candidate point for the target point, with the candidate point for the target point derived based on the offset as the new reference point, N times (N≧1) until a predetermined condition is satisfied.
Citation Information
Patent Citations
Convolutional Neural Network Based Landmark Tracker
JP2022532039A