Image processing device, image processing method, and program

By generating nearby cross-sectional images through out-of-plane translation and rotation of reference cross-sections, the technology addresses the insufficient variation in training data for anatomical landmark estimation, enhancing inference performance and accuracy in medical image analysis.

JP7837185B2Active Publication Date: 2026-03-30CANON KK +1
View PDF 8 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-16
Publication Date
2026-03-30

AI Technical Summary

Technical Problem

Conventional data augmentation techniques for medical images fail to achieve sufficient inference performance due to insufficient variation in training data for estimating anatomical landmarks using machine learning.

Method used

The technology involves acquiring three-dimensional images and generating nearby cross-sectional images through out-of-plane translation and rotation of reference cross-sections, using these images along with anatomical landmark positions to create training data for constructing a learning model that estimates landmark positions on cross-sectional images.

Benefits of technology

This approach enhances the inference performance in estimating anatomical landmarks by increasing the variety of training data, leading to improved accuracy in landmark detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007837185000001
    Figure 0007837185000001
  • Figure 0007837185000002
    Figure 0007837185000002
  • Figure 0007837185000003
    Figure 0007837185000003
Patent Text Reader

Abstract

To provide a technology to improve inference performance when executing inference of an anatomical landmark (feature point) position in a medical image on the basis of machine learning.SOLUTION: An image processing device in the present disclosure includes: a data acquisition unit for acquiring data including a three-dimensional image obtained by capturing a subject, reference cross section parameters indicating a predetermined reference cross section in the three-dimensional image, and an anatomical landmark position in the three-dimensional image; a nearby cross sectional image acquisition unit for acquiring a nearby cross sectional image for stipulating a nearby cross section of the predetermined reference cross section on the basis of the reference cross section parameters; a position acquisition unit for acquiring the anatomical landmark position in the nearby cross sectional image; and a learning unit for acquiring a learning model for estimating the anatomical landmark position on a cross sectional image from the cross sectional image created by using information including the combination of the nearby cross sectional image and the anatomical landmark position acquired by the position acquisition unit as learning data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology disclosed herein relates to an image processing apparatus, an image processing method, and a program.

Background Art

[0002] In medical image diagnosis, the position of a predetermined anatomical landmark (feature point) in a cross-sectional image showing a two-dimensional cross-section of a subject is estimated, and based on the estimation result, the shape of the subject is measured or various diagnostic indices are calculated. In recent years, with the development of machine learning technologies such as deep learning, when a large amount of learning data with sufficient variations can be used, the anatomical landmark position can be estimated with high accuracy. However, unlike the processing of images for which a large amount of learning data can be collected, medical images generally have the problem that there is little learning data available for the variations of the subject. For this reason, techniques for augmenting (data augmentation) learning data have been proposed.

[0003] In Patent Document 1, a learning image is determined in advance, and while interactively acquiring an image using an ultrasonic probe, an image having a similarity of a certain level or more with the learning image is acquired as an augmented learning image. Further, in Patent Document 2, data augmentation of learning data is performed by rotating a template image for detecting nodules of a subject.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, even if the cross-sectional images for training are augmented using the conventional techniques described above, sufficient inference performance may still not be obtained due to insufficient variation in the training data.

[0006] In light of the above, the technology disclosed herein aims to provide a technology that improves the inference performance when estimating the location of anatomical landmarks (feature points) in medical images based on machine learning. [Means for solving the problem]

[0007] The image processing device relating to the technology disclosed herein is A data acquisition unit that acquires data including a three-dimensional image of a subject, a reference cross-sectional parameter representing a predetermined reference cross-section in the three-dimensional image, and the position of an anatomical landmark in the three-dimensional image. Based on the aforementioned reference cross-sectional parameters, This is obtained by moving the aforementioned reference cross-section outside the aforementioned reference cross-section. The cross section near the predetermined reference cross section Based on the nearby cross-sectional parameters and the three-dimensional image, A nearby cross-sectional image acquisition unit that acquires nearby cross-sectional images, Based on the location of the anatomical landmark and the relationship between the predetermined reference cross section and the nearby cross section, the location of the anatomical landmark in the image of the predetermined reference cross section is approximated on the nearby cross section. A position acquisition unit that acquires the position of anatomical landmarks in the aforementioned nearby cross-sectional image, Using information including the pair of the aforementioned nearby cross-sectional image and the anatomical landmark position acquired by the position acquisition unit as training data, construction A learning unit that acquires a learning model for estimating the position of anatomical landmarks on a cross-sectional image from the cross-sectional image, The present invention includes an image processing apparatus characterized by comprising:

[0008] Furthermore, the image processing device related to the technology disclosed herein is An input data acquisition unit that acquires an input 3D image and an input reference cross-section parameter representing a predetermined reference cross-section in the input 3D image, An input cross-sectional image acquisition unit acquires an input cross-sectional image based on the input three-dimensional image and the input reference cross-sectional parameters, A learning model acquisition unit that acquires a learning model for estimating the position of anatomical landmarks on a cross-sectional image from the cross-sectional image, An estimation unit that estimates the position of anatomical landmarks in the input cross-sectional image using the learning model acquired by the learning model acquisition unit, Equipped with, The aforementioned learning model, A 3D image of the subject, A reference cross-section parameter representing a predetermined reference cross-section in the aforementioned three-dimensional image, The location of anatomical landmarks in the aforementioned three-dimensional image, Based on the aforementioned reference cross-section parameters, the reference cross-section is obtained by moving it outside the reference cross-section. The cross section near the predetermined reference cross section Based on the nearby cross-sectional parameters and the 3D image, the data was obtained. A nearby cross-sectional image based on the aforementioned reference cross-sectional parameters, Based on the location of the anatomical landmark and the relationship between the predetermined reference cross-section and the nearby cross-section, the location of the anatomical landmark in the image of the predetermined reference cross-section was approximated on the nearby cross-section. The position of the anatomical landmark in the aforementioned cross-sectional image and The pair of the nearby cross-sectional image obtained based on and the anatomical landmark location is used as training data. construction This is a learning model. The present invention includes an image processing apparatus characterized by the following features.

[0009] Furthermore, the image processing device related to the technology disclosed herein is A data acquisition unit that acquires data including a three-dimensional image of a subject, a reference cross-sectional parameter representing a predetermined reference cross-section in the three-dimensional image, and the position of an anatomical landmark in the three-dimensional image. By rotational movement of the predetermined reference cross section around an axis passing through at least one of the anatomical landmark locations on the predetermined reference cross section, This is obtained by moving the aforementioned reference cross-section outside the aforementioned reference cross-section. The cross section near the predetermined reference cross section Based on the nearby cross-sectional parameters and the three-dimensional image, A nearby cross-sectional image acquisition unit that acquires nearby cross-sectional images, Based on the location of the anatomical landmark and the relationship between the predetermined reference cross-section and the nearby cross-section, the location of the anatomical landmark in the image of the predetermined reference cross-section was approximated on the nearby cross-section. The aforementioned nearby cross-sectional image in The location of the aforementioned anatomical landmark and The aforementioned nearby cross-sectional image Using information including sets as training data construction A cross-sectional image was created. a learning unit that obtains a learning model for estimating the positions of anatomical landmarks on the cross-sectional image An image processing apparatus characterized by including the above components is provided.

[0010] Furthermore, the present disclosure can also be regarded as an image processing method including at least a part of the above processing, a program for causing a computer to execute these methods, or a computer-readable recording medium on which such a program is non-temporarily recorded. Each of the above configurations and processes can be combined with each other to constitute the present invention as long as no technical contradiction occurs.

Advantages of the Invention

[0011] According to the technology of the present disclosure, it becomes possible to improve the inference performance when estimating the positions of anatomical landmarks (feature points) in medical images based on machine learning.

Brief Description of the Drawings

[0012] [Figure 1] A diagram showing a schematic configuration of an image processing apparatus according to an embodiment. [Figure 2] A flowchart of a process executed by an image processing apparatus according to an embodiment [Figure 3] A flowchart of a learning model acquisition process executed by an image processing apparatus according to an embodiment [Figure 4] A diagram schematically showing a reference cross-section and anatomical landmarks in an embodiment [Figure 5] A diagram schematically showing an approximation process of anatomical landmarks in an embodiment [Figure 6] A diagram showing a schematic configuration of an image processing apparatus according to an embodiment [Figure 7] A flowchart of a process executed by an image processing apparatus according to an embodiment [Figure 8] A flowchart of a process executed by an image processing apparatus according to an embodiment

Embodiments for Carrying Out the Invention

[0013] A preferred embodiment of the technology disclosed herein will be described below with reference to the drawings. Note that the drawings are provided solely for the purpose of illustrating the structure or configuration, and the dimensions of the illustrated components do not necessarily reflect actual dimensions. Furthermore, the same reference numerals are used for identical components or elements in each drawing, and redundant explanations will be omitted below.

[0014] The image processing apparatus according to the embodiment described below provides a function for estimating the position of a predetermined anatomical landmark of an object on a predetermined reference cross-section for observing the object in an input three-dimensional image. The input image to be processed is a medical image, i.e., an image of a subject (such as the human body) taken or generated for the purpose of medical diagnosis, examination, research, etc., and is typically an image acquired by an imaging system called a modality. For example, ultrasound images obtained by an ultrasound diagnostic device, X-ray CT images obtained by an X-ray CT device, and MRI images obtained by an MRI device may be the images to be processed. Anatomical landmarks are characteristic points of the object to be observed in the image. The image processing apparatus performs data augmentation (data enrichment) of training data suitable for estimating the position of anatomical landmarks on a reference cross-section of a three-dimensional image, and generates a learning model by machine learning.

[0015] The following explanation will detail a specific example of an image processing device, using a transesophageal 3D ultrasound image of the mitral valve region of the heart as the input 3D image, and estimating the anatomical landmarks of the mitral valve region, which is the target of observation. Note that the mitral valve is just one example; other parts of the body may also be the target of observation.

[0016] <First Embodiment> The image processing device according to the first embodiment estimates the positions of anatomical landmarks defined on a reference cross-section from an input 3D image, which is an input image, and a 2D reference cross-section defined on the image. The image processing device also generates a learning model for estimating the positions of anatomical landmarks from training data.

[0017] Here, the reference cross-section is a two-dimensional cross-section suitable for observing the target object in a three-dimensional image of the subject. In this embodiment, it is a cross-section (hereinafter referred to as "plane A") that allows simultaneous observation of the mitral valve and aortic valve in the heart. In other words, the reference cross-section is a cross-section that includes the anatomical landmark to be observed. Plane A is an example of a predetermined reference cross-section. Note that the reference cross-section may be any other plane depending on the purpose of observing the target object.

[0018] In this embodiment, as an example, the anatomical landmarks are four points, "Ao," "A," "P," and "Nadir," which are included in "plane A." Figures 4A to 4D schematically show plane A set in the 3D image and the anatomical landmarks defined on plane A. In Figure 4A, plane A is a predetermined cross section (402a) cut out from a 3D image (401) of the mitral valve region. Figure 4B shows a 2D cross-sectional image of plane A 402a. In Figure 4B, anatomical landmark 404 is Ao, anatomical landmark 405 is A, anatomical landmark 406 is Nadir, and anatomical landmark 407 is P. Anatomical landmark Ao is the position where plane A and the aortic valve annulus intersect, and anatomical landmarks A and P are the positions where plane A and the mitral valve annulus intersect. Also, anatomical landmark Nadir is the lowest point of the mitral valve annulus on plane A. Note that the anatomical landmarks used to estimate the position may be other anatomical landmarks depending on the purpose of observation.

[0019] In this embodiment, when augmenting training data, inter-observer variability (inter-observer variability) is used. If the displacement (perturbation) is of a similar magnitude to that of the reference cross-section (rver variability), then an image of a nearby cross-section (neighboring cross-section) displaced out of the plane of the reference cross-section can also be considered as a training image. Based on this, the image processing device of this embodiment first applies out-of-plane translation and / or rotation within a predetermined range to the reference cross-section to obtain an image of a nearby cross-section different from the reference cross-section (neighboring cross-section image). Here, out-of-plane translation and rotation refers to a coordinate transformation in which at least one point on the reference cross-section changes to a position outside the reference cross-section due to the translation and rotation. A typical example is translation in the direction of the normal of the reference cross-section, in which case all points that were originally on the reference cross-section move to a position outside the reference cross-section. Next, the positions of anatomical landmarks on the neighboring cross-section image are obtained based on the positions of anatomical landmarks defined on the reference cross-section image. Specifically, the anatomical landmarks defined on the reference cross-section image are projected onto the neighboring cross-section image. In this way, under the constraint that the displacement (perturbation) from the reference cross-section is within a predetermined range, a large number of pairs of cross-sectional images and anatomical landmarks can be obtained as training data. The image processing device of this embodiment performs machine learning using the training data augmented in this way to generate a training model. Furthermore, this embodiment can also be applied to estimating the position of any point on the reference cross-section, in addition to the four points mentioned above.

[0020] In this embodiment, it is assumed that the anatomical landmark lies on the reference cross-section, but the three-dimensional position of the anatomical landmark on the input three-dimensional image does not necessarily have to lie on the reference cross-section. That is, the shortest distance between the anatomical landmark and the cross-section does not necessarily have to be zero, as long as the difference is small enough that it can be considered to be substantially on the cross-section. Even in such cases, if the approximate position of the anatomical landmark on the reference cross-sectional image can be defined, the processing of this embodiment can be executed without any problems.

[0021] The configuration and processing of the image processing apparatus of this embodiment will be described below with reference to Figure 1. Figure 1 is a block diagram showing an example configuration of an image processing system (also called a medical image processing system) including the image processing apparatus of this embodiment. The image processing system 1 comprises an image processing apparatus 10 and a database 22. The image processing apparatus 10 is connected to the database 22 via a network 21 in a manner that allows communication. The network 21 includes, for example, a LAN (Local Area Network) or a WAN (Wide Area Network).

[0022] Database 22 holds and manages multiple images and information used in the processing described below. The information managed in database 22 includes input images (images to be processed) used for anatomical landmark estimation processing by the image processing device 10, and training data for generating a learning model. Alternatively, the information managed in database 22 may include information about the learning model generated from the training data instead of the learning data. The learning model information may be stored in the internal storage (ROM 32 or storage unit 34) of the image processing device 10 instead of database 22. The image processing device 10 can retrieve data held in database 22 via the network 21.

[0023] The image processing device 10 includes a communication interface 31, a read-only memory 32, a random access memory 33, a storage unit 34, an operation unit 35, a display unit 36, and a control unit 37.

[0024] The communication IF31 is a communication unit consisting of a LAN card or the like, which enables communication between an external device (e.g., database 22) and the image processing device 10. The ROM32 is composed of non-volatile memory or the like, and stores various programs and data. The RAM33 is composed of volatile memory or the like, and is used as work memory to temporarily store programs and data that are currently running. The storage unit 34 is composed of an HDD (Hard Disk Drive) or the like, and stores various programs and data. The operation unit 35 is a key It consists of a board, mouse, touch panel, etc., and inputs instructions from the user (for example, a doctor or medical technologist) into various devices.

[0025] The display unit 36 ​​is composed of a display or the like and displays various information to the user. The control unit 37 is composed of a CPU (Central Processing Unit) or the like and provides overall control of the processing in the image processing device 10. Functionally, the control unit 37 includes an inference unit 41, a learning model acquisition unit 42, and a display processing unit 51. The inference unit 41 consists of an input data acquisition unit 43, an input cross-sectional image generation unit 44, and a landmark estimation unit 45. The learning model acquisition unit 42 includes a learning data acquisition unit 46, an augmented parameter calculation unit 47, a learning cross-sectional image generation unit 48, a landmark approximation unit 49, and a learning unit 50. The control unit 37 may also include a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), or an FPGA (Field-Programmable Gate Array).

[0026] The learning process may be performed by a separate device from the image processing device 10, and the learning model acquisition unit 42 of this device may be configured to acquire the learning model stored in the database 22 or the storage unit 34. Alternatively, the inference unit may be a separate device, and the model acquisition unit may simply generate the learning model and store it in the storage unit 34 or the database 22.

[0027] The input data acquisition unit 43 acquires a pair of the input image to be processed (an image in which the position of anatomical landmarks is unknown) and the reference cross-sectional parameters in the input image from the database 22. The input image is an image of a subject acquired by various modalities, and in this embodiment, it is a three-dimensional ultrasound image of the heart. The reference cross-sectional parameters are a vector of a total of six parameters: three parameters for the center position and three normal vector parameters, and in this embodiment, these are parameters representing plane A. The input image may be acquired directly from the modality. In this case, the image processing device 10 may be implemented in the console of the modality (imaging system). The reference cross-sectional parameters may be obtained as a result of manual setting by the user, or they may be automatically estimated based on the input image.

[0028] The input cross-sectional image generation unit 44 uses the input image acquired by the input data acquisition unit 43 and the reference cross-sectional parameters to generate a two-dimensional cross-sectional image by cutting out the input image using the reference cross-sectional parameters.

[0029] The landmark estimation unit 45 uses the learned model obtained by the learned model acquisition unit 42 to estimate the coordinate values ​​of anatomical landmarks from the two-dimensional cross-sectional image obtained by the input cross-sectional image generation unit 44.

[0030] The training data acquisition unit 46 acquires training data from the database 22. Each sample constituting the training data consists of pixel value information of a 3D image for training, reference cross-sectional parameters, and coordinate value information of anatomical landmarks. Here, the coordinate value information of anatomical landmarks is assumed to be expressed as 3D coordinate values, and as mentioned above, the anatomical landmarks consist of four points called Ao, A, P, and Nadir. It is assumed that all of these anatomical landmarks lie on the cross-section expressed by the reference cross-sectional parameters, and their coordinates are transformed into 2D coordinates on the cross-section and stored in the RAM 33.

[0031] The augmented parameter calculation unit 47 calculates cross-sectional parameters for multiple neighboring cross-sections for each learning sample, based on the learning data acquired by the learning data acquisition unit 46, under the constraint that each learning sample is both near and outside the reference cross-section. In other words, the augmented parameter calculation unit 47 adds translation and rotation within a predetermined range to the reference cross-sectional parameters, without including in-plane movement or rotation of the reference cross-section, to create new parameters (augmented parameters, neighboring cross-sectional parameters). Calculate the meter.

[0032] The learning cross-sectional image generation unit 48 generates a two-dimensional cross-sectional image using each learning sample acquired by the learning data acquisition unit 42 and the augmentation parameters calculated by the augmentation parameter calculation unit 47. That is, for each three-dimensional image contained in each learning sample, "number of augmentation parameters + 1" two-dimensional cross-sectional images are generated. The reason for the one extra image compared to the number of augmentation parameters is to include the reference cross-section itself in the learning sample. The cross-sectional image generation process by the learning cross-sectional image generation unit 48 is the same as the cross-sectional image generation process executed by the input cross-sectional image generation unit 44, except that the input data is learning data.

[0033] The landmark approximation unit 49 approximates the position of the anatomical landmark in the cross-section represented by the augmented parameters, using the anatomical landmark acquired by the training data acquisition unit 46 and the augmented parameters calculated by the augmented parameter calculation unit 47. The anatomical landmark whose approximate position has been calculated in this way is then transformed into a two-dimensional coordinate on the cross-section. Similarly, the anatomical landmarks of each training sample are also transformed into two-dimensional coordinates on the reference cross-section using the same method.

[0034] The learning unit 50 uses the two-dimensional cross-sectional images generated by the learning cross-sectional image generation unit 48 and the two-dimensional coordinates of anatomical landmarks acquired by the landmark approximation unit 49 to learn (construct) a learning model that estimates the two-dimensional coordinates of anatomical landmarks from the two-dimensional cross-sectional images.

[0035] Based on the results calculated by the landmark estimation unit 45, the display processing unit 51 displays the input image and the estimated anatomical landmarks in the image display area of ​​the display unit 36 ​​in a display format that can be easily viewed by the user of the image processing device 10.

[0036] Each component of the image processing device 10 functions according to a computer program. For example, the control unit 37 (CPU) reads and executes a computer program stored in the ROM 32 or storage unit 34, etc., using RAM 33 as the work area, thereby realizing the function of each component. Some or all of the functions of the components of the image processing device 10 may be realized by using dedicated circuits. In addition, some of the functions of the components of the control unit 37 may be realized by using a cloud computer. For example, a computing device located in a different location from the image processing device 10 may be connected to the image processing device 10 via a network 21 so as to be communicative, and the functions of the components of the image processing device 10 or the control unit 37 may be realized by the image processing device 10 and the computing device sending and receiving data.

[0037] Next, an example of the processing of the image processing device 10 in Figure 1 will be explained using flowcharts in Figures 2 and 3.

[0038] (Step S201: Obtaining input data) In step S201, the user instructs the system to acquire an image via the operation unit 35. The input data acquisition unit 43 then acquires a pair of input 3D images specified by the user and input reference cross-section parameters representing reference cross-sections for the input 3D images from the database 22 as input data and stores it in the RAM 33.

[0039] Any known method may be used to specify the input 3D image and input reference cross-section parameters. For example, the user may specify an image and manually set the input reference cross-section parameters for the specified image. Alternatively, the input reference cross-section parameters may be automatically estimated for the image specified by the user based on predetermined criteria. Furthermore, the system may automatically select an image that satisfies predetermined conditions from a group of images specified by the user and automatically estimate the reference cross-section parameters for the selected image, or it may accept input of a reference cross-section by the user. In addition to obtaining input data from database 22, input images may also be obtained from the ultrasound images acquired moment by moment by the ultrasound imaging device, and input reference cross-sectional parameters may be set for the obtained input images. In this case as well, as with obtaining data from database 22, both manual setting and automatic estimation of input reference cross-sectional parameters are applicable.

[0040] Alternatively, the system may be configured to acquire a two-dimensional cross-sectional image as the input image to be processed. In this case, it is not necessary to acquire the input reference cross-sectional parameters, and the processing in step S203 does not need to be performed.

[0041] (Step S202: Obtaining the trained model) In step S202, the learning model acquisition unit 42 functions as a learning data acquisition unit 46, an augmented parameter calculation unit 47, a learning cross-sectional image generation unit 48, a landmark approximation unit 49, and a learning unit 50, respectively, to construct a learning model.

[0042] The process of this step will be described in detail below using the flowchart in Figure 3.

[0043] (Step S2021: Acquisition of training data) In step S2021, the training data acquisition unit 46 acquires the 3D image, reference cross-sectional parameters, and coordinate values ​​of anatomical landmarks for each training sample from the database 22 and stores them in the RAM 33. As shown in Figure 4B, in this embodiment, four points, Ao, A, P, and Nadir, are used as anatomical landmarks.

[0044] Furthermore, to enhance the robustness of the learning model, the training data should ideally consist of images taken of multiple different patients. However, it is acceptable to include images of the same patient taken at different times.

[0045] (Step S2022: Calculation of augmented parameters) In step S2022, the augmented parameter calculation unit 47 calculates multiple cross-sectional parameters (augmented parameters) based on the reference cross-sectional parameters acquired by the learning data acquisition unit 46, under the constraint that the cross-sectional parameters are in the vicinity of the reference cross-sectional parameters but outside of the reference cross-sectional parameters. That is, the augmented parameter calculation unit 47 calculates new parameters that define the nearby cross-sectional parameters by adding translation and rotation within a predetermined range to the reference cross-sectional parameters, without including in-plane movement or rotation of the reference cross-sectional parameters.

[0046] In this embodiment, "translation along the normal vector" is defined as parallel movement, and "rotation around the major axis" is defined as rotation. These two operations are used to translate and rotate the reference cross-section parameters and increase their value. Here, translation along the normal vector is the operation of translating the reference cross-section in the direction normal to plane A (Figure 4B) (perpendicular to the plane of the paper). Rotation around the major axis is the operation of rotating plane A around a predetermined axis extending on the reference cross-section, in this case the major axis of plane A (403a in Figure 4A) as the central axis. In this embodiment, a movement range of ±5 mm is set for translation along the normal vector, and a rotation range of ±5 degrees is set for rotation around the major axis. By randomly sampling within these ranges a predetermined number of times (for example, 20 times), a number of augmented parameters corresponding to the number of samples can be obtained.

[0047] Furthermore, instead of randomly sampling the augmentation parameters, they may be sampled by moving them in predetermined step sizes within the set movement and rotation ranges. Alternatively, to avoid bias due to random sampling, a constraint may be imposed that the distance between the center positions of the augmentation parameters must be greater than or equal to a predetermined value. In this way, by making at least one of the translation and rotation movements a movement at predetermined intervals, a less biased augmentation can be achieved. You can obtain the parameters.

[0048] Furthermore, augmentation may be performed not by rotation around the major axis as exemplified, but by rotation around other axes such as the minor axis. Also, augmentation may be performed by rotation around at least two or more axes passing through a predetermined reference cross-section (for example, one augmentation may be performed by rotation around the minor axis, and another by rotation around the major axis). In addition, instead of fixed axes such as the major and minor axes, the rotation axis may be dynamically changed according to the landmark position for each training sample. For example, the rotation axis may be set to pass through a predetermined landmark on the reference cross-section. For example, the rotation may be performed around an axis connecting landmark A (point 405 in Figure 4B) and landmark P (point 407 in Figure 4B). Note that the rotation axis may be set to pass through at least one landmark on the predetermined reference cross-section. In this way, the positional relationship of landmarks can be taken into consideration when defining the rotation axis. Furthermore, since the position of the landmark on the rotation axis is the same in the reference cross-section and the augmented neighboring cross-sectional image, the positions of these landmarks can be used in the augmented cross-sectional image as well as the original landmark positions. In other words, there is no need to approximate the positions of the landmarks. This reduces the load associated with data augmentation and allows for more accurate identification of landmark locations in nearby cross-sectional images.

[0049] In addition, the data augmentation methods are not limited to the methods described above; they may also include other known augmentation techniques commonly used in image processing, such as translation, rotation, scaling, and affine transformation within a reference cross-section. Alternatively, augmentation may be performed by modifying the pixel values ​​of the image itself, such as by increasing or decreasing the overall pixel values ​​through luminance conversion, contrast leveling, or noise addition. By doing so, the amount of training data obtained as a result of augmentation will increase further, and it is expected that the performance of the training model obtained in the subsequent training process will be further improved.

[0050] (Step S2023: Generating 2D cross-sectional images of training samples) In step S2023, the learning cross-sectional image generation unit 48 functions as a neighboring cross-sectional image acquisition unit that acquires neighboring cross-sectional images that define the cross-sections in the vicinity of a predetermined reference cross-section based on reference cross-sectional parameters. The learning cross-sectional image generation unit 48 uses the 3D image of the learning sample acquired in step S2021 and the augmented parameters calculated in step S2022 to generate a 2D cross-sectional image (neighboring cross-sectional image) by cutting out the neighboring cross-section represented by the augmented parameters from the 3D image.

[0051] As explained in step S2022, multiple augmentation parameters (e.g., 20) are calculated for each training sample. In this step, two-dimensional cross-sectional images are generated using the original reference cross-sectional parameters and the augmentation parameters. That is, if there are 20 augmentation parameters, 21 two-dimensional cross-sectional images are generated in total, including the original reference cross-sectional parameters. Therefore, if the number of training samples is, for example, 60 cases, this step will generate 21 × 60 = 1260 two-dimensional cross-sectional images.

[0052] The process of generating a 2D cross-sectional image by applying cross-sectional parameters to a 3D image can be performed using known image processing techniques. Specifically, the cross-sectional parameters are used to calculate which pixel (voxel) in the original 3D image corresponds to each pixel in the output 2D cross-sectional image, and the pixel values ​​of the voxels are assigned to the pixels in the 2D cross-sectional image. When assigning pixel values, any known interpolation method such as nearest neighbor interpolation, linear interpolation, or bicubic interpolation can be used.

[0053] (Step S2024: Approximation of anatomical landmarks) In step S2024, the landmark approximation area 49 is the anatomical reference area in the nearby cross-sectional image. It functions as a position acquisition unit that obtains the location of landmarks. Specifically, the landmark approximation unit 49 uses the anatomical landmarks acquired by the training data acquisition unit 46 and the augmented parameters calculated by the augmented parameter calculation unit 47 to approximate the location of the anatomical landmarks in the nearby cross-section represented by the augmented parameters. The anatomical landmarks for which the approximate location has been calculated are transformed into 2D coordinates on the nearby cross-section. In addition, in this step, the anatomical landmarks acquired by the training data acquisition unit 46 are also transformed into 2D coordinates on the reference cross-section.

[0054] In this embodiment, the location of an anatomical landmark in a nearby cross-sectional image is determined by projecting the location of the anatomical landmark on the reference cross-section. Specifically, the anatomical landmark on the reference cross-section is projected perpendicularly to the nearby cross-section, and its position is considered to be the location of the anatomical landmark in the nearby cross-section. That is, the coordinates of the foot of the perpendicular line drawn from the anatomical landmark on the reference cross-section to the nearby cross-section are taken as the approximate location of the anatomical landmark in the nearby cross-section.

[0055] Figures 5A to 5C show schematic diagrams of the approximation process for anatomical landmarks. In Figure 5A, plane 502a represents plane A, and axis 503a represents the major axis on plane A. In Figure 5B, plane 505a represents a nearby cross-section obtained by rotating plane A by a predetermined angle around the major axis. Axis 503b is the major axis and corresponds to axis 503a in Figure 5A.

[0056] The approximation process for anatomical landmarks will be explained with reference to Figure 5C. Image 502b is a cross-sectional image of plane A, and there are four anatomical landmarks 510a, 511a, 512a, and 513a on the cross-section. When this cross-sectional image of plane A 502b is viewed from above (viewed in the direction indicated by arrow 514 in Figure 5A), plane A is represented as a single line segment 502c. The anatomical landmarks 510a, 511a, 512a, and 513a are then represented as points 510b, 511b, 512b, and 513b on line segment 502c, respectively. From this viewpoint, the nearby cross-section 505a is represented as a single line segment 505b that is inclined with respect to plane A 502c. In this case, the projection of anatomical landmarks corresponds to the process of dropping points on plane A 502c perpendicularly onto the nearby cross section 505b, as shown by the dotted arrows extending from points 510b, 511b, 512b, and 513b, respectively.

[0057] Furthermore, the approximate position of anatomical landmarks can be obtained using methods other than the projection described above. For example, the approximate position may be set so that the 2D coordinates on the reference cross section and the 2D coordinates on the nearby cross section are the same. This approximate position can be calculated by transforming the coordinates of the anatomical landmark using the translation and rotation translation set when calculating the augmentation parameters. That is, for example, if the augmentation parameters are obtained by rotating by 2.5 degrees, the approximate position of the anatomical landmark can also be set to a position rotated by 2.5 degrees. In the case of simple projection, if the rotation angle around the major axis becomes large, the projected position approaches the rotation axis, which may result in the approximate position being set to a position that deviates from the appropriate landmark position in the nearby cross section. By applying the same parameters as the augmentation parameters, the distance between the rotation axis and the anatomical landmark is maintained, thus avoiding this problem.

[0058] (Step S2025: Building the Learning Model) In step S2025, the learning unit 50 acquires a learning model that estimates the location of anatomical landmarks on a cross-sectional image from the cross-sectional image, which is created using information including pairs of nearby cross-sectional images and the locations of anatomical landmarks as training data. Specifically, the learning unit 50 learns a learning model that estimates the 2D coordinates of anatomical landmarks from a 2D cross-sectional image using pairs of 2D cross-sectional images consisting of a reference cross-sectional image and nearby cross-sectional images, and 2D coordinate values ​​of anatomical landmarks defined on the 2D cross-sectional image. Here, the 2D cross-sectional images of all training samples, including augmented parameters, calculated in step S2023, and the step Input the 2D coordinates of anatomical landmarks whose approximate positions were calculated using S2024. Any known technique can be used for training, such as Principal Component Analysis (PCA) or convolutional neural networks (CNNs) like VGG16 or ResNet.

[0059] As described above, the process from step S2021 to step S2025 is executed to obtain the learning model (step S202).

[0060] (Step S203: Generation of a 2D cross-sectional image of the input image) In step S203, the input cross-sectional image generation unit 44 functions as an input cross-sectional image acquisition unit and uses the input 3D image acquired in step S201 and the input reference cross-sectional parameters to generate a 2D cross-sectional image by cutting out the 3D image using the reference cross-sectional parameters. The process of applying the reference cross-sectional parameters to the 3D image to generate a 2D cross-sectional image is the same as the process described in step S2023.

[0061] (Step S204: Estimation of anatomical landmark coordinates) In step S204, the landmark estimation unit 45 uses the learned model obtained in step S202 to estimate the coordinate values ​​of anatomical landmarks from the cross-sectional image obtained in step S203. As described above, the learned model estimates the two-dimensional coordinate values ​​of anatomical landmarks in the cross-sectional image. Therefore, the two-dimensional coordinate values ​​are converted to three-dimensional coordinate values ​​using the reference cross-sectional parameters obtained by the input data acquisition unit 41. In this way, estimated values ​​of anatomical landmark coordinate values ​​can be obtained for input images where the coordinate values ​​of anatomical landmarks are unknown.

[0062] (Step S205: Displaying the estimated results) In step S205, the display processing unit 51 displays the input image and estimated anatomical landmarks within the image display area of ​​the display unit 36 ​​in a display format that allows for easy viewing, based on the results calculated in step S204.

[0063] Furthermore, if the purpose is analysis or measurement based on the location of anatomical landmarks, the display processing in step S205 is not mandatory, and the system may simply store the estimated anatomical landmark information in a storage device.

[0064] According to this embodiment, in machine learning for detecting anatomical landmarks from 2D reference cross-sectional images, images of cross-sections outside the plane of the reference cross-section are extracted from the 3D image, and landmark locations are projected onto them to augment the data. This effectively increases the amount of training data, thereby improving the performance of the learning model.

[0065] (Extreme Variation 1-1) The following describes modifications of the above embodiments. In the first embodiment, an example was shown in which the processing target was a three-dimensional ultrasound image of the cardiac mitral valve region. However, the technology disclosed herein can also be implemented when using images of organs other than the heart or images obtained using other modalities.

[0066] Examples of applications to images other than the cardiac mitral valve region or images other than ultrasound include detecting specific structures such as the diencephalon and corpus callosum from brain MRI images. When detecting these structures, the axial section passing through a position where the structure can be visually observed is the reference section, and the point cloud that depicts the central position and contour of these structures becomes the anatomical landmark. The position of the anatomical landmark is then estimated by processing similar to that of the first embodiment.

[0067] Thus, this modification allows the technology disclosed in this document to be applied to images from modalities other than 3D ultrasound images.

[0068] (Variations 1-2) In the first embodiment, the example shown was that the two-dimensional cross-sectional image extracted from the three-dimensional image by the processing in steps 203 and S2023 was a plane in the space of the three-dimensional image. However, the technology disclosed herein can also be implemented when extracting images other than planes.

[0069] Examples of extracting a 2D cross-sectional image from an image other than a plane include using a curved cross-section or a free-form surface defined by a spline function. In this case, extracting a 2D cross-sectional image from a 3D image can be done using known techniques. The curved cross-section or free-form surface is then represented as a 2D image unfolded on a 2D plane. Even if the reference cross-section is such a 2D unfolded image, the calculation of the augmented parameters performed in step S2022 can be done in the same way as in the first embodiment, although a specific method of calculation may be used. For example, if the reference cross-section is set as a curved surface that follows a predetermined structural surface of the subject, the calculation of the augmented parameters may be performed under the constraint of "following a predetermined structural surface".

[0070] Thus, according to this modified version, the reference cross-section used to augment the training data is not necessarily limited to being a plane. This allows the technology disclosed hereto be applied even when the reference cross-section is not a plane in the original 3D image, such as when observing coronary arteries by unfolding them into a 2D image.

[0071] <Second Embodiment> Next, a second embodiment will be described. In the following description, components and processes similar to those in the first embodiment will be denoted by the same reference numerals, and detailed explanations will be omitted. The image processing apparatus according to the second embodiment is, like the first embodiment, a device that estimates anatomical landmarks defined on a reference cross-section from an input three-dimensional image and a two-dimensional reference cross-section defined on the image. In the first embodiment, there was only one reference cross-section (plane A only). However, the technology disclosed hereon can also be applied when there are two or more reference cross-sections. In this embodiment, similar to the first embodiment, a three-dimensional ultrasound image of the mitral valve region will be used as the subject, and an example will be given in which, in addition to plane A, plane B, which is a cross-section orthogonal to plane A, is used as the reference cross-section.

[0072] Schematic diagrams of plane B are shown in Figures 4C and 4D. As shown in Figure 4C, plane B is a cross-section that intersects plane A (Figure 4A) at its long axis 403b and is perpendicular to plane A. Also, as shown in Figure 4D, there are two anatomical landmarks on plane B, and these landmarks are located where plane B intersects the mitral valve annulus. In the diagrams, point 409 represents point AL and point 410 represents point PM.

[0073] Next, the configuration and processing of the image processing system 2 of this embodiment will be explained using Figure 6. Compared with the device configuration diagram of the first embodiment shown in Figure 1, the image processing device 20 of the image processing system 2 has an additional landmark assignment unit 69. Also, because the number of reference cross-sections has increased from one to two, the processing content of each other processing unit has been changed from the first embodiment. The processing content of each processing unit will be explained below.

[0074] The processing of the input data acquisition unit 63 and the learning data acquisition unit 66 is the same as that of the input data acquisition unit 43 and the learning data acquisition unit 46 in the first embodiment, except that the reference cross-sectional parameters acquired are for two cross-sectional surfaces, A and B.

[0075] The processing of the input cross-sectional image generation unit 64 is the same as that of the input cross-sectional image generation unit 44 in the first embodiment. However, it is possible to distinguish whether the generated cross-sectional image is plane A or plane B. The difference lies in the fact that labels are assigned and these labels are stored together with the cross-sectional images.

[0076] The processing of the landmark estimation unit 65 is the same as that of the landmark estimation unit 45 in the first embodiment. However, this embodiment differs from the first embodiment in that there are two reference cross-sections, plane A and plane B. Therefore, as the learning model is constructed independently for plane A and plane B in the learning model acquisition unit 62, the estimation of anatomical landmarks is also performed independently for each plane.

[0077] The processing of the augmented parameter calculation unit 67 is the same as that of the augmented parameter calculation unit 47 in the first embodiment. The only difference is that augmented parameters are calculated for two cross-sections, A-surface and B-surface.

[0078] The processing of the learning cross-sectional image generation unit 68 is the same as the processing of the learning cross-sectional image generation unit 48 in the first embodiment. However, it differs in that a label is assigned to the generated cross-sectional image that can identify whether it is a nearby cross-section augmented from plane A or a nearby cross-section augmented from plane B, and the label and cross-sectional image are stored as a pair.

[0079] The landmark assignment unit 69 identifies whether the anatomical landmarks acquired by the learning data acquisition unit 66 are located on reference cross-section A or B, and assigns either the A-plane or B-plane label to each landmark.

[0080] The processing of the landmark approximation unit 70 is the same as the processing of the landmark approximation unit 49 in the first embodiment. However, it differs in that, by referring to the assignment results of the landmark assignment unit 69, anatomical landmarks assigned to plane A are projected onto an augmented nearby cross-section from plane A, and anatomical landmarks assigned to plane B are projected onto an augmented nearby cross-section from plane B.

[0081] The processing of the learning unit 71 is the same as the processing of the learning unit 50 in the first embodiment. However, in the first embodiment there was only one reference cross-section, surface A, whereas in this embodiment there are two reference cross-sections, surface A and surface B. In this embodiment, the learning of the two cross-sections is performed independently of each other.

[0082] The processing of the display processing unit 72 is the same as the processing of the display processing unit 51 in the first embodiment. However, it differs from the first embodiment in that there are two reference cross-sections, A-plane and B-plane, and the display processing is performed according to each of the A-plane and B-plane.

[0083] Next, an example of the processing of the image processing device 20 in Figure 6 will be explained using flowcharts in Figures 7 and 8.

[0084] (Step S701: Obtaining input data) The process in step S701 is the same as the process in step S201 in the first embodiment, except that the acquired reference cross-section parameters represent two reference cross-sections, A-plane and B-plane.

[0085] (Step S702: Acquisition of the learning model (for multiple cross-sections)) In step S702, the learning model acquisition unit 62 functions as a learning data acquisition unit 66, an augmented parameter calculation unit 67, a learning cross-sectional image generation unit 68, a landmark assignment unit 69, a landmark approximation unit 70, and a learning unit 71, respectively. The learning model acquisition unit 62 then constructs a learning model from the learning data acquired from the database 22.

[0086] The acquisition of the learning model is performed for each of the two reference cross-sections, A and B. The details of this step will be explained below using the flowchart in Figure 8.

[0087] (Step S7021: Acquisition of training data) The process in step S7021 is the same as the process in step S2021 in the first embodiment, except that the acquired reference cross-section parameters represent two reference cross-sections, A-plane and B-plane.

[0088] (Step S7022: Calculation of augmented parameters (for multiple cross-sections)) The process in step S7022 is the same as the process in step S2022 in the first embodiment, except that augmented parameters are calculated for two reference cross-sections, surface A and surface B. The augmented parameters in this step are assigned labels to identify whether they were augmented from surface A or surface B, and are stored in pairs with each augmented parameter.

[0089] (Step S7023: Landmark Assignment) In step S7023, the landmark assignment unit 69 identifies whether the anatomical landmarks acquired by the learning data acquisition unit 66 are located on reference plane A or plane B, and assigns either the A-plane or B-plane label to each landmark. Note that the assignment destination for a single anatomical landmark (e.g., point Ao) must be the same plane (A-plane or B-plane) across all cases in the learning data.

[0090] As shown in Figures 4B and 4D, in this embodiment, the assignment of each anatomical landmark to each reference plane is predetermined. Specifically, landmarks Ao (point 404), A (point 405), Nadir (point 406), and P (point 407) are assigned to plane A. Landmarks PM (point 409) and AL (point 410) are assigned to plane B. In this step, labels are assigned by applying the above assignments as they are.

[0091] Furthermore, this step can be performed even if it is not predetermined which reference plane each anatomical landmark is assigned to. In this case, the assignment cost to each reference plane is calculated for each training data set, and the reference plane with the minimum average cost across all cases is selected as the destination for the anatomical landmark. Here, the assignment cost can be, for example, the distance between each anatomical landmark and its respective reference plane. Alternatively, the contrast of the image around the projection position when the anatomical landmark is projected onto each reference plane can be used as the assignment cost. In this case, the higher the contrast (i.e., the more likely it is to be near an edge), the lower the assignment cost is determined to be. In this way, the processing of this embodiment can be performed even if the training data does not contain information on which anatomical landmark is assigned to which plane.

[0092] (Step S7024: Generation of 2D cross-sectional images of training samples (multiple cross-sections)) The process in step S7024 is the same as the process in step S2024 in the first embodiment, except that a label is added to the generated two-dimensional cross-sectional image to identify whether it is a nearby cross-section originating from plane A or plane B.

[0093] (Step S7025: Approximation of anatomical landmarks (multiple cross-sections)) The process in step S7025 is the same as the process in step S2024 of the first embodiment, except that the assignment results calculated in step S7023 are referenced and each anatomical landmark is projected onto an augmented neighboring cross section from the reference cross section to its respective assignment destination. Anatomical landmarks assigned to plane A are projected onto an augmented nearby section from plane A, and anatomical landmarks assigned to plane B are projected onto an augmented nearby section from plane B.

[0094] (Step S7026: Building a learning model (for multiple sections)) The process in step S7026 is the same as the process in step S7025 in the first embodiment, except that there are two reference cross-sections, surface A and surface B.

[0095] In this step, the learning of planes A and B is performed independently. Specifically, the learning model for plane A is constructed using reference cross-sections, nearby cross-sections, and anatomical landmarks that are labeled as "plane A" as training data. Then, the learning model for plane B is constructed using reference cross-sections, nearby cross-sections, and anatomical landmarks that are labeled as "plane B" as training data.

[0096] Furthermore, instead of constructing the learning models for surface A and surface B completely independently, it is possible to share part of the learning process. For example, multi-task learning in a convolutional neural network (CNN) can be used. In this case, the weights are shared up to the convolutional layer of the CNN, and then fully connected layers for estimating anatomical landmarks for surface A and fully connected layers for estimating anatomical landmarks for surface B are connected independently. Alternatively, it is also possible to define surface A and surface B as channels of the input image and construct the learning model as a multi-channel CNN. In this way, "image features specific to cardiac ultrasound images" that exist in both surface A and surface B are learned together, so the constructed learning model is expected to capture the image features more accurately.

[0097] As described above, the process from step S7021 to step S7026 is followed by the execution of the learning model acquisition process (step S702).

[0098] (Step S703: Generation of 2D cross-sectional images of the input image (for multiple cross-sections)) The process in step S703 is the same as the process in step S203 in the first embodiment, except that a label is added to the generated two-dimensional cross-sectional image to identify whether the cross-section is surface A or surface B.

[0099] (Step S704: Estimation of anatomical landmark coordinates (for multiple cross-sections)) The process in step S704 is the same as the process in step S204 in the first embodiment, except that it estimates the coordinates of anatomical landmarks for two 2D cross-sectional images, plane A and plane B. Through this process, anatomical landmarks on plane A and plane B are estimated, respectively.

[0100] (Step S705: Displaying the estimated results) In step S705, the display processing unit 48 displays the input image and estimated anatomical landmarks within the image display area of ​​the display unit 36 ​​in a display format that allows the user of the image processing device 2 to easily view them, based on the results calculated in step S704.

[0101] The display format could be, for example, two-dimensional cross-sectional images of plane A and plane B, as schematically illustrated in Figures 4B and 4D, with the locations of anatomical landmarks drawn on top of them as circles or dots. Any other display format is acceptable as long as it allows the user to easily visualize the estimated anatomical landmarks.

[0102] According to this embodiment, even when there are multiple reference cross-sections instead of just one, a learning model can be constructed to estimate the location of anatomical landmarks present on any of the reference cross-sections.

[0103] In this embodiment, the case where there are two reference cross-sections was used as an example, but the process can be carried out similarly even if there are three or more reference cross-sections. For example, even if there is a plane (plane C) that is perpendicular to both planes A and B, or if there is a reference cross-section that is parallel to any of these planes A, B, or C, the assignment destination of the anatomical landmark can be determined by the process in step S7023 and the process can be carried out similarly. Also, in this embodiment, the case where the reference cross-sections are perpendicular to each other was used as an example, but the process of this embodiment can be carried out in the same way even if they are not perpendicular to each other.

[0104] <Third Embodiment> Next, a third embodiment will be described. In the following description, the same reference numerals will be used for components and processes as in the first and second embodiments, and detailed explanations will be omitted. The image processing apparatus according to the third embodiment is, like the first embodiment, a device that estimates the positions of anatomical landmarks defined on a reference cross-section from an input three-dimensional image and a two-dimensional reference cross-section defined on the image. In the first embodiment, as shown in Figure 5C, anatomical landmarks were augmented by projecting the anatomical landmarks on the reference cross-section onto an augmented neighboring cross-section. However, the method of approximating landmarks is not limited to this method; for example, it is also possible to estimate (or correct) the positions of landmarks on the neighboring cross-section based on various image processing techniques.

[0105] The device configuration and processing flowchart of the image processing apparatus according to this embodiment are the same as those of the first embodiment. However, the landmark approximation processing (step S2024) is different. The differences from the first embodiment are described below.

[0106] (Step S2024: Landmark approximation) In this embodiment, the landmark approximation unit 49 uses pattern matching as the landmark approximation process to determine the position on a nearby cross-section that corresponds to an anatomical landmark on a reference cross-section.

[0107] The specific processing performed by the landmark approximation unit 49 will now be explained. First, using the reference cross-sectional image calculated in step S2023, a predetermined square area of ​​size is defined centered on the location of each anatomical landmark. This is called a template. By performing template matching (template matching) on ​​each neighboring cross-sectional image calculated in step S2023, the position of each template in the neighboring cross-sectional image is calculated. The positions obtained in this way are the approximate positions of each anatomical landmark.

[0108] In this case, the search range of template matching may be limited based on the position of anatomical landmarks on the reference cross section. For example, the approximate position of a landmark obtained by the method shown in the first embodiment may be defined as the initial value of the approximate position, and template matching may be performed within a predetermined search range centered on this initial value. This reduces the possibility that other feature points may be misrecognized as the landmark. Furthermore, if a suitable approximate position is not found by template matching, the initial value may be adopted as the approximate position. This process corresponds to the process of correcting the approximate position of a landmark obtained by the method shown in the first embodiment using image information. Therefore, the landmark approximation unit 49 functions as a position correction unit and corrects the anatomical landmark position in each of the multiple neighboring cross-sectional images using the anatomical landmark position identified by template matching using the cross-sectional image of the reference cross section.

[0109] Note that the template may be a 3D image instead of a 2D image. The template may be a rectangular prism shape that includes the pixel values ​​on both the near and far sides of the cross-section. Alternatively, it may be a spherical or ellipsoidal template centered on the landmark position on the reference cross-section. In this case, template matching is also performed in a three-dimensional space that takes into account the pixel values ​​on both the near and far sides of the neighboring cross-sectional images. Furthermore, the positional relationship with other anatomical landmarks may be considered when performing template matching. For example, the distance value between the anatomical landmark itself and the remaining anatomical landmarks may be used as a regularization term to suppress large changes in the distance value (i.e., changes that can be considered to disrupt the positional relationship between anatomical landmarks). By doing so, the accuracy of template matching can be improved.

[0110] Furthermore, it may be determined whether a suitable approximate location for an anatomical landmark does not exist in the nearby cross-sectional image (i.e., the anatomical landmark is lost), and the processing in this step may be divided according to the result of that determination. In this case, if no location with a similarity of a certain level or higher is found in template matching, it is determined that the landmark is lost in the nearby cross-sectional image. If it is determined that the landmark is lost, that nearby cross-section is not used (it is excluded from subsequent processing). By doing so, inappropriate nearby cross-sections with unclear anatomical landmarks can be excluded from the training data, thereby improving the quality of the training data. Alternatively, without excluding nearby cross-sections, the positions of anatomical landmarks on nearby cross-sections may be determined independently of the positions of anatomical landmarks on the reference cross-sectional image. In this case, for example, for multiple anatomical landmarks, the positions on nearby cross-sections of anatomical landmarks determined to be lost are determined so that the relative positional relationship between the landmarks is the same as the relative positional relationship in the reference cross-section. Also, nearby cross-sections are excluded only if all anatomical landmarks are determined to be lost. By doing so, the number of nearby cross-sections that are excluded can be minimized.

[0111] Alternatively, landmark approximation can be performed by registering a reference cross-sectional image with neighboring cross-sectional images. This registration process generates a deformation field representing the correspondence between the images. This deformation field is an example of pixel correspondence, indicating which position in the neighboring cross-sectional image corresponds to a specific position in the reference cross-sectional image. This deformation field can then be used to perform coordinate transformations on anatomical landmarks in the reference cross-sectional image and neighboring cross-sectional images. Various known techniques can be used for registration, including iterative optimization methods using image similarity as a cost, and machine learning-based methods. Furthermore, any known technique can be used to represent the deformation field, such as rotation matrices, affine matrices, or FFD (Free Form Deformation).

[0112] According to this embodiment, it is possible to calculate the approximate position of anatomical landmarks in a nearby cross-sectional image based on the image information of a reference cross-sectional image and a nearby cross-sectional image. This is expected to have the effect of enabling appropriate approximation even when, for example, the shape of the subject as seen on the cross-section differs significantly between the reference cross-section and the nearby cross-section, and it is not possible to calculate the appropriate approximate position of anatomical landmarks through projection alone.

[0113] <Fourth Embodiment> Next, a fourth embodiment will be described. In the following description, components and processes similar to those in the first to third embodiments will be denoted by the same reference numerals, and detailed explanations will be omitted. The image processing apparatus according to the fourth embodiment, like the first embodiment, is a device that estimates the positions of anatomical landmarks defined on a reference cross-section from an input three-dimensional image and a two-dimensional reference cross-sectional image extracted from the input image. In the first embodiment, when the learning model was constructed in step S2025, no particular selection was made, and all the training data generated in the preceding processes was used. However, the method of constructing the learning model is not limited to this method. For example, It is also possible to display augmented results to the user, obtain instructions from the user indicating whether or not to adopt the augmented results, and build a learning model using the training data of the augmented results that were deemed "adoptable."

[0114] The device configuration and flowchart of the image processing apparatus 10 according to this embodiment are the same as those of the first embodiment. However, the learning model construction process (step S2025) is different. The differences from the process in the first embodiment will be explained below.

[0115] (Step S2025: Building the Learning Model) In step S2025, the image processing device 10 controls the display processing device 48 to perform display processing for the user to confirm the augmented cross-section (nearby cross-section). Next, the user inputs a determination result of whether or not to adopt the augmented result displayed by the operation unit 35, and the image processing device 10 accepts this input. Then, the image processing device 10 constructs a learning model using only the reference cross-section and the nearby cross-section that has been determined to be adopted. The learning model construction process is the same as in step S2025 in the first embodiment.

[0116] The processing related to displaying the augmented results and obtaining the user's judgment result can be performed, for example, as follows. First, the display processing unit 48 performs a cross-sectional image display processing on the display unit 36, sequentially switching and displaying nearby cross-sectional images in response to the user's mouse wheel operation on the operation unit 35. At this time, the display processing unit 48 displays the reference cross-sectional image that will be augmented in the display area adjacent to the display area of ​​the nearby cross-sectional image on the display unit 36. Then, anatomical landmarks projected onto the nearby cross-sectional image by the processing in step S2023 are also displayed on each of these cross-sectional images. The user compares the reference plane image and the nearby cross-sectional image displayed on the display unit 36 ​​to determine whether the nearby cross-sectional image can be used as training data. The user inputs the judgment result by operating the operation unit 35, for example, by pressing the "OK" or "NG" button on the screen with the mouse. Alternatively, the operation of determining "OK" or "NG" may be assigned to a specific key on the keyboard, and the user may input the judgment result by pressing that key on the keyboard.

[0117] In step S2025, the user may not only decide whether to accept or reject the training data, but also correct the position of anatomical landmarks on nearby cross-sectional images. In this case, if the user determines that the position of an anatomical landmark on a nearby cross-sectional image is inappropriate, they may correct the position of the landmark, for example, by dragging it with the mouse.

[0118] According to this embodiment, instead of using all of the augmented training data generated by the image processing device 10 to build the learning model, the user can select and discard training data based on their judgment. This allows for the exclusion of nearby cross-sectional images that are unsuitable as training data, thereby improving the accuracy of anatomical landmark estimation.

[0119] (Variation 4-1) The following describes a modification of the fourth embodiment. In the fourth embodiment, the decision of whether or not to adopt the training data as augmentation parameters was made based on the user's judgment, but this decision may be made automatically. For example, random cross-sectional parameters may be generated near the reference cross-section, and if the similarity between the resulting 2D cross-sectional image and the reference cross-sectional image is greater than or equal to a predetermined value, the cross-sectional parameters may be adopted as augmentation parameters. Furthermore, when deciding whether or not to adopt a parameter as augmentation parameters, weighting based on the distance from the reference cross-section may be used in addition to the similarity between cross-sectional images.

[0120] In addition, there is the automatic determination described in this modified example and the one described in step S2025 of the fourth embodiment. It is also possible to combine this with manual judgment by the user. In this case, the image processing device 10 may, for example, present only those that have been judged as "rejected" by the automatic judgment to the user, and the user may make the final decision on whether or not to reject the augmented parameters based on the user's actions.

[0121] This approach reduces the effort required from users due to manual decision-making while simultaneously improving the performance of the learning model.

[0122] Although these embodiments have been described above, the technology disclosed herein is not limited to these, and can be modified or transformed within the scope of the technical concept of the disclosed technology. Furthermore, each of the above embodiments and each of its modifications may be combined as appropriate.

[0123] <Other Embodiments> The present invention can also be realized by supplying a program that implements one or more of the functions of the above-described embodiments to a system or device via a network or storage medium, and by having one or more processors in the computer of that system or device read and execute the program. It can also be realized by a circuit (e.g., an ASIC) that implements one or more functions.

[0124] Furthermore, the disclosed technology can take the form of, for example, a system, apparatus, method, program, or recording medium (storage medium). Specifically, it may be applied to a system consisting of multiple devices (e.g., a host computer, interface devices, imaging devices, web applications, etc.) or to an apparatus consisting of a single device.

[0125] Furthermore, it goes without saying that the object of the present invention is achieved as follows: a recording medium (or storage medium) containing software program code (computer program) that realizes the functions of the embodiments described above is supplied to a system or device. Such storage medium is, needless to say, a computer-readable storage medium. The computer (or CPU or MPU) of the system or device then reads and executes the program code stored on the recording medium. In this case, the program code read from the recording medium itself realizes the functions of the embodiments described above, and the recording medium containing that program code constitutes the present invention. [Explanation of symbols]

[0126] 10 Image processing device, 37 Control unit, 43 Input data acquisition unit, 48 Learning cross-sectional image generation unit, 49 Landmark approximation unit, 50 Learning unit

Claims

1. A data acquisition unit that acquires data including a three-dimensional image of a subject, a reference cross-sectional parameter representing a predetermined reference cross-section in the three-dimensional image, and the position of an anatomical landmark in the three-dimensional image. A nearby cross-sectional image acquisition unit acquires a nearby cross-sectional image based on a nearby cross-sectional parameter that represents a cross-section in the vicinity of a predetermined reference cross-section, obtained by moving the reference cross-section outside the reference cross-section based on the reference cross-sectional parameter, and a three-dimensional image. A position acquisition unit that acquires the position of an anatomical landmark in a nearby cross-sectional image by approximating the position of the anatomical landmark in the image of the predetermined reference cross-section onto the nearby cross-section, based on the position of the anatomical landmark and the relationship between the predetermined reference cross-section and the nearby cross-section, A learning unit that acquires a learning model for estimating the position of an anatomical landmark on a cross-sectional image from the cross-sectional image, which is constructed using information including the pair of the nearby cross-sectional image and the position of the anatomical landmark acquired by the position acquisition unit as learning data, An image processing apparatus characterized by comprising:

2. The image processing apparatus according to claim 1, wherein the nearby cross-sectional image acquisition unit acquires nearby cross-sectional parameters representing the nearby cross-section obtained by moving the reference cross-section outside the reference cross-section by at least one of rotational movement and translational movement, and acquires the nearby cross-sectional image based on the nearby cross-sectional parameters and the three-dimensional image.

3. The nearby cross-sectional image acquisition unit acquires a plurality of nearby cross-sectional images that define a plurality of cross-sections in the vicinity of the predetermined reference cross-section based on the reference cross-sectional parameter. The position acquisition unit acquires the positions of anatomical landmarks in each of the plurality of nearby cross-sectional images. The image processing apparatus according to claim 1 or 2.

4. The image processing apparatus according to any one of claims 1 to 3, characterized in that the nearby cross-sectional image is a cross-sectional image that defines a cross-section outside the predetermined reference cross-section.

5. The image processing apparatus according to claim 2, wherein the position acquisition unit acquires the position of the anatomical landmark in the nearby cross-sectional image by approximating the position of the anatomical landmark in the image of a predetermined reference cross-section onto the nearby cross-section by approximating at least one of the following: projection by a perpendicular line drawn down to the nearby cross-section; rotational movement corresponding to the rotational movement that acquires nearby cross-sectional parameters representing the nearby cross-section; and alignment of the reference cross-sectional image and the nearby cross-sectional image.

6. The image processing apparatus according to any one of claims 1 to 5, characterized in that the position acquisition unit acquires the position of the anatomical landmark in the nearby cross-sectional image by projecting the position of the anatomical landmark acquired by the data acquisition unit onto the nearby cross-sectional image.

7. The image processing apparatus according to claim 1 or 2, characterized in that the learning unit learns information including a set of images of a predetermined reference cross-section and the location of an anatomical landmark as learning data.

8. The nearby cross-sectional image acquisition unit acquires a plurality of nearby cross-sectional images by moving at least one of the translational and rotational movements of the predetermined reference cross-section. The image processing apparatus according to any one of claims 1 to 7, characterized in that the position acquisition unit approximates the position of the anatomical landmark acquired by the data acquisition unit to a position on the nearby cross-sectional image based on the movement of at least one of the units, thereby acquiring the position of the anatomical landmark in each of the plurality of nearby cross-sectional images.

9. The image processing apparatus according to claim 8, characterized in that the translation is a translation outside the predetermined reference cross-section, and the rotational translation is a rotational translation about a predetermined axis extending on the predetermined reference cross-section.

10. The image processing apparatus according to claim 9, characterized in that the predetermined axis is an axis passing through at least one of the anatomical landmark locations on the predetermined reference cross-section.

11. The predetermined axis is at least two or more axes passing through the predetermined reference cross-section, The aforementioned nearby cross-sectional image acquisition unit selects one of the two or more axes for each rotational movement and performs the rotational movement. The image processing apparatus according to feature 9.

12. The image processing apparatus according to any one of claims 8 to 11, characterized in that at least one of the movements is a movement at a predetermined interval.

13. An input data acquisition unit that acquires an input 3D image and an input reference cross-section parameter representing a predetermined reference cross-section in the input 3D image, An input cross-sectional image acquisition unit acquires an input cross-sectional image based on the input three-dimensional image and the input reference cross-sectional parameters, An estimation unit that estimates the position of anatomical landmarks in the input cross-sectional image using the learning model acquired by the learning unit, An image processing apparatus according to any one of claims 1 to 12, characterized by comprising:

14. The aforementioned predetermined reference cross-section is at least two or more reference cross-sections, The system further includes an assignment unit that assigns the anatomical landmark location to one of the reference cross-sections of the at least two reference cross-sections, The data acquired by the data acquisition unit includes the anatomical landmark locations assigned by the assignment unit. The learning unit acquires the learning model using a pair of cross-sectional images representing the vicinity of the reference cross-section and positions based on the anatomical landmark locations assigned by the assignment unit as learning data. The image processing apparatus according to any one of claims 1 to 13.

15. The image processing apparatus according to any one of claims 1 to 14, characterized in that the position acquisition unit acquires the position of an anatomical landmark in the nearby cross-sectional image by identifying the position of an anatomical landmark in the nearby cross-sectional image that corresponds to the position of an anatomical landmark in the cross-sectional image of the predetermined reference cross-section, based on the correspondence between pixels in the cross-sectional image of the predetermined reference cross-section and pixels in the nearby cross-sectional image.

16. The image processing apparatus according to claim 15, further comprising a position correction unit that corrects the position of anatomical landmarks in the nearby cross-sectional image using the position of anatomical landmarks obtained by the position acquisition unit as a reference, and using the position of anatomical landmarks identified by template matching using a template determined based on the predetermined reference cross-section.

17. The image processing apparatus according to claim 16, wherein the learning unit excludes from the learning data any nearby cross-sectional images in which the anatomical landmark positions on the nearby cross-sectional images corresponding to the anatomical landmark positions on the cross-sectional images of the predetermined reference cross-sections cannot be identified in the template matching.

18. A display unit that displays the nearby cross-sectional image acquired by the nearby cross-sectional image acquisition unit, An acquisition unit that acquires a determination result of whether or not to adopt the nearby cross-sectional image displayed on the display unit as training data. Furthermore, The learning unit removes nearby cross-sectional images from the learning data that indicate the determination result will not be used as learning data. The image processing apparatus according to any one of claims 1 to 17.

19. An input data acquisition unit that acquires an input 3D image and an input reference cross-section parameter representing a predetermined reference cross-section in the input 3D image, An input cross-sectional image acquisition unit acquires an input cross-sectional image based on the input three-dimensional image and the input reference cross-sectional parameters, A learning model acquisition unit that acquires a learning model for estimating the position of anatomical landmarks on a cross-sectional image from the cross-sectional image, An estimation unit that estimates the position of anatomical landmarks in the input cross-sectional image using the learning model acquired by the learning model acquisition unit, Equipped with, The aforementioned learning model, A 3D image of the subject, A reference cross-section parameter representing a predetermined reference cross-section in the three-dimensional image, The position of anatomical landmarks in the aforementioned three-dimensional image, A nearby cross-sectional image based on the reference cross-sectional parameters obtained by moving the reference cross-section outside the reference cross-section based on the reference cross-sectional parameters, a nearby cross-sectional image based on the reference cross-sectional parameters obtained based on the three-dimensional image, Based on the anatomical landmark position and the relationship between the predetermined reference cross section and the nearby cross section, the anatomical landmark position in the nearby cross section image obtained by approximating the anatomical landmark position in the image of the predetermined reference cross section onto the nearby cross section is obtained. This is a learning model constructed using the pair of the nearby cross-sectional image obtained based on and the anatomical landmark location as training data. An image processing apparatus characterized by the following:

20. A data acquisition unit that acquires data including a three-dimensional image of a subject, a reference cross-sectional parameter representing a predetermined reference cross-section in the three-dimensional image, and the position of an anatomical landmark in the three-dimensional image. A nearby cross-sectional image acquisition unit acquires a nearby cross-sectional image based on a nearby cross-sectional parameter representing a nearby cross-section obtained by moving the predetermined reference cross-section outside the predetermined reference cross-section by rotating the predetermined reference cross-section around an axis passing through at least one of the anatomical landmark positions on the predetermined reference cross-section, and a three-dimensional image. A learning unit that acquires a learning model for estimating the position of an anatomical landmark on a cross-sectional image from a cross-sectional image, constructed using information including a pair of the position of the anatomical landmark in the nearby cross-sectional image and the nearby cross-sectional image, obtained by approximating the position of the anatomical landmark in the image of the predetermined reference cross-section on the nearby cross-section based on the position of the anatomical landmark and the relationship between the predetermined reference cross-section and the nearby cross-section, as learning data. An image processing apparatus characterized by comprising:

21. A step of acquiring data including a three-dimensional image of a subject, a reference cross-sectional parameter representing a predetermined reference cross-section in the three-dimensional image, and the position of an anatomical landmark in the three-dimensional image. A step of acquiring a nearby cross-sectional image based on a nearby cross-sectional parameter that represents a cross-section in the vicinity of a predetermined reference cross-section, obtained by moving the reference cross-section outside the reference cross-section based on the reference cross-sectional parameter, and the three-dimensional image. The steps include obtaining the position of the anatomical landmark in the nearby cross-sectional image by approximating the position of the anatomical landmark in the image of the predetermined reference cross-section onto the nearby cross-section, based on the position of the anatomical landmark and the relationship between the predetermined reference cross-section and the nearby cross-section, The steps include: obtaining a learning model that estimates the location of an anatomical landmark on a cross-sectional image from the cross-sectional image, which is constructed using information including the pair of the nearby cross-sectional image and the acquired location of the anatomical landmark as training data; Image processing methods including [specific details omitted].

22. The steps include obtaining an input three-dimensional image and input reference cross-section parameters that represent a predetermined reference cross-section in the input three-dimensional image, The steps include: acquiring the input three-dimensional image and an input cross-sectional image based on the input reference cross-sectional parameters; The steps include obtaining a learning model that estimates the location of anatomical landmarks on a cross-sectional image from the cross-sectional image, The steps include: using the acquired learning model to estimate the location of anatomical landmarks in the input cross-sectional image; Includes, The aforementioned learning model, A 3D image of the subject, A reference cross-section parameter representing a predetermined reference cross-section in the three-dimensional image, The position of anatomical landmarks in the aforementioned three-dimensional image, A nearby cross-sectional image based on the reference cross-sectional parameters obtained by moving the reference cross-section outside the reference cross-section based on the reference cross-sectional parameters, a nearby cross-sectional image based on the reference cross-sectional parameters obtained based on the three-dimensional image, Based on the anatomical landmark position and the relationship between the predetermined reference cross section and the nearby cross section, the anatomical landmark position in the nearby cross section image obtained by approximating the anatomical landmark position in the image of the predetermined reference cross section onto the nearby cross section is obtained. This is a learning model constructed using the pair of the nearby cross-sectional image obtained based on and the anatomical landmark location as training data. An image processing method characterized by the following:

23. A step of acquiring data including a three-dimensional image of a subject, a reference cross-sectional parameter representing a predetermined reference cross-section in the three-dimensional image, and the position of an anatomical landmark in the three-dimensional image. A step of acquiring a nearby cross-sectional image based on a nearby cross-sectional parameter representing a nearby cross-section obtained by moving the predetermined reference cross-section outside the predetermined reference cross-section by rotating the predetermined reference cross-section around an axis passing through at least one of the anatomical landmark positions on the predetermined reference cross-section, and a three-dimensional image. An image processing method comprising the step of obtaining a learning model for estimating the position of an anatomical landmark on a cross-sectional image from a cross-sectional image, which is constructed using information including a pair of the position of the anatomical landmark in a nearby cross-sectional image and the nearby cross-sectional image obtained by approximating the position of the anatomical landmark in an image of a predetermined reference cross-section on the nearby cross-section, based on the position of the anatomical landmark and the relationship between the predetermined reference cross-section and the nearby cross-section, as learning data.

24. A program that causes a computer to execute the image processing method described in any one of claims 21 to 23.

Citation Information

Patent Citations

  • A method and system for classifying and locating pulmonary nodules based on multi-slice CT images

    CN109523521A

  • Unsupervised three-dimensional medical image registration method and system based on neural network

    CN110599528A

  • Hierarchical segmentation method and device for tissue structure in medical image, equipment and medium

    CN113822845A

  • A method for detecting polyps in a 3D image volume

    JP2008515595A

  • Medical image generation apparatus

    JP2019118694A