Image processing device, image processing method, and program
The image processing device improves contour extraction accuracy in medical images by using subspaces from divided training data sets and the BPLP method, addressing inaccuracies in existing techniques and reducing computational costs.
Patent Information
- Application Number
- JP2025005976
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-01-16
- Publication Date
- 2026-01-14
- Estimated Expiration
- 2039-11-26
AI Technical Summary
Existing image processing techniques for medical diagnosis, particularly those using statistical models, struggle with inaccurate contour extraction when faced with images significantly different from training data, leading to inappropriate inference results due to insufficient generalization ability.
An image processing device that estimates shape information by acquiring pixel value information and shape information of a region of interest, utilizing multiple subspaces generated from divided training data sets, selecting an appropriate subspace based on the input image, and employing a method like BPLP to estimate contour information without iterative processing.
This approach enhances the accuracy of contour extraction from medical images by reducing the risk of inappropriate inference and lowering computational costs, ensuring precise estimation of region shapes.
Smart Images

Figure 0007799095000010 
Figure 0007799095000011 
Figure 0007799095000012
Abstract
Description
[Technical Field]
[0001] The present invention relates to an image processing device, an image processing method, and a program. [Background technology]
[0002] In the medical field, diagnoses are performed using images acquired by various imaging devices (modalities), such as ultrasound diagnostic imaging devices. Information such as the area, volume, and dimensions of a region of interest captured in an image is used for diagnosis. To calculate the area of a region, it is necessary to extract the contour of the region from the image (i.e., to estimate contour information, which is information representing the contour shape). However, when this region extraction work is performed manually, it requires a great deal of labor from the operator, which is a problem. Therefore, various techniques for automatic or semi-automatic region extraction from images have been proposed to reduce the operator's labor.
[0003] One example is a method using images and correct contour information of a region of interest in the images collected for a large number of cases as training data, as shown in the following prior art. This technology estimates contour information of a region of interest in an unknown input image (i.e., extracts the contour) based on the results of statistical analysis of the training data. Non-Patent Document 1 discloses a technology for constructing a statistical model called an Active Appearance Model based on pixel value information of the image in the training data and coordinate value information of a point cloud representing the contour of the region of interest in the image. Furthermore, it also discloses a technology for estimating contour information of a region of interest in the input image using iterative processing based on a gradient method, using an evaluation value representing the similarity between pixel value information of an unknown input image and pixel value information of an image in the statistical model. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] Elco Oost,et.al.”Active Appearance Models in Medical Image Processing” Medical Imaging Technology.vol.27.No3.2009 May. Summary of the Invention [Problem to be solved by the invention]
[0005] In estimation processes using statistical models, plausible inference results are obtained when an image similar to the sample data set included in the training data is input. However, in rare cases, inappropriate inference results may be output when an image far removed from the sample data set is input. In general, when generating a statistical model, it is considered desirable to have as much variation in the training data as possible. This is because it is expected that the greater the variation in the training data, the higher the generalization ability of the statistical model will be, making it able to handle a greater amount of unknown data. However, the more variation in the training data is increased and the generalization ability of the statistical model is improved, the more likely it is that the inappropriate inference results mentioned above will be output.
[0006] The present invention has been made in view of the above-mentioned problems, and has an object to provide a technique that can extract information about the shape of a region of interest from a medical image with higher accuracy.
[0007] In addition to the above-mentioned object, the present specification also provides operational effects that cannot be obtained by conventional techniques, which are derived from the configurations shown in the below-described embodiments of the invention. This can be positioned as one of the other purposes of the disclosure of the book. [Means for solving the problem]
[0008] A first aspect of the present invention is an image processing device that estimates shape information, which is information about the shape of a region of interest, from an image, comprising: an image acquisition means that acquires a first image to be processed; The sample data includes pixel value information, which is data arranging principal component scores of the images, and shape information of the region of interest, which are obtained by projecting each image of the plurality of images for learning onto a subspace obtained by principal component analysis of the plurality of images for learning, and different Two or moreThe image processing device includes a subspace information acquisition means for acquiring information on two or more subspaces generated using a data set, a subspace information selection means for selecting information on a first subspace from the information on the two or more subspaces based on pixel value information of the first image, and an estimation means for estimating shape information of a region of interest in the first image from the pixel value information of the first image using the information on the first subspace. vinegar The present invention provides an image processing device characterized by:
[0009] A second aspect of the present invention is an image processing method for estimating shape information, which is information about the shape of a region of interest, from an image, comprising: acquiring a first image to be processed; Image acquisition process and, The sample data includes pixel value information, which is data arranging principal component scores of the images, and shape information of the region of interest, which are obtained by projecting each image of the plurality of images for learning onto a subspace obtained by principal component analysis of the plurality of images for learning, and different Two or more Obtain information on two or more subspaces generated using a dataset Subspace information acquisition process and selecting information of a first subspace from information of the two or more subspaces based on pixel value information of the first image. Subspace information selection process and estimating shape information of a region of interest in the first image from pixel value information of the first image using information of the first subspace. Estimation process And, with vinegar The present invention provides an image processing method comprising the steps of:
[0010] A third aspect of the present invention is a program for causing a computer to function as each means of the image processing device according to the first aspect, or a program for causing a computer to perform each means of the image processing method according to the second aspect. process Run Rup and a computer-readable storage medium non-temporarily storing such a program. [Effects of the Invention]
[0011] According to the present invention, information about the shape of a region of interest can be extracted from a medical image with higher accuracy. [Brief explanation of the drawings]
[0012] [Figure 1] FIG. 2 is a diagram showing the functional configuration of the image processing apparatus. [Figure 2] 10 is a flowchart showing an example of a processing procedure of the image processing device. [Figure 3] FIG. 1 is a diagram showing an example of an ultrasound image of the heart. [Figure 4] 1 is a diagram showing an example of a point cloud representing the contour of a region of interest in an ultrasound image of the heart. DETAILED DESCRIPTION OF THE INVENTION
[0013] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. However, the components described in the embodiments are merely examples, and the technical scope of the present invention is determined by the claims, and is not limited to the individual embodiments described below.
[0014] An image processing device according to an embodiment of the present invention provides a function for estimating information about the shape of a region of interest (ROI) from an input image. The input image to be processed is a medical image, i.e., an image of a subject (such as a human body) photographed or generated for the purpose of medical diagnosis, examination, research, etc., and is typically an image acquired by an imaging system called a modality. For example, ultrasound images obtained by an ultrasound diagnostic device, X-ray CT images obtained by an X-ray CT device, MRI images obtained by an MRI device, etc. can be processed. The input image may be a two-dimensional image or a three-dimensional image. The image may be of one time phase or of multiple time phases. A region of interest is a portion of an image, such as an anatomical structure (organ, blood vessel, bone, etc.) or a lesion. The region of interest can be arbitrarily selected. Shape information about the region of interest (also referred to as "shape information" in this specification) is information that represents the spatial (or geometric) characteristics of the region of interest in the image. Examples of shape information include contour information about the region of interest (information representing the contour shape), the positions of feature points on the region of interest, the relative positions (e.g., spacing) between feature points, the length, area, and volume of the region of interest, and the shape of the region of interest (circle, ellipse, triangle, etc.). Shape information includes not only primary information obtained directly from the image, but also secondary information generated using the primary information (e.g., the difference in the positions of feature points, the area ratio between two time phases, etc.).
[0015] One of the features of an image processing device according to an embodiment of the present invention is that it acquires information on two or more subspaces generated using different data sets, selects information on a subspace appropriate for an input image from the information on the two or more subspaces, and estimates shape information. When training data is divided into multiple data sets, the expressive power and generalization ability of the statistical model (subspace) may be inferior compared to when statistical analysis is performed on the entire training data, but the risk of outputting inappropriate estimation results is reduced. It is preferable to use data as training data, in which each sample data contains pixel value information of a training image and shape information (ground truth data) of a region of interest within that training image. Using such training data makes it possible to obtain a model that accurately reflects the correlation between the pixel value information of an image and the shape information of a region of interest, thereby improving estimation accuracy.
[0016] A specific example of an image processing device according to an embodiment of the present invention will be described in detail below, taking as an example a case where the region of the left atrium of the heart is extracted from a two-dimensional ultrasound image.
[0017] The image processing device according to this embodiment has a function of automatically extracting contour information of the left atrium, which is a region of interest, from an input image. The image processing device according to this embodiment selects an appropriate subspace from two or more subspaces (statistical models) constructed from training data according to the input image to be processed, and estimates shape information of the region of interest in the input image using the selected subspace.
[0018] The configuration and processing of the image processing device of this embodiment will be described below with reference to Fig. 1. Fig. 1 is a block diagram showing an example configuration of an image processing system (also called a medical image processing system) including the image processing device of this embodiment. The image processing system includes an image processing device 10 and a database 22. The image processing device 10 is communicably connected to the database 22 via a network 21. The network 21 includes, for example, a LAN (Local Area Network) or a WAN (Wide Area Network).
[0019] The database 22 stores and manages multiple images. The images managed by the database 22 include images whose contour information of the region of interest is unknown and images whose contour information of the region of interest is known. The former images are used for contour extraction processing by the image processing device 10. The latter images are associated with contour information (ground truth data) and are used as training data. It is desirable that these images have been spatially normalized to a certain extent with respect to the image size and the region of interest depicted in the image. Spatial normalization is an operation of aligning the size (dimension per pixel) and position in the image to a certain standard. The information managed by the database 22 may include information on subspaces used in the contour extraction processing (e.g., information on estimated matrices). The subspace information may be stored in an internal memory (ROM 32 or memory unit 34) of the image processing device 10 instead of in the database 22. The image processing device 10 can acquire data stored in the database 22 via the network 21.
[0020] The image processing device 10 includes a communication IF (Interface) 31 (communication unit), a ROM (Read Only Memory) 32, a RAM (Random Access Memory) ry) 33, a memory unit 34, an operation unit 35, a display unit 36, and a control unit 37.
[0021] The communication IF 31 (communication unit) is configured with a LAN card or the like, and realizes communication between an external device (e.g., the database 22) and the image processing device 10. The ROM 32 is configured with a non-volatile memory or the like, and stores various programs and various data. The RAM 33 is configured with a volatile memory or the like, and is used as a work memory for temporarily storing programs and data currently being executed. The storage unit 34 is configured with an HDD (Hard Disk Drive) or the like, and stores various programs and various data. The operation unit 35 is configured with a keyboard, mouse, touch panel, etc., and inputs instructions from a user (e.g., a doctor or a medical technician) to various devices.
[0022] The display unit 36 is configured with a display or the like and displays various information to the user. The control unit 37 is configured with a CPU (Central Processing Unit) or the like and controls the overall processing in the image processing device 10. The control unit 37 has, as its functional configuration, an image acquisition unit 51, a subspace information acquisition unit 52, a contour information estimation unit 53, and a display processing unit 54. The control unit 37 may also have a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), an FPGA (Field-Programmable Gate Array), or the like.
[0023] The image acquisition unit 51 acquires a first image (an image with unknown contour information) to be processed from the database 22. In other words, the image acquisition unit 51 corresponds to an example of an image acquisition means for acquiring a first image. The first image is an image of a subject acquired by various modalities. The image acquisition unit 51 may acquire the first image directly from the modality. In this case, the image processing device 10 may be implemented in the console of the modality (imaging system). In this embodiment, an example is described in which the first image is a two-dimensional ultrasound image, but other types of images may also be used. The method of this embodiment is also applicable to images that are two or more dimensional (such as multiple two-dimensional images, two-dimensional moving images, three-dimensional still images, multiple three-dimensional images, or three-dimensional moving images). Furthermore, it is applicable regardless of the type of modality.
[0024] The subspace information acquisition unit 52 acquires training data from the database 22 and performs two or more statistical analyses using this training data. Each piece of sample data constituting the training data is data configured to include pixel value information of a training image and shape information of a region of interest in that training image (details will be described later). Then, the subspace information acquisition unit 52 calculates subspace information (e.g., information on the bases constituting the subspace) that represents the distribution of the sample data group from the results of the statistical analysis. At this time, the subspace information acquisition unit 52 divides the training data into two or more data sets and generates information on the two or more subspaces using different data sets. In other words, the subspace information acquisition unit 52 corresponds to an example of subspace information acquisition means that acquires information on two or more subspaces generated using different data sets from the training data.
[0025] The subspace information selection unit 55 selects first subspace information suitable for estimating contour information of the region of interest from the first image, based on the first image acquired by the image acquisition unit 51 and the two or more pieces of subspace information acquired by the subspace information acquisition unit 52. In other words, the subspace information selection unit 55 corresponds to an example of a subspace information selection means that selects first subspace information from information on two or more subspaces, based on pixel value information of the first image.
[0026] The contour information estimation unit 53 estimates contour information of the left atrium in the first image acquired by the image acquisition unit 51 through calculation using the first subspace information selected by the subspace information selection unit 55. That is, the contour information estimation unit 53 is an example of an estimation means that estimates shape information of the region of interest in the first image from pixel value information of the first image using the first subspace information. be.
[0027] Based on the results calculated by the contour information estimation unit 53, the display processing unit 54 displays the first image and the contour information of the estimated region of interest in the image display area of the display unit 36 in a display format that makes the contour information easily visible.
[0028] Each component of the image processing device 10 functions according to a computer program. For example, the control unit 37 (CPU) uses the RAM 33 as a work area to read and execute a computer program stored in the ROM 32 or the storage unit 34, thereby realizing the function of each component. Note that some or all of the functions of the components of the image processing device 10 may be realized using dedicated circuits. Furthermore, some functions of the components of the control unit 37 may be realized using a cloud computer. For example, a calculation device located at a different location from the image processing device 10 may be communicably connected to the image processing device 10 via the network 21, and the image processing device 10 and the calculation device may transmit and receive data to each other, thereby realizing the functions of the components of the image processing device 10 or the control unit 37.
[0029] Next, an example of processing by the image processing device 10 in FIG. 1 will be described with reference to FIG. 2. FIG. 2 is a flowchart showing an example of the processing procedure of the image processing device 10. In this embodiment, an example will be described in which the left atrium region of the heart is set as the region of interest. However, this embodiment can also be applied to cases in which the region of interest is set as other parts including the left ventricle, right ventricle, and right atrium, or a region combining a plurality of these regions.
[0030] (Step S101: Acquire and display image) In step S101, when the user issues an instruction to acquire an image via the operation unit 35, the image acquisition unit 51 acquires the first image specified by the user from the database 22 and stores it in the RAM 33. At this time, the display processing unit 54 may also display the first image in the image display area of the display unit 36. An example of the first image is shown in FIG. 3. FIG. 3 shows an example in which the first image is an apical four-chamber view in an ultrasound image of the heart. In the following description, it is assumed that the number of pixels in the x direction that make up the first image is Nx, and the number of pixels in the y direction is Ny. In other words, the total number of pixels that make up the first image is Nx × Ny.
[0031] (Step S102: Acquiring two or more pieces of subspace information) In step S102, the subspace information acquisition unit 52 acquires pixel value information (e.g., image data) of a plurality of images different from the first image and correct contour information of the region of interest in the image from the database 22. That is, the subspace information acquisition unit 52 acquires training data used to calculate subspace information. Note that, in order to increase the robustness of the statistical model, it is desirable that the plurality of images used as training data be composed of images taken of different patients. However, images taken of the same patient at different times may also be included. Alternatively, data augmentation (increasing the size of training data) may be performed by, for example, deforming or processing the images.
[0032] Next, the subspace information acquisition unit 52 generates data that combines pixel value information and contour information of the region of interest for each of the multiple images acquired as learning data. Here, the pixel value information of the image and the contour information of the region of interest will be described in detail with reference to FIGS. 3 and 4.
[0033] As shown in Fig. 3, the image handled in this embodiment is an image composed of Nx x Ny pixels. In this embodiment, pixel value information of an image refers to a column vector in which the pixel values of each pixel are arranged in the raster scan order of the image. That is, with the upper left corner of the image shown in Fig. 3 set as the origin (0,0), the pixel value at the location of pixel (x, y) is represented by I(x, y), and pixel value information a of the image is defined as follows:
number
[0034] Next, FIG. 4 shows an example in which the region of interest is the left atrium, and the contour information of the region is represented by a group of nine points (p1 to p9). In this embodiment, each of these points is represented by (x, y) using coordinate values on the x and y axes, with the upper left corner of the image being the origin (0, 0). Taking point p1 as an example, the coordinates of point p1 are (x1, y1). In this embodiment, the contour information of the region of interest refers to a column vector that lists the x and y coordinate values of multiple feature points arranged on the contour of the region of interest. That is, in the example shown in FIG. 4, contour information b representing the left atrium is defined as follows: Note that even when the "shape information" is the position of a feature point on the region of interest, b can be defined in the same format as above, with the coordinate values of the feature points listed, such as b={x1, y1}. Furthermore, if the "shape information" is the length, area, or volume of the region of interest, then b can be defined based on these values, such as b=length, b=area, and b=volume.
number
[0035] After acquiring the pixel value information a and the contour information b, the subspace information acquisition unit 52 further connects these column vectors to generate one column vector that combines the two pieces of information. That is, the subspace information acquisition unit 52 generates a column vector that includes an element corresponding to the pixel value information a and an element corresponding to the contour information b (shape information) of the region of interest. In the example shown in Figures 3 and 4, the following data c is generated.
number
[0036] Here, since the pixel value information of the image and the contour information of the region of interest have different variances, at least one of the data may be weighted. In this case, weighting may be performed according to the total variance of the pixel value information and contour information of the learning data so that the total variances are equal or balanced by a predetermined value. Alternatively, data may be weighted using weights set by the user via the operation unit 35.
[0037] The subspace information acquisition unit 52 generates data c by linking pixel value information and contour information for each of the acquired images (multiple learning sample data) in the above manner. The subspace information acquisition unit 52 then divides the generated data group into two or more data sets and performs statistical analysis on each data set. For the statistical analysis, a known method such as principal component analysis (PCA) can be used, and methods such as kernel PCA or weighted PCA may also be used. By performing PCA, it is possible to calculate the mean vector and eigenvectors for data c in which pixel value information and contour information are integrated, as well as the eigenvalues corresponding to each eigenvector. That is, by using the mean vector (c) of data c in which pixel value information and contour information are linked, the eigenvector e, and the coefficient g for each eigenvector e, a point d in the subspace can be expressed by the following formula: This point d represents one learning sample data. The coefficient g is the principal component score.
number
[0038] Here, L represents the number of eigenvectors used in the calculation, and its specific value may be determined based on the cumulative contribution rate that can be calculated from the eigenvalues. For example, a cumulative contribution rate of 95% may be set in advance as a threshold, and the number of eigenvectors whose cumulative contribution rate is 95% or more may be calculated and set as L. Alternatively, the number of eigenvectors set by the user may be set as L.
[0039] As the final process of step S102, the subspace information acquisition unit 52 acquires subspace information corresponding to each data set from the results of the statistical analysis of each data set, and stores the acquired information in the RAM 33. The subspace information is information that defines a subspace, and includes, for example, information on a plurality of eigenvectors and a mean vector that constitute the subspace. In this manner, in this embodiment, two or more pieces of subspace information having different bases are acquired based on the training data.
[0040] It should be noted that the calculation process of the subspace information in this step is data processing independent of the contour extraction (contour estimation) of the first image. Therefore, the calculation process of the subspace information may be performed in advance, and the calculated subspace information may be saved in a storage device (for example, the database 22 or the storage unit 34). In this case, in step S102, the subspace information acquisition unit 52 may read two or more pieces of subspace information from the storage device and store them in the RAM 33. Calculating the subspace information in advance has the effect of shortening the processing time required for estimating the region of interest in the first image. It should be noted that the calculation process of the subspace information may be performed by a device other than the image processing device 10.
[0041] In addition, in this embodiment, an example has been shown in which data in which pixel values of the entire image are arranged in raster scan order is used as pixel value information, but other feature amounts related to pixel values (feature amounts related to the texture of the image, etc.) may also be used as pixel value information. Furthermore, data in which pixel values of a partial image representing a part of the image are arranged may also be used as pixel value information, or data in which only pixel values around the contour of a region of interest in the image are arranged may also be used as pixel value information.
[0042] Alternatively, data obtained by projecting an image onto a subspace obtained by principal component analysis of multiple training images may be used as pixel value information for the image. For example, after calculating pixel value information a for each of multiple training images, principal component analysis may be performed on the data set containing pixel value information a, and the vector of principal component scores for each image may be used as new pixel value information a' for the image instead of a. If the total number of multiple images used as training data is smaller than the number of pixels constituting the image, the number of dimensions of pixel value information a' will be smaller than the number of dimensions of pixel value information a. Therefore, using pixel value information a' has the effect of reducing the computational cost of statistical analysis of data combining pixel value information and contour information. Furthermore, computational cost may be further reduced by setting a threshold for the cumulative contribution rate to reduce the number of dimensions of the principal component scores (i.e., the number of eigenvectors).
[0043] On the other hand, while the coordinate values of the point cloud representing the contour of the region of interest are used as the contour information of the region of interest, other values may also be used. For example, the contour information may be information obtained by arranging the results of calculating a level set function representing the region of interest (e.g., a signed distance value from the contour, with the inside of the region of interest being negative and the outside being positive) for each pixel of the image in raster scan order. Alternatively, a label image or mask image that distinguishes the region of interest from the rest may also be used as the contour information. Furthermore, similar to the method for calculating pixel value information a' from pixel value information a described above, values calculated using principal component analysis for contour information b may also be used as new contour information b'. In other words, by performing principal component analysis on a data set containing only contour information b corresponding to multiple training images and setting the vector of principal component scores for each image as new contour information b', computational costs can be further reduced.
[0044] Here, a specific example of a method for grouping training data into two or more datasets will be described. For example, images acquired from the database 22 and contour information of the regions of interest in the images may be subjected to a predetermined image transformation, and then the data may be grouped for each transformation. That is, two or more datasets may be created by performing data augmentation (increasing the size of the training data) on the image data of the multiple images and the point cloud representing the contours of the regions of interest. More specifically, two or more datasets may be created by generating data that has been subjected to multiple operations such as translation, rotation, and scaling, and then grouping data with the same or similar transformation parameters into the same group. Alternatively, the training data may be grouped based on patient information (e.g., height, weight, or medical history) of the subject appearing in the image. Alternatively, the training data may be grouped based on differences in time phase (diastole vs. systole), cross-sectional direction (four-chamber vs. two-chamber), examination type or imaging protocol, modality model, hospital, or photographer. Furthermore, the grouping may be performed based on a combination of these conditions. Furthermore, the grouping may be performed using any classification that affects the shape or image quality of the regions of interest.
[0045] Here, we will discuss the effect of dividing training data into two or more sets, generating two or more sets of subspace information, and selecting one subspace from among them. As shown in Equation (4), the results obtained from principal component analysis express the subspace as a linear sum of a mean vector and an eigenvector. Considering the principal component analysis of the pixel values of each pixel in an image, it is believed that in the aforementioned subspace, an image that is plausible as image data (an image close to the training data) is obtained in the peripheral area of the training data for the analysis. However, for coordinates in the subspace that are far from the training data, backprojection from the subspace to the original image space may result in an unnatural image. This is an issue that arises because a linear model is used to model the non-linear relationship between pixel values of different images. If such points exist in the subspace, the output may be inappropriate when performing estimation processing using that subspace.
[0046] In general, it is desirable for training data to be subjected to statistical analysis to have as much variation as possible, such as the shape of the region of interest and where the region of interest is located in the image. This is because the greater the variation in this data, the greater the expressive power of the subspace obtained from the statistical analysis, making it possible to accommodate a greater amount of unknown data. However, the greater the expressive power of the subspace, the greater the risk of the subspace containing the aforementioned unnatural images. In consideration of this problem, this embodiment employs a method in which, rather than generating a single subspace using all sample data in the training data, the training data is divided into multiple data sets each containing relatively similar sample data, and a subspace is then generated for each of these data sets. In other words, this method enables the generation of subspaces specialized for groups of training sample data with certain characteristics according to the data classification method. By dividing the training data into multiple data sets, the expressive power of the subspace for each set is reduced compared to generating a single subspace as a whole, but this is expected to have the effect of reducing the risk of unnatural images being included in the subspace. In the method of this embodiment, when an unknown input image is given, it is important to select the subspace information to use for estimation. Therefore, in the next step S103, subspace information suitable for the estimation process is selected.
[0047] (Step S103: Selection of subspace information) In step S103, the subspace information selection unit 55 selects one piece of subspace information based on the first image acquired by the image acquisition unit 51 and the two or more pieces of subspace information acquired by the subspace information acquisition unit 52, and stores the result in the RAM 33. That is, the subspace information (first subspace information) that is more suitable for estimating the contour information of the region of interest in the first image is selected. As the spatial information, one piece of subspace information is selected from two or more pieces of subspace information based on the first image.
[0048] The process of selecting subspace information suitable for estimating contour information of the region of interest from two or more pieces of subspace information can be realized by using the similarity between the pixel value information of the first image and the pixel value information of the sample data included in each data set. That is, the subspace information selection unit 55 may identify a data set including sample data similar to the first image based on the pixel value information of the image, and select subspace information generated from the identified data set.
[0049] Here, a specific method for identifying one dataset optimal for the first image from two or more datasets includes a method of evaluating the difference between the first image and each sample image. The difference between two images may be evaluated, for example, by the sum of squared differences (SSD) of pixel values. For example, the sample data with the smallest SSD is considered to be the sample data most similar to the first image (referred to as similar sample data), and the dataset including this similar sample data is identified as the dataset optimal for the estimation process. In this way, by using the subspace information generated from the dataset including the sample data most similar to the first image, it is expected that the accuracy of the estimation process in the subsequent step S104 will be improved. However, this method has the problem that the calculation cost will be enormous if the total number of sample data is enormous.
[0050] Therefore, in this embodiment, the subspace information selection unit 55 performs principal component analysis on pixel value information of sample data of each data set, and then identifies the optimal data set by the subspace method using the results of the principal component analysis and the pixel value information of the first image. This method can reduce the calculation cost of the above-mentioned identification process.
[0051] The procedure for processing using the subspace method is as follows: First, a principal component analysis is performed on pixel value information for each of two or more datasets to generate a subspace. Next, for each of the subspaces, a projection and backprojection process of the first image onto the subspace is performed to calculate the reconstruction error. That is, the pixel value information of the first image is projected onto the subspace, and then the sum of squared pixel value differences (i.e., reconstruction error) is calculated between the image obtained by backprojecting the pixel value information of the first image onto the subspace and the first image. The dataset corresponding to the subspace with the smallest reconstruction error is the dataset containing sample data similar to the first image; in other words, the dataset providing the subspace that can most accurately represent the first image. This method allows the dataset corresponding to the subspace information most suitable for the subsequent process in step S104 to be identified.
[0052] Although the example shown here uses the sum of squared differences of pixel values to evaluate the similarity between images, other indices such as the sum of absolute differences (SAD) may also be used.
[0053] Furthermore, the method is not limited to selecting one optimal dataset from multiple datasets, but may instead select all datasets that satisfy a predetermined condition. For example, all datasets whose pixel value information reconstruction error is smaller than a predetermined threshold may be selected. Alternatively, a predetermined number of datasets with the smallest reconstruction error may be selected. When multiple datasets are selected, in step S104, the contour information estimation unit 53 may execute a contour information estimation process using the subspace information of each dataset and integrate the results. For example, the average or median of the contour information estimated using each subspace information may be obtained.
[0054] (Step S104: Estimation of contour information of the region of interest) In step S104, the contour information estimation unit 53 estimates contour information of the region of interest in the first image, based on the first image acquired in step S101 and the first subspace information selected in step S103. The estimation result is stored in the RAM 33.
[0055] Any method for estimating the contour information may be used. For example, Non-Patent Document 1 discloses a technique for estimating contour information of a region of interest in an input image based on iterative processing using a gradient method, using an evaluation value that indicates the consistency between pixel value information of an unknown input image and pixel value information of an image in subspace information (statistical model). The contour information estimation unit 53 may estimate the contour information using such iterative processing using a gradient method.
[0056] However, when using iterative processing to estimate contour information, it is necessary to repeat the iterative processing until the change in a predefined energy function before and after the iteration becomes sufficiently small or until a certain number of iterations is reached, which incurs a certain amount of computational cost.In addition, because the number of iterations (the timing at which the processing converges) varies depending on the input data, there is also the issue that it is difficult to grasp in advance the computation time required for the processing due to variations in the overall computation time.
[0057] Therefore, in this embodiment, when estimating the contour information of the region of interest, a method that does not use iterative processing is applied, thereby realizing a contour information estimation process with lower calculation costs.
[0058] Here, we will summarize the problem of estimating contour information representing a predetermined region of interest (left atrium) when an unknown input image (ultrasound image) is given. The first subspace information selected in the previous step S103 is information on a subspace that associates pixel value information of the ultrasound image with point cloud information representing the contour of the left atrium shown in the image. When dealing with an unknown input image, an ultrasound image is given, so pixel value information a corresponding to the image can be calculated using the same procedure as in step S102. In other words, the proposition in this step is to estimate contour information b corresponding to the input image based on the first subspace information calculated from training data and pixel value information a related to the input image.
[0059] Non-Patent Document 2 below discloses a technique for estimating information representing the pose of an object appearing in an image from pixel value information of an unknown image, by utilizing subspace information related to data that combines pixel value information of multiple images and information representing the pose of the object appearing in the image. In other words, this technique is a technique for interpolating data in which a certain loss occurs based on the results of statistical analysis of training data without the loss. This technique is called the back projection for lost pixels (BPLP) method.
[0060] [Non-Patent Document 2] Toshiyuki Amano,et.al.”An appearance based fast linear pose estimation” MVA 2009 IAPR Conference on Machine Vision Applications.2009 May 20-22.
[0061] In order to apply the BPLP method, it is necessary to identify which parts of the input data are known information and which parts are unknown information (missing parts) at the time of estimation. In this embodiment, pixel value information of the first image, which is the input image, is set as known information, and contour information of the region of interest in the first image is set as unknown information. That is, the problem setting is changed so that the information on the pixel values of the image in Non-Patent Document 2 is set as pixel value information a, and the information representing the posture of the object is set as contour information b, and the proposition of this step is solved based on the method shown below.
[0062] The contour information estimation unit 53 calculates a vector f including elements corresponding to pixel value information a of the first image, which is known information, and elements corresponding to contour information b of the region of interest in the first image, which is unknown information, using the following formula. Because the calculation based on the following formula does not involve repeated calculations, it is possible to achieve contour information estimation processing with lower calculation costs than conventional gradient methods. Furthermore, since the factors that affect the variation in calculation time for this calculation are the input pixel value information and the number of dimensions of the contour information to be estimated, it is expected that the variation in processing time will be small once the size of the input image is determined.
number
[0063] Here, f′ on the right side represents input data with a certain defect, and in this embodiment, as shown in equation (3), it is a vector in which the elements of vector f corresponding to the contour information, which is unknown information, are set to 0.
number
[0064] E is a matrix representing the subspace defined by the first subspace information selected in step S103. When L eigenvectors are e1, e2, . . . , eL, the matrix E is given by, for example, E = [e1, e2, . . . , eL]. Σ is a square matrix in which diagonal elements corresponding to pixel value information, which is known information, are set to 1 and the other elements are set to 0. In other words, Σ is a matrix in which elements of a unit matrix corresponding to contour information, which is unknown information, are set to 0. In this embodiment, Σ is a square matrix with a side length of Nx × Ny + 9 × 2 (the number of dimensions of pixel value information a + the number of dimensions of contour information b), in which the last 18 diagonal elements are 0 and the other diagonal elements are 1.
number
[0065] In this embodiment, even if the input image is arbitrary, the part set as unknown information does not change (it is always the part of the contour information b). Therefore, at the time when the subspace information is calculated in step S102, E (E T ΣE) -1 E T (Part of the calculations performed in the estimation process) can be calculated in advance. T ΣE) -1 E T The calculation results may be stored as subspace information in a storage device (for example, the database 22 or the storage unit 34). T ΣE) -1 E T If the result of E(E T ΣE) -1 E T Some of them (for example, E T ΣE) -1The calculation results of the above part (part (a)) may be stored in a storage device (for example, the database 22 or the storage unit 34) as subspace information, and the calculation of the remaining part may be performed by the subspace information acquisition unit 52 or the contour information estimation unit 53.
[0066] Finally, the contour information estimation unit 53 extracts information corresponding to contour information from the estimation result f (in this embodiment, the last 18 elements) and stores it in the RAM 33.
[0067] If pixel value information a' is used instead of pixel value information a in step S102, it is necessary to replace I(0,0) to I(Nx,Ny) in equation (6) with values calculated using the same method as the pixel value information a'. More specifically, if pixel value information a' is used for learning, In the case of principal component scores related to pixel value information of an image, a projection process is performed based on the input image and information on a subspace constructed using only pixel value information of the training data to calculate the principal component scores of the input image. Furthermore, when estimation process is performed on values such as contour information b' as the target of estimation process, the contour information of the estimated result is not in xy coordinates on the contour line, so the estimated values are converted into coordinate values and stored in RAM 33. More specifically, in the case of contour information b' being principal component scores related to the contour information of the training data, a back projection process is performed based on the principal component scores and information on a subspace constructed using only the contour information of the training data.
[0068] The contour information of the estimated region of interest may be used as a contour estimation result as is, or may be used as a rough extraction result to be provided as an initial value for a more precise extraction process. In the latter case, the contour information estimation unit 53 uses the contour information obtained by the above process as an initial value and performs precise contour extraction using a known method for precisely extracting an object contour.
[0069] (Step S105: Display image based on contour information estimation result) In step S105, the display processing unit 54 displays the first image, which is the input image, and the contour information of the region of interest estimated by the contour information estimation unit 53 within the image display area. At this time, the estimated contour information and the first image may be displayed superimposed. By performing the superimposed display, it is possible to easily visually confirm how well the estimated contour information matches the first image. In this embodiment, the contour information of the region of interest is a discrete point group obtained by sampling the contour of the region, and therefore, may be displayed after interpolating between adjacent points using a known technique such as spline interpolation.
[0070] If the purpose is to analyze or measure a region of interest, the process of step S105 is not necessarily required, and the configuration may be such that the estimated contour information is simply saved.
[0071] Furthermore, in the above description, the image is a two-dimensional image, but similar processing can be performed even if the image is a three-dimensional image (volume image).
[0072] In this embodiment, an example has been shown in which the coordinate values of the point cloud representing the left atrium are used as the contour information of the region of interest, but the contour information may also be a combination of coordinate values of point clouds representing two or more regions, such as the left ventricle, right ventricle, and right atrium in addition to the left atrium. In this case, by performing statistical analysis on not only the coordinate values of the point cloud representing the left atrium region but also the coordinate values of all point clouds of the two or more regions in step S102, the coordinate values of all point clouds of the two or more regions can be simultaneously estimated in step S104.
[0073] According to this embodiment, by selecting an optimal subspace from a plurality of subspaces according to the first image and estimating contour information of the region of interest, it is possible to provide the user with a more accurate contour extraction result. Furthermore, according to this embodiment, by estimating contour information, which is shape information of the region of interest, by a matrix operation using a matrix E representing a subspace constructed from training data, it is possible to provide the user with a more accurate contour extraction result at a lower calculation cost.
[0074] Although the present embodiment has been described above, the present invention is not limited to this embodiment, and can be modified and changed within the scope of the claims.
[0075] (Variation 1) In the above embodiment, an example was shown in which the images treated as input images or learning data were two-dimensional ultrasound images, but these images may also be images (three-dimensional spatiotemporal images) in which time-series data of two-dimensional images are connected. In other words, a two-dimensional ultrasound image video (two or more frames of two-dimensional ultrasound images) may also be handled as an input image. In this case, a plurality of frames consecutive in time may be used. A three-dimensional spatiotemporal image may be constructed using a set of images of a specific frame (phase) rather than a set of images of a frame.
[0076] Taking a moving image of an ultrasound image of the heart as an example, a 3D image composed of two time phase images, a frame image of the diastole (ED (End-Diastole) phase) and a frame image of the systole (ES (End-Systole) phase), may be used. In this case, the method of the above embodiment may be performed based on pixel value information of the two frame images of the diastole and systole and the contour information corresponding to each of the images. By performing processing similar to that of the above embodiment, when an unknown 3D image is input, contour information corresponding to all frames of the 3D image can be simultaneously estimated. Note that any known method may be used to identify the frames of the heart during diastole and systole. For example, the frames may be manually identified by a user from the moving image, or the frames may be automatically identified from information such as an electrocardiogram corresponding to the moving image.
[0077] When constructing a three-dimensional image from multiple frames as described above, in order to implement the method of the first embodiment, all pixel values of all frames that serve as input images are set as pixel value information. That is, for example, when a three-dimensional image is constructed from images of the diastole (di) and the systole (sys), the pixel value information a should be as follows:
number
[0078] where I di (x,y) represents the pixel value of the diastolic image, and I sys (x, y) represents the pixel value of the image during systole. In other words, if the input image is an image composed of a set of images of multiple frames (time phases), data obtained by concatenating pixel value information of the images of all frames can be used as the pixel value information of the input image.
[0079] Furthermore, with regard to contour information, the method of the above embodiment can be implemented by using data obtained by concatenating point cloud information representing contour information of the region of interest corresponding to each image of all frames as the contour information of the region of interest corresponding to the input image. That is, contour information b can be set as follows. Here, too, those with the suffix di are the coordinates of points representing the contour information of the region of interest in the image during diastole, and those with the suffix sys are the coordinates of points representing the contour information of the region of interest in the image during systole.
number
[0080] At this time, in step S101, the image acquisition unit 51 acquires a pair of diastolic and systolic images to be processed as the first image. In addition, in step S102, the subspace information acquisition unit 52 performs the processing described in the first embodiment on multiple pairs of diastolic and systolic images different from the first image. That is, statistical analysis is performed on data c, which is a combination of pixel value information a (Equation (8)) of the image pair and correct contour information b (Equation (9)) of the region of interest in the image pair, to acquire subspace information. Then, in step S103, the contour information estimation unit 53 estimates contour information of the region of interest in the first image. In this way, by setting the pixel value information and contour information and then performing processing similar to that of the above embodiment, contour information corresponding to all frames of an unknown three-dimensional image can be simultaneously estimated for that three-dimensional image.
[0081] As described in the above embodiment, principal component analysis may be performed on data containing only pixel value information of the image, and the principal component scores corresponding to each image may be calculated and used as new pixel value information for each image. When pixel value information is configured using values on a pixel-by-pixel basis as in Equation (8), As the number of frames in a 3D image increases, the number of dimensions of the data becomes enormous, resulting in increased computational costs. Therefore, by using principal component scores with fewer dimensions as pixel value information, computational costs can be significantly reduced. Similarly, for contour information, principal component scores calculated using principal component analysis of the coordinate values of points representing the region of interest can be used as new contour information.
[0082] In this way, extracting two or more frame images, such as diastole and systole, from a video sequence to construct a 3D image and simultaneously estimating the contour information of these two or more frame images has the following two advantages: These effects are not only obtained when combining diastole and systole images, but are also obtained by using a combination of multiple images of different time phases.
[0083] The first benefit is the effect of stabilizing the estimation accuracy of the region of interest. In other words, using multiple frames improves the robustness of the estimation process. Ultrasound diagnostic images often have variations in image quality (contrast and noise levels) between frames. That is, even if the contour of the region of interest is clearly depicted and highly visible in one frame of the video, the contour may be less visible in another frame. Therefore, when extracting a frame from a video and performing region estimation processing, a low-quality frame may be the target of processing. However, performing region extraction processing on a low-quality image may result in reduced extraction accuracy. In such cases, a method can be considered to improve the region extraction accuracy of a low-quality image by using a high-quality image that is correlated with the low-quality image. While it is necessary for the input 3D image to include high-quality frames, performing region extraction processing based on multiple frames is expected to improve the average extraction accuracy and reduce variation in extraction accuracy. In other words, improved robustness of the estimation process can be expected.
[0084] The second advantage is that it enables region estimation processing that takes into account the relationship between the position and size of regions of interest between different frames. Taking an example of an image capturing cardiac movement, it is conceivable that there is a correlation between the size of the atria and ventricles and the locations of their anatomical structures in the image between diastole and systole. However, if region-of-interest estimation processing is performed independently for different frames, the processing is performed without considering the relationship between the frames, which may result in a loss of relationship between frames, such as the size relationship between the estimated regions. For example, even if the area of a region of interest during diastole should be larger than that during systole, the estimated result during diastole may be smaller. By expanding the above embodiment to handle images from multiple frames, a subspace that takes into account the relationship between regions of interest between different frames can be obtained, thereby obtaining output results that maintain the original relationship between the frames for the estimation results between different frames.
[0085] In this modification, contour information of the region of interest is estimated by matrix calculation using a matrix representing a subspace constructed from training data, so that highly accurate contour extraction results can be provided to the user at a lower computational cost. Furthermore, the combinations constituting the three-dimensional image may be not only combinations of different frames (phases) in a single examination, but also combinations of past and current images. Furthermore, not only combinations of different times, but also combinations of two-dimensional images in multiple cross-sectional directions (for example, four-chamber and two-chamber views, or long-axis and short-axis views) may be used. Furthermore, combinations of cross-sectional images in multiple frames and directions may also be used.
[0086] (Variation 2) In the above embodiment and modification 1, pixel value information of an input image is set as known information, and contour information of a region of interest corresponding to the input image is set as unknown information, and the unknown information is estimated by the BPLP method. However, the known information and the unknown information may be set differently. .
[0087] For example, in the above embodiment, if at least part of the contour information of the region of interest is known in advance in addition to the pixel value information of the input image, the contour information may also be set as known information, and the remaining contour information may be set as unknown information. For example, the position of the annulus in the region of interest may be detected using a method such as known feature point extraction, and this position may be set as known information. Alternatively, the user may manually set the position of the annulus on the image via the operation unit 35. For example, if point p1 in the contour information b is the annulus position, (x1, y1) can be given as known information. That is, in step S103, the contour information estimation unit 53 assigns the coordinates of the acquired annulus position to the elements corresponding to x1 and y1 in equation (6), changes the values of the diagonal elements corresponding to x1 and y1 in equation (7) to 1, and then estimates the coordinates of p2 to p9 using equation (5).
[0088] In addition, in the first modification, pixel value information of a three-dimensional image is set as known information, and contour information of a region of interest corresponding to each frame of the three-dimensional image is set as unknown information, and the unknown information is estimated using the BPLP method. Even in this case, if at least part of the contour information of a region of interest in a certain frame is known, the contour information may be set as known information. For example, when a three-dimensional image composed of images in the diastole and systole is input, and the contour information in the diastole is known in advance or specified by the user, only the contour information in the systole may be set as unknown information. Increasing the amount of known information in this way increases the information that can be used in the estimation process, thereby improving the accuracy of the estimation process for the unknown information.
[0089] In another example, pixel value information of an image may be estimated by the BPLP method based on contour information input by a user via the operation unit 35 or contour information generated using subspace information of the contour information of a region of interest, etc. In this case, by setting the contour information as known information and the pixel value information of the image as unknown information, estimation processing can be performed using the same procedure as in the above embodiment.
[0090] In this modified example, unknown information is estimated by matrix calculation using a matrix representing a subspace constructed from training data, so the effect of being able to provide estimation results to the user at low computational cost is maintained.
[0091] The above is one example of an embodiment, but the present invention is not limited to the embodiment described above and shown in the drawings, and can be practiced by making appropriate modifications within the scope that does not change the gist of the present invention.
[0092] <Other embodiments> Furthermore, the disclosed technology can be embodied as, for example, a system, a device, a method, a program, or a recording medium (storage medium), etc. Specifically, it may be applied to a system consisting of multiple devices (for example, a host computer, an interface device, an imaging device, a web application, etc.), or it may be applied to an apparatus consisting of a single device.
[0093] Needless to say, the object of the present invention can be achieved by the following: That is, a recording medium (or storage medium) on which is recorded program code (computer program) of software that realizes the functions of the above-described embodiments is supplied to a system or device. The storage medium is, of course, a computer-readable storage medium. Then, the computer (or CPU or MPU) of the system or device reads and executes the program code stored on the recording medium. In this case, the program code itself read from the recording medium realizes the functions of the above-described embodiments, and the program The recording medium on which the program code is recorded constitutes the present invention. [Explanation of symbols]
[0094] 10 Image processing device 51 Image acquisition unit 52 Subspace information acquisition unit 53 Contour information estimation unit 54 Display processing section 55 Subspace information selection unit
Claims
1. An image processing device that estimates shape information, which is information about the shape of a region of interest, from an image, image acquisition means for acquiring a first image to be processed; a subspace information acquisition means for acquiring information on two or more subspaces generated using two or more different data sets, the subspaces including sample data including pixel value information, which is data that lists principal component scores of the images, and shape information of a region of interest, obtained by projecting each of the plurality of images for training onto a subspace obtained by performing principal component analysis on the plurality of images for training; a subspace information selection means for selecting information of a first subspace from information of the two or more subspaces based on pixel value information of the first image; and an estimation unit that estimates shape information of a region of interest in the first image from pixel value information of the first image using information of the first subspace.
2. The image processing apparatus according to claim 1 , wherein the information on the subspace represents a distribution of sample data included in the data set.
3. 3. The image processing device according to claim 1, wherein the subspace information selection means identifies a dataset including sample data similar to the first image based on similarity between pixel value information of the first image and pixel value information of sample data included in each dataset, and selects information of a subspace generated from the identified dataset as information of the first subspace.
4. 4. The image processing apparatus according to claim 3, wherein the subspace information selection means evaluates the similarity between the first image and each data set based on a reconstruction error when pixel value information of the first image is projected and backprojected onto a subspace related to pixel value information of each data set. 。
5. 4. The image processing device according to claim 1, wherein the information on the subspace is information obtained by performing statistical analysis on a plurality of sample data included in the data set.
6. the statistical analysis is principal component analysis; 6. The image processing apparatus according to claim 5, wherein the information on the subspace includes information on a mean vector and a plurality of eigenvectors obtained by principal component analysis.
7. The image processing device according to any one of claims 1 to 6, characterized in that the subspace information acquisition means acquires information on the two or more subspaces from a storage device that stores information on pre-generated subspaces.
8. 8. The image processing device according to claim 1, wherein the first image is an image configured from a set of a plurality of images at different time phases.
9. 9. The image processing device according to claim 8, wherein the estimation means uses data obtained by concatenating pixel value information of the plurality of images at different time phases as pixel value information of the first image, and uses data obtained by concatenating shape information corresponding to each of the plurality of images at different time phases as shape information corresponding to the first image.
10. the first image is an image of the heart; 10. The image processing device according to claim 1, wherein the region of interest is an atrium or a ventricle.
11. 11. The image processing apparatus according to claim 10, wherein the first image is an image consisting of a set of an image of the heart in a diastolic period and an image of the heart in a systolic period.
12. 12. The image processing apparatus according to claim 1, wherein the pixel value information is data in which pixel values of an image are arranged.
13. 13. The image processing device according to claim 1, wherein the shape information is contour information of the region of interest.
14. An image processing method for estimating shape information, which is information about the shape of a region of interest, from an image, comprising: an image acquisition step of acquiring a first image to be processed; a subspace information acquisition step of acquiring information on two or more subspaces generated using two or more different datasets, the subspaces including sample data including pixel value information, which is data that lists principal component scores of the images, and shape information of the region of interest, obtained by projecting each of the plurality of images for training onto a subspace obtained by performing principal component analysis on the plurality of images for training; a subspace information selection step of selecting information of a first subspace from information of the two or more subspaces based on pixel value information of the first image; an estimation step of estimating shape information of a region of interest in the first image from pixel value information of the first image using information of the first subspace.
15. A program for causing a computer to execute each step of the image processing method according to claim 14.
Citation Information
Patent Citations
Image pattern identification apparatus
JP2012022540A
Method and System for Precise Segmentation of the Left Atrium in C-Arm Computed Tomography Volumes
US20130129170A1