Information processing apparatus and information processing method
The information processing device accurately estimates the ultrasound probe's orientation using virtual feature points, overcoming occlusion and structural complexity issues in existing methods.
Patent Information
- Application Number
- JP2024130995
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-07
- Publication Date
- 2026-02-20
AI Technical Summary
Existing ultrasound probe orientation estimation methods face challenges such as occlusion when using cameras and markers, and adding magnetic sensors complicates the probe's structure.
An information processing device that estimates the orientation of an ultrasound probe by acquiring images, estimating the positions of pre-set virtual feature points on the probe, and determining its attitude based on these points without additional sensors.
Accurately estimates the orientation of the ultrasound probe with high precision without the need for additional sensors, addressing occlusion issues and simplifying the probe's structure.
Smart Images

Figure 2026028510000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to an information processing device and an information processing method. [Background technology]
[0002] In an examination using an ultrasound diagnostic device, an ultrasound probe, which is an examination device, is placed on the surface of the subject's body to obtain an ultrasound image of the inside of the subject. In order to identify which part of the subject is being examined during the examination, a technique is known in which the spatial coordinates of the examination device are sensed when the subject is scanned.
[0003] In the technology disclosed in Patent Document 1, the position and orientation of the ultrasound probe are detected using a camera and a human-shaped marker attached to the inspection device.
[0004] In the technology disclosed in Patent Document 2, the position and orientation of the ultrasonic probe are detected using a magnetic sensor, an acceleration sensor, and the like. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Japanese Patent Application Publication No. 2020-127629 [Patent Document 2] JP 2019-136273 A Summary of the Invention [Problem to be solved by the invention]
[0006] However, when using the camera and marker disclosed in Patent Document 1, if the marker is hidden by another object, i.e., occlusion occurs, it is not possible to estimate the posture. Furthermore, if an additional sensor such as a magnetic sensor is provided to solve the problem of occlusion as disclosed in Patent Document 2, the structure of the ultrasound probe becomes complicated.
[0007] Therefore, an object of the present invention is to provide an information processing device that can accurately estimate the orientation of an inspection device without providing a sensor on the inspection device. [Means for solving the problem]
[0008] In order to achieve the above object, the information processing device of the present invention includes an acquisition means for acquiring an image of an inspection device, a position estimation means for estimating positions of a plurality of virtual feature points that have been set in advance on the surface of the inspection device based on the image acquired by the acquisition means, and an attitude estimation means for estimating an attitude of the inspection device based on the positions of the plurality of virtual feature points estimated by the position estimation means. [Effects of the Invention]
[0009] According to the present invention, the orientation of the inspection device can be estimated with high accuracy without providing a sensor on the inspection device. [Brief explanation of the drawings]
[0010] [Figure 1] 1 is a diagram illustrating an example of the configuration of an ultrasound diagnostic apparatus according to a first embodiment. [Figure 2] 1 is a diagram showing an installation configuration of an ultrasound diagnostic apparatus according to a first embodiment. [Figure 3] 10 is a flowchart illustrating a processing flow of position and orientation estimation according to the first embodiment. [Figure 4] 3A and 3B are diagrams illustrating virtual feature points of the ultrasound probe according to the first embodiment. [Figure 5] FIG. 2 is a diagram illustrating learning data of a deep learning model according to the first embodiment. [Figure 6] 5A to 5C are diagrams illustrating coordinate estimation results of virtual feature points according to the first embodiment. [Figure 7] 5A to 5C are diagrams illustrating coordinate estimation results of virtual feature points according to the first embodiment. [Figure 8] FIG. 1 is a diagram illustrating an example of the configuration of an ultrasound probe according to a first embodiment. [Figure 9] FIG. 2 is a diagram showing a display example of an ultrasound diagnostic image according to the first embodiment. [Figure 10] FIG. 1 is a diagram illustrating an example of the configuration of an ultrasound probe according to a first embodiment. [Figure 11] 10 is a flowchart showing a processing flow of position and orientation estimation according to Modification 2 of the first embodiment. [Figure 12] 10 is a flowchart showing an automatic identification flow of an ultrasound probe according to a second modification of the first embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0011] Hereinafter, embodiments of the present invention will be described in detail.
[0012] [First embodiment] A first embodiment will be described below, in which an example of an ultrasound diagnostic device as a medical image diagnostic device including an ultrasound diagnostic device main body as an information processing device will be described.
[0013] (Configuration of Ultrasound Diagnostic Equipment) Fig. 1 is a block diagram showing an example of the configuration of an ultrasound diagnostic apparatus 100 as a medical image diagnostic apparatus according to the first embodiment. As shown in Fig. 1, the ultrasound diagnostic apparatus 100 includes an ultrasound diagnostic apparatus main body 1 as an information processing device, an ultrasound probe 2 as an examination device, a camera 3 as an imaging means, a display 4 as a display means, and a control panel 5 as an instruction means. The ultrasound diagnostic apparatus main body 1 has a housing containing various control units, a power supply, and a computer equipped with a communication interface as an information processing device.
[0014] The ultrasound diagnostic device main body 1 is equipped with an internal bus 15, a transmitting / receiving circuit 11, a signal processing circuit 12, an image generating circuit 13, a camera control circuit 14, a processing circuit 6, a memory 7, a non-volatile memory 8 as a storage medium, a communication interface 9, and a power supply 10. Each component connected to the internal bus 15 is configured to be able to exchange data with each other via the internal bus 15.
[0015] (ultrasonic probe) The ultrasonic probe 2 is an example of an inspection device according to this embodiment, and is an ultrasonic probe that transmits and receives ultrasonic waves with its tip end in contact with the surface of a subject. The ultrasonic probe 2 has multiple piezoelectric vibrators and is connected to the ultrasonic diagnostic apparatus main body 1. The ultrasonic probe 2 generates ultrasonic waves using the multiple piezoelectric vibrators based on control signals supplied from the ultrasonic diagnostic apparatus main body 1, receives reflected waves from the subject, and converts them into electrical signals (echo signals). For example, the ultrasonic probe 2 may be any type of ultrasonic probe, such as a sector type, linear type, or convex type. The ultrasonic probe 2 may also be a one-dimensional ultrasonic probe in which multiple piezoelectric vibrators are arranged in a row. Alternatively, the ultrasonic probe 2 may be an ultrasonic probe in which multiple piezoelectric vibrators of the one-dimensional ultrasonic probe are mechanically oscillated. Alternatively, the ultrasonic probe may be a two-dimensional ultrasonic probe in which multiple piezoelectric vibrators are arranged in a two-dimensional lattice pattern.
[0016] (camera) FIG. 2 is a diagram showing the installation configuration of an ultrasound diagnostic apparatus main body 1, an ultrasound probe 2, a camera 3, a display 4, and a control panel 5 according to this embodiment. In this embodiment, the camera 3 is mainly used to acquire an external image for identifying an examination site when examining a subject using the ultrasound probe 2. Specifically, the camera 3 captures an external image including the examination site and the ultrasound probe 2 when examining a subject using the ultrasound probe 2. The camera 3 is installed, for example, at the tip of an arm attached to the ultrasound diagnostic apparatus main body 1, and can be used to capture images of the surroundings of the ultrasound diagnostic apparatus 100. The camera 3 can also be installed separately from the ultrasound diagnostic apparatus main body 1, such as on the ceiling, and may be installed anywhere as long as it can capture images of the ultrasound probe 2, which is the subject of the examination.
[0017] The camera 3 has the configuration of a typical camera, including an imaging optical system, an imaging element, a CPU, an image processing circuit, ROM, RAM, and at least one communication I / F. Light beams from a subject are focused on an imaging element such as a CCD or CMOS sensor by an imaging optical system including optical elements such as lenses, thereby capturing an image. The imaging optical system includes a lens group, and the camera 3 may also include a lens drive control circuit that controls zoom and focus by driving the lens group along the optical axis. The electrical signal output from the imaging element is converted into digital image data by an A / D converter, and various image processes are performed by an image processing circuit before being output to an external device. At least a portion of the image processing performed by the image processing circuit may be output to the external device via the communication I / F and then processed by a processing circuit in the external device.
[0018] In this embodiment, the camera 3 uses an imaging element that mainly receives light in the visible light region to capture an image. However, examples of the camera 3 are not limited to this, and the camera 3 may be a camera that receives light in the infrared light region to capture an image, or may be equipped with a camera that receives light in multiple wavelength regions such as visible light and infrared light to capture an image. Furthermore, the camera 3 may be a stereo camera that can measure distance in addition to external images, or a camera equipped with a TOF (Time Of Flight) sensor for distance measurement. Hereinafter, the image captured by the camera 3 will be referred to as a camera image.
[0019] (display device) The display 4 includes a display device such as an LCD, and displays images input from the ultrasound diagnostic apparatus main body 1, menu screens, a graphical user interface (GUI), and the like. Specifically, the display 4 displays images stored in the memory of the ultrasound diagnostic apparatus main body 1 and images recorded in the non-volatile memory on the display device. The display 4 also displays ultrasound images, camera images, body mark images, probe mark images, and body part identification results. Here, the body mark image is an image that simply represents the shape of the body and is commonly used in ultrasound diagnostic apparatuses. The probe mark image is a mark superimposed on the body mark image and is provided for the purpose of identifying at a glance the angle at which the ultrasound probe 2 is in contact with the tangent plane of the body.
[0020] (User Interface) The control panel 5 is composed of a keyboard, a trackball, switches, dials, a touch panel, etc. The control panel 5 accepts various input operations by the examiner using these operating members, such as instructions to perform imaging using the ultrasound probe 2 or the camera 3, instructions to display various images, image switching, mode designation, and various setting instructions. The accepted input operation signals are input to the ultrasound diagnostic device main body 1 and reflected in various controls. If the control panel 5 is a touch panel, it may be integrated with the display 4, and the examiner can perform various settings and operations of the ultrasound diagnostic device main body 1 by touching and dragging buttons displayed on the display.
[0021] When the examiner operates the freeze button while a signal is received from the ultrasound probe 2 and the ultrasound image in the memory of the ultrasound diagnostic device main body 1 is being updated, the signal from the ultrasound probe 2 stops, and the update of the ultrasound image memory is temporarily stopped. At this time, the signal from the camera 3 may also stop, and the update of the camera image memory may be temporarily stopped. For example, when the freeze button is operated while the update of the camera image memory has been stopped, a signal is received again from the ultrasound probe 2, and the update of the ultrasound image memory is started, and the update of the camera image is also started in the same way. When the examiner operates the confirm button while a single ultrasound image has been fixed by pressing the freeze button, the ultrasound image is saved in non-volatile memory. The freeze button and confirm button may be provided on the control panel 5 instead of the ultrasound probe 2.
[0022] (Control circuit) In the ultrasound diagnostic apparatus 100, each processing function is stored in the form of a computer-executable program in a nonvolatile memory 8 serving as a storage medium. The transmitting and receiving circuitry 11, the signal processing circuitry 12, the image generating circuitry 13, the camera control circuitry 14, and the processing circuitry 6 are processors that realize the function corresponding to each program by reading and executing the program from the nonvolatile memory 8. In other words, each circuit that has read each program has the function corresponding to the read program.
[0023] The term "processor" used in the above description refers to circuits such as a central processing unit (CPU) or a graphics processing unit (GPU). Alternatively, the term "processor" refers to circuits such as an application specific integrated circuit (ASIC), a programmable logic device (e.g., a simple programmable logic device (SPLD), a complex programmable logic device (CPLD), and a field programmable gate array (FPGA)). A processor realizes its functions by reading and executing a program stored in a memory. Instead of storing a program in a memory, the processor may be configured so that the program is directly embedded in the circuit. In this case, the processor realizes its functions by reading and executing the program embedded in the circuit. Note that each processor in this embodiment is not limited to being configured as a single circuit, but may also be configured as a single processor by combining multiple independent circuits.
[0024] The memory 7 is, for example, a RAM (such as a volatile memory using semiconductor elements). The processing circuit 6 controls each component of the ultrasound diagnostic apparatus main body 1 using the memory 7 as a work memory in accordance with, for example, a program stored in the non-volatile memory 8. The processing circuit 6 has a control function 6a, an image acquisition function 6b, and an estimation function 6c. The control function 6a performs various processes related to capturing ultrasound images and displaying the ultrasound images. The image acquisition function 6b performs various processes related to acquiring images captured by the camera. The estimation function 6c estimates the position and orientation of the ultrasound probe 2 using the camera image.
[0025] The nonvolatile memory 8 stores image data, subject data, and various programs for operating each circuit including the processing circuit 6. The nonvolatile memory 8 is configured by, for example, a hard disk or a read-only memory (ROM).
[0026] The transmission / reception circuit 11 has at least one communication interface for supplying power to the ultrasonic probe 2, transmitting control signals, receiving echo signals, etc. The transmission / reception circuit 12 supplies a control signal for causing the ultrasonic probe 2 to transmit an ultrasonic beam based on a control signal from the processing circuit 6, for example. Furthermore, the transmission / reception circuit 12 receives reflected wave signals, i.e., echo signals, from the ultrasonic probe 2, performs phased addition on the received signals, and outputs the signal obtained by the phased addition to the signal processing circuit 12.
[0027] The signal processing circuit 12 includes a B-mode processing circuit, a Doppler mode processing circuit, a color Doppler mode processing circuit, etc. The B-mode processing circuit visualizes amplitude information of the received signals supplied from the transmitting / receiving circuit 11 using known processing to generate B-mode signal data. The Doppler mode processing circuit extracts Doppler shift frequency components from the received signals supplied from the transmitting / receiving circuit 11 using known processing and further performs FFT (Fast Fourier Transform) processing, etc. to generate Doppler signal data of blood flow information. The color Doppler mode processing circuit visualizes blood flow information based on the received signals supplied from the transmitting / receiving circuit 11 using known processing to generate color Doppler mode signal data. The signal processing circuit 12 outputs the various generated data to the image generation circuit 13.
[0028] The image generation circuit 13 generates two-dimensional or three-dimensional ultrasound images of the scan area by known processing based on the data supplied from the signal processing circuit 12. For example, the image generation circuit 13 generates volume data of the scan area from the supplied data. From the generated volume data, the image generation circuit 13 generates two-dimensional ultrasound image data by MPR processing (multiplanar reconstruction) and three-dimensional ultrasound image data by volume rendering processing. Examples of ultrasound images include B-mode images, Doppler mode images, color Doppler mode images, and M-mode images.
[0029] The camera control circuit 14 has at least one communication interface for supplying power to the camera 3, transmitting and receiving control signals, and transmitting and receiving image signals. The camera 3 may have a power supply for standalone operation without receiving power from the ultrasound diagnostic device main body 1. The camera control circuit 14 can also control various imaging parameters of the camera 3, such as zoom, focus, and aperture value, by transmitting control signals to the camera 3 via the communication interface. The camera 3 may also be equipped with a pan head capable of automatic pan-tilt, receive pan-tilt control signals, and be configured to be able to control the position and orientation by pan-tilt driving. The configuration of the ultrasound diagnostic device 100 according to this embodiment has been described above. With this configuration, the inference function 6c estimates the position and orientation of the ultrasound probe 2 based on the image acquired by the camera 3.
[0030] (Ultrasound probe position and orientation estimation) Figure 3 is a flowchart showing the process of estimating the position and orientation of the ultrasound probe 2 using the estimation function 6c. Video images are acquired from the camera at each time, and the estimation results of the position and orientation of the ultrasound probe are determined using the flowchart in Figure 3, and the position and orientation of the ultrasound probe at that time are stored in memory 7. In Figure 3 and the following explanations of other figures that represent processing flows, the symbol "S" means a step.
[0031] The position and orientation estimation flow in FIG. 3 starts when the user inputs an instruction to start an examination, release the examination from freezing, etc., through the control panel 5.
[0032] In step S201, the image acquisition function 6b as an acquisition means acquires an external image (image data) including the ultrasound probe 2 from the camera 3 via the communication interface of the camera control circuit 14. Since the angle of view of the camera 3 has been adjusted in the previous step so as to include the subject, the image may be captured as is, but at least one of pan control, tilt control, zoom control, etc. may be performed so that the angle of view is more easily detected by the ultrasound probe 2.
[0033] (Virtual feature points on the ultrasound probe) 4 is a diagram showing the appearance of the ultrasound probe 2, the cable 2a connecting the ultrasound diagnostic apparatus main body 1 and the ultrasound probe 2, and a virtual feature point K set in advance at an arbitrary position on the ultrasound probe 2. A plurality of black dots drawn on the ultrasound probe 2 represent a plurality of virtual feature points K (K1, K2, ... KN) (N is the total number of set feature points). Information on the virtual feature points K is stored in the non-volatile memory 8. The information on the virtual feature points K here refers to three-dimensional coordinate information (Xi, Yi, Zi) (i = 1, 2, ... N) of each point (K1, K2, ... KN) of the virtual feature points K relative to an arbitrary three-dimensional coordinate origin (0,0,0).
[0034] (Deep learning model) In step S202, a coordinate estimation process (position estimation step) is performed to estimate the coordinates of any virtual feature points set on the ultrasound probe from the input image. Each feature point is estimated using a trained deep learning model that outputs a two-dimensional array, such as an autoencoder.
[0035] (Training data) Training data is required to train a deep learning model. The training data here refers to training image data showing the image of the ultrasound probe 2 and annotation data that indicates the coordinates of the virtual feature point K in each image. The training image data and annotation data can be obtained, for example, by using computer graphics (CG). Using CG makes it possible to create a wide variety of training image data and, at the same time, to accurately calculate the coordinates of the virtual feature point K on the image.
[0036] It is also possible to acquire learning data by capturing learning image data of the ultrasound probe 2 with a camera and simultaneously determining the position and orientation directly by attaching an AR marker to the ultrasound probe 2. It is also possible to calculate the position and orientation in advance by attaching an inertial sensor with a built-in gyro sensor and acceleration sensor to the ultrasound probe.
[0037] The training image data is array data of size H x W x C, which is the image of the ultrasound probe 2. Here, H is the height of the image, W is the width of the image, and C is the number of color channels (1 for a grayscale image, 3 for an RGB color image).
[0038] The annotation data used for training is, for example, an H x W x N array of data when the training image data size is as described above. Here, the number of channels, N, is the total number of virtual keypoints mentioned above. In other words, the annotation data is a collection of two-dimensional arrays with the same height and width as the training image data, totaling N virtual feature points. Each channel of the annotation data, 1 to N, represents the position of feature points K1 to KN, respectively. Figure 5 shows the position of virtual feature point K in the training image data and the annotation data corresponding to the i-th feature point (annotation data channel i (i = 1 to N)). When the coordinates of feature point Ki in the training image data are (xi, yi), channel i of the annotation data becomes an array of numerical values (heat map) that decays concentrically from the peak at coordinate (xi, yi). A heat map is a two-dimensional array with values from 0 to 1, with the peak value being 1 and values far enough away from the peak being 0. The annotation data can be considered a collection of such heat maps centered on the position of feature point Ki.
[0039] The height and width of the annotation data array size can be changed to any size as long as the number of channels, N, is fixed. For example, by setting the array size of the annotation data to H / 2, W / 2, or N, the number of parameters in the deep learning model can be reduced compared to when the height and width are the same as those of the training image data. As a result, it is expected that the time required to train the model and the calculation time during inference will be reduced.
[0040] (Coordinate estimation processing) In step S202, the camera image acquired in S201 is input into the above-mentioned deep learning model, and processing is performed to estimate the coordinates of the feature points of the ultrasound probe 2. The output of the model is H x W x N array data, similar to the annotation data. Channel i (i = 1 to N) is a heat map representing the probability of existence of feature point i (0 to 1), and the larger the value, the higher the probability that the feature point exists at the coordinate. In step S203, peak detection is performed from each channel using any method, and the coordinates of virtual feature point K on the image are estimated.
[0041] The feature points that can be estimated by peak detection are only a portion of the N virtual feature points K. It may be impossible to predict the coordinates of feature points located on the back side of the object relative to the camera or hidden by another object, such as a human hand. Figures 6 and 7 show heat maps of highly reliable feature points and unreliable feature points, respectively. In peak detection in step S202, a threshold can be set to determine whether a coordinate value indicates a peak. For example, assume that the reliability threshold for determining a peak is 0.4. If, as a result of peak detection, the coordinate value indicating the peak value is 0.9 in Figure 6 and 0.3 in Figure 7, the coordinates of the feature point at that index are estimated in Figure 6, but not in Figure 7. As described above, in step S202, the estimation function 6c as a position estimation means calculates the reliability of the virtual feature point and estimates its coordinates based on the reliability.
[0042] Hereinafter, the position of a virtual feature point K estimated from an image is referred to as an estimated feature point Pj (1≦j≦N), and the set of indexes j of the estimated feature points is referred to as J. The coordinates of the multiple virtual feature points K (estimated feature points Pj) estimated by this coordinate estimation process are two-dimensional coordinates.
[0043] The feature points estimated in step S202 do not include the coordinates of all of the pre-defined virtual feature points K. It may be impossible to predict the coordinates of feature points located behind the object relative to the camera or hidden by another object, such as a human hand. To complement the coordinates of feature points that could not be predicted, the process of step S203 can include any inter-frame processing using inference results from past times. For example, if the coordinates of feature point K at time t = t0 cannot be inferred, they can be complemented by using the coordinates at time t = t0-1. As another method, filtering such as a Kalman filter can be performed using statistical information on coordinate values from the start time t = 0 to t0-1 of camera image acquisition and a physical model. In this way, the coordinates of at least some of the multiple virtual feature points K on the image are identified in step S202.
[0044] The image for which coordinate estimation is performed in step S202 may be the camera image acquired in S201, or may be an image that has undergone any image processing, such as edge enhancement or noise reduction filtering. Alternatively, an ROI including the ultrasound probe can be selected from the camera image. The ROI may be manually selected by an operator using the control panel 5 or a touch panel function on the display 4. Alternatively, an area including the ultrasound probe can be detected by image processing, and an ROI can be adaptively determined for each frame at each time. Image processing here may be performed using a method such as template matching that uses known image features, or any object detection method using deep learning. When performing the coordinate estimation process in step S202 and the process of identifying feature point coordinates in step S203, the processing circuit 6 functions as a position estimation means. The above-described process can be considered as a process in which the position estimation means estimates the positions of multiple virtual feature points previously set on the surface of the ultrasound probe 2 as an inspection device, based on an image acquired by the image acquisition function 6b as an acquisition means.
[0045] (Determining coordinates by thresholding) In step S204, it is determined whether the number of feature points whose coordinates are identified in step S202 is equal to or greater than a preset threshold value Th. If the number of feature points whose coordinates are identified in step S202 is equal to or greater than the preset threshold value Th, the process proceeds to step S205.
[0046] In the feature point matching in step S205, which will be described later, the position and orientation of the ultrasound probe 2 relative to the position of the camera 3 are estimated. In estimating the position and orientation of the ultrasound probe 2, the coordinates of the estimated feature point Pj estimated in step S202 and the three-dimensional coordinate information (Xi, Yi, Zi) of each point (K1, K2, ..., KN) of the virtual feature point K are used. In this case, the threshold Th can be set to the minimum number of feature points required to solve the problem, depending on the problem setting for position and orientation identification. Furthermore, to perform position and orientation estimation more robustly, a number greater than the minimum number may be set. If the number of feature points whose coordinates are identified in step S202 is less than the preset threshold Th, step S205 is skipped, and inter-frame correction processing is performed in step S206, in which correction is performed using information from the previous time.
[0047] (Posture estimation processing) In S205, a posture estimation process (posture estimation step) is performed to estimate the position and posture of the ultrasound probe 2 relative to the position of the camera 3 in three-dimensional coordinates. The posture estimation process (posture estimation step) uses the coordinates of the estimated feature point Pj (1≦j≦N) estimated in S202 and the three-dimensional coordinate information (Xi, Yi, Zi) of each point (K1, K2, ..., KN) of the virtual feature point K. First, among the coordinate information (Xi, Yi, Zi) of the virtual feature point K, information corresponding to the index J of the estimated feature point is read into a work memory. The position and posture of the ultrasound probe 2 can be estimated by solving the Perspective-n-Point Problem (PnP problem) to find the external parameters (rotation vector, translation vector) of the camera. It is also possible to use a more advanced algorithm related to the PnP problem or an algorithm that removes outliers, such as Random Sample Consensus (RANSAC). When performing this posture estimation process, the processing circuit 6 functions as a posture estimation unit.
[0048] In step S206, the position and orientation are specified. Specifying the position and orientation is equivalent to obtaining the rotation vector and translation vector described above. In other words, in this step, the rotation vector and translation vector estimated by performing feature point matching in step S205 can be used as is.
[0049] In step S206, any inter-frame correction process can be applied. As described above, if the number of feature points is less than the preset threshold Th in step S204, inter-frame correction process is performed using information from the previous time. Also, if the estimated position and orientation have changed significantly compared to the position and orientation estimated at the previous time, it is possible to discard the estimation results at that time and apply the position and orientation information from the previous time. It is also possible to perform filtering such as a Kalman filter using statistical information on the position and orientation from the start time t=0 to t0-1 of camera image acquisition and a physical model.
[0050] In step S207, it is determined whether or not to end the position and orientation estimation of the ultrasound probe 2 instructed by the user. If the process is not to be ended, an image is captured again from the camera 3 at the next time, and new position and orientation estimation is performed based on that image. If the user inputs an instruction to end the examination or freeze the examination from the control panel 5, the position and orientation estimation ends. The above-mentioned position and orientation estimation process is also referred to as an orientation estimation step. The position and orientation estimation process can also be said to be a process in which the orientation estimation means estimates the orientation of the inspection device based on the positions of multiple virtual feature points estimated by the position estimation means.
[0051] Generally, the ultrasound probe 2 is held by the user's hand during the examination, and therefore the ultrasound probe is never captured in its entirety by the camera during the examination, and is always hidden by the user's hand, a so-called occlusion state. When detecting the position and orientation of the ultrasound probe 2 from the camera image, this type of occlusion problem is a major issue. However, in the method of determining the position and orientation by matching feature points as in this embodiment, the position and orientation can be detected as long as the number of estimated feature points P does not fall below the threshold value Th in step S204.
[0052] As described above, according to the first embodiment, the image acquisition function 6b acquires an image including the ultrasound probe 2 under examination. The estimation function 6c acquires the coordinates of at least some of the virtual feature points preset on the ultrasound probe based on the image, and estimates the position and orientation of the ultrasound probe by matching the coordinates with the three-dimensional coordinates of the virtual feature points. Therefore, the ultrasound diagnostic apparatus 100 according to the first embodiment can realize a method that is robust against occlusion and can accurately estimate the orientation of the inspection device without providing a sensor on the inspection device.
[0053] [Variation 1] (Specifying the scanning direction using orientation marks) Generally, the image display method used during ultrasound examinations has a set orientation for up, down, left, and right so that the user can see the image most naturally during the examination. Therefore, the user holds the ultrasound probe 2 in a predetermined orientation and scans the subject's body. As shown in Figure 8, the ultrasound probe 2 has a logo mark 2a that indicates the brand, an operation button 2b that can be used to switch image acquisition on and off at hand, and an orientation mark 2c with a protrusion that indicates the direction. These marks allow the user to determine the correct orientation of the ultrasound probe 2.
[0054] Figure 9 shows an example of an acquired ultrasound diagnostic image 20 displayed on the display 4. Guide marks 21 are displayed on the screen, and the user scans the ultrasound probe 2 so that the orientation mark 2c is aligned with the guide mark 21, and saves the ultrasound diagnostic image 20. Marks attached to the ultrasound probe 2, such as the logo mark 2a, guide mark 21, and orientation mark 2c, are also referred to as landmarks.
[0055] However, a user who is unfamiliar with the examination may perform the examination with the ultrasound probe 2 facing in the wrong direction, and may end up saving an image with the left and right reversed.
[0056] Figure 10 is a view of the ultrasound probe 2 in Figure 8 as seen from the side without the logo mark 2a and operation buttons 2b. As can be seen from Figures 8 and 10, the ultrasound probe 2 has a substantially symmetrical shape. Therefore, when the orientation in Figure 8 is the front and the orientation in Figure 10 is the back, if the position and orientation of the ultrasound probe 2 are obtained from the image acquired by the camera 3, the index of the feature points may erroneously estimate the front or back.
[0057] In the first embodiment, the virtual feature point K is set in advance at an arbitrary position on the ultrasound probe 2, but the feature point on the ultrasound probe 2 can also be set at a portion with asymmetric features, such as a logo mark 2a, an operation button 2b, or an orientation mark 2c. The estimation function 6c can perform processing to determine on which surface other symmetric feature points are located, depending on whether or not the coordinates of an asymmetric feature point are detected from the estimated heat map in the peak detection of step S202 in Fig. 3. This processing enables more robust estimation of the position and orientation of the ultrasound probe 2.
[0058] [Variation 2] (Switching deep learning models for multiple probes) As mentioned above, there are various types of ultrasonic probes 2, such as sector type, linear type, and convex type. These types are used depending on the purpose of the examination, and each has different frequency characteristics and the shape of the contact surface with the body surface.
[0059] The ultrasound diagnostic device 100 may be equipped with a plurality of types of ultrasound probes 2. In such a case, the estimation function 6c preferably includes a deep learning model and three-dimensional coordinate information related to the coordinate estimation means (position estimation means) and the position and orientation estimation means (orientation estimation means) for each ultrasound probe 2 having a different shape. That is, a virtual feature point K is set for each different ultrasound probe 2, and images and annotation data for learning are created. Then, the deep learning model and three-dimensional coordinate information of the virtual feature points learned for each ultrasound probe are stored in the non-volatile memory 8.
[0060] Fig. 11 is a flowchart showing the processing of Modification 2. It differs from the flowchart of Fig. 3 in that step S200 is added. As shown in Fig. 11, the estimation function 6c reads out a model and three-dimensional coordinate information for performing position and orientation estimation in step S200 according to the type of ultrasound probe 2 selected.
[0061] In step S200, the user can select a deep learning model and three-dimensional coordinate information of virtual feature points. When starting a new examination, the user selects the ultrasound probe 2 to be used according to the purpose of the examination from the control panel 5. Then, the estimation function 6c reads out the model and three-dimensional coordinate information for position and orientation estimation corresponding to the ultrasound probe 2 selected by the user.
[0062] Alternatively, in step S200, the type of ultrasound probe 2 can be automatically selected by identifying the ultrasound probe 2 from the camera image using a deep learning model and three-dimensional coordinate information of virtual feature points. For example, the ultrasound probe 2 can be identified from the camera image by performing object detection using deep learning. Specifically, a model that has learned the appearances of multiple ultrasound probes 2 from the image can be used in a deep learning device for object detection such as SSD (Single Shot Multibox Detector) or YOLO (YouLook Only Once). The learned model, such as the SSD or YOLO described above, is stored in advance in non-volatile memory 8, which serves as a storage medium.
[0063] FIG. 12 shows a flowchart for automatically identifying the type of ultrasound probe 2 using a deep learning model in step S200. First, in step S301, the estimation function 6c reads the deep learning model for identifying the type of ultrasound probe 2 from the non-volatile memory 8 to the working memory. Next, in step S302, the image acquisition function 6b acquires a camera image. Next, in step S303, the acquired camera image is used as input to perform object detection of the ultrasound probe 2 using the deep learning model, thereby identifying which ultrasound probe 2 is captured by the camera. In step S304, it is determined whether identification of the ultrasound probe 2 has been completed from the camera image. If identification has been completed, in step S305, a model and three-dimensional coordinate information for estimating the position and orientation of the ultrasound probe 2 are read out, and the process ends. If identification has not been successful, the process returns to step S302 again, and the camera image at the next time is acquired.
[0064] As described above, in this modification, the ultrasound diagnostic device 100 is provided with different deep learning models and three-dimensional coordinate information of virtual feature points for each of the multiple types of ultrasound probe 2. In other words, it can be said that the ultrasound diagnostic device main body 1 (information processing device) has multiple types of position estimation means and posture estimation means corresponding to the types of ultrasound probe 2 (examination device). This makes it possible to estimate the position and posture of the ultrasound probe 2 even when the user uses different ultrasound probes 2 depending on the purpose of the examination. [Explanation of symbols]
[0065] 1. Ultrasound diagnostic device 2 Ultrasound probes 6 Processing circuit 6a Control Functions 6b Image acquisition function 6c Estimation Function
Claims
1. an acquisition means for acquiring an image of the inspection device; a position estimation means for estimating positions of a plurality of virtual feature points preset on the surface of the inspection device based on the image acquired by the acquisition means; and an attitude estimation unit that estimates the attitude of the inspection device based on the positions of the plurality of virtual feature points estimated by the position estimation unit.
2. 2. The information processing apparatus according to claim 1, wherein said position estimating means estimates coordinates of said plurality of virtual feature points.
3. 3. The information processing device according to claim 2, wherein the coordinates of the plurality of virtual feature points are two-dimensional coordinates in the image, and the posture estimation means estimates three-dimensional coordinates of the inspection device using the two-dimensional coordinates in the image, thereby estimating the posture.
4. 3. The information processing apparatus according to claim 2, wherein the position estimation means estimates coordinates of at least some of a plurality of virtual feature points preset on the surface of the inspection device.
5. 5. The information processing apparatus according to claim 4, wherein the position estimation means calculates a reliability of the virtual feature point based on the image, and estimates the coordinates based on the reliability.
6. 6. The information processing apparatus according to claim 5, wherein said position estimation means excludes virtual feature points whose reliability is lower than a predetermined threshold value, and estimates the coordinates of virtual feature points that have not been excluded.
7. The information processing apparatus according to claim 1 , wherein the plurality of virtual feature points include virtual feature points set on landmarks of the inspection device.
8. The information processing device according to claim 7 , wherein the landmark indicates information related to a direction in which the inspection device is scanned.
9. 2. The information processing apparatus according to claim 1, further comprising a plurality of types of said position estimation means and said attitude estimation means corresponding to a plurality of types of said inspection devices.
10. 2. The information processing apparatus according to claim 1, wherein the posture estimation means further estimates the position of the inspection device relative to the imaging means that has imaged the inspection device.
11. A medical image diagnostic apparatus comprising: the inspection device; and the information processing apparatus according to claim 1 .
12. The information processing device according to claim 1 , wherein the image is a moving image.
13. The information processing apparatus according to claim 1 , wherein the inspection device is an ultrasonic probe.
14. an acquisition step of acquiring an image of the inspection device; a position estimation step of estimating positions of a plurality of virtual feature points set in advance on the surface of the inspection device based on the image acquired in the acquisition step; an attitude estimation step of estimating an attitude of the inspection device based on the positions of a plurality of virtual feature points estimated in the position estimation step.
15. A program capable of executing the information processing method according to claim 14 on a computer.
16. A computer-readable storage medium storing the program according to claim 15.
Citation Information
Patent Citations
JP136273A
Image processing device
JP2020127629A