Predicting the likelihood that an individual will have one or more diseases
By identifying and grouping regions of interest in multiple ultrasound images using machine learning, the system addresses low detection rates and computational challenges, enhancing disease prediction accuracy and reducing false positives in resource-limited environments.
Patent Information
- Application Number
- JP2023532490
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-01-22
- Filing Date
- 2021-11-29
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2041-11-29
AI Technical Summary
Existing ultrasound-based disease detection methods face challenges due to low detection rates, computational complexity, and time constraints, particularly in resource-limited environments, limiting the effectiveness of traditional and deep learning approaches for identifying lesions in ultrasound images.
A processing system that acquires multiple ultrasound images, identifies regions of interest, groups them based on similarity without spatial alignment, and uses machine learning to generate predictive indicators of disease likelihood, reducing computational burden while maintaining accuracy.
The proposed approach significantly reduces false positive rates and maintains high precision in disease identification by processing groups of regions of interest, offering a cost-effective and efficient mechanism for disease prediction in resource-constrained settings.
Smart Images

Figure 0007809114000001 
Figure 0007809114000002 
Figure 0007809114000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to the field of ultrasound imaging, and in particular to identifying disease lesions revealed in ultrasound images. [Background technology]
[0002] Ultrasound is increasingly being used in the detection and diagnosis of various cancers, such as breast cancer, thyroid cancer, and liver cancer. In particular, early detection of cancers using ultrasound is usually performed by identifying the lesion in ultrasound images. Unfortunately, the detection rate of early cancers is relatively low due to several factors, such as a shortage of sufficiently experienced sonographers and time constraints for performing ultrasound imaging screening.
[0003] Clinical decision support (CDS) systems using computer-implemented ultrasound-based lesion detection techniques have been employed to assist clinicians in performing ultrasound image screening. Such systems could assist clinicians by predicting the presence or absence of one or more lesions, thereby reducing the likelihood of missed detections and serving as a double-read / confirmation to increase diagnostic confidence.
[0004] Generally, there are two types of methods for ultrasound-based disease detection (automated or computer-implemented): traditional disease detection methods based on image processing steps and deep learning approaches. Traditional approaches are generally not considered robust and / or flexible enough for a large portion of the population because they rely on rule-based approaches and specific assumptions. Deep learning approaches, which rely less on such strong assumptions, have shown superior accuracy in object / disease detection, but suffer from high computational complexity and processing time.
[0005] In the paper "3D tumor detection in automated breast ultrasound using deep convolutional neutral network" by Li Yanfeng et al., published in Medical Physics, vol. 47, no. 11, a tumor detection method was proposed for automated breast ultrasound (ABUS). ABUS is known as an automated ultrasound device for breast screening. ABUS uses 3D ultrasound technology to acquire operator-independent volume images of the breast, generate coronal view segments from the acquired 3D volume, and display such segments. Summary of the Invention [Problem to be solved by the invention]
[0006] Thus, there is a continuing desire to improve mechanisms for predicting the likelihood that an individual will have one or more disease conditions. [Means for solving the problem]
[0007] The invention is defined by the claims.
[0008] According to an example according to one aspect of the present invention, a processing system for predicting a likelihood that an individual has one or more disease states is provided, the processing system being configured to: acquire multiple ultrasound images of the individual; identify regions of interest in the multiple ultrasound images, each region of interest being a portion of one of the multiple ultrasound images that represents an area of the individual with a potential disease state; group the regions of interest from different ultrasound images of the multiple ultrasound images based on similarities between the regions of interest; and process each group of regions of interest using a machine learning method to generate a predictive indicator of a likelihood that the group of regions of interest contains the disease state for the individual. The multiple ultrasound images include a time series of ultrasound images. In some embodiments, the time series of ultrasound images is a video, e.g., a cine-loop, of ultrasound images. Identifying the regions of interest that represent areas with a potential disease state includes identifying the presence or absence of one or more regions of interest. For some ultrasound images, there may be no identified regions of interest in the respective ultrasound images, and for some ultrasound images, there may be one or more identified regions of interest in the respective ultrasound images.
[0009] The present disclosure proposes techniques for identifying the presence or absence of one or more pathologies in an individual, and specifically for generating one or more predictive indicators of the likelihood that a pathology is present. Ultrasound images of a patient are processed to identify regions of interest (i.e., portions of the respective ultrasound images) that contain potential pathologies—i.e., candidate areas predicted to represent pathologies. A region of interest is thus a portion of an ultrasound image and is itself an image (but of smaller size than the ultrasound image). A region of interest contains a potential pathology if the area of the individual represented by the region of interest is predicted to contain—i.e., is likely to contain—the potential pathology. Mechanisms for identifying regions of interest that contain potential pathologies are well known to those skilled in the art and include, for example, machine learning methods, edge detection algorithms, image segmentation techniques, and the like.
[0010] The regions of interest are then grouped or clustered based on their similarity to produce groups of regions of interest. In this manner, portions of different ultrasound images that correspond to each other (i.e., contain the same potential pathology) are grouped together. Thus, regions of interest are considered similar to each other (and grouped accordingly) if they are predicted to contain the same potential pathology of an individual. Furthermore, the grouping or clustering is performed based on the similarity of the regions of interest without requiring the relative spatial relationship between the regions of interest. In this manner, the grouping or clustering can be performed without spatially aligning the multiple ultrasound images. This is particularly advantageous when the multiple ultrasound images are not spatially aligned. For example, when multiple two-dimensional ultrasound images are manually acquired by a handheld ultrasound probe, the relative spatial relationship between these two-dimensional ultrasound images is unknown, unlike those acquired by ABUS (automated chest ultrasound). Conventionally, such multiple ultrasound images can first be combined into a three-dimensional (3D) volume using various spatial alignment techniques (e.g., spatial registration, 3D construction), and then further signal processing (e.g., disease detection) can be performed based on the spatially aligned images or fragments. However, spatial alignment or 3D can be difficult in some scenarios. In response, it is proposed herein to group multiple regions of interest based on their similarity without relying on prior spatial registration or alignment across multiple images, in other words, without relying on knowledge of the relative spatial relationships between the multiple images or between regions of interest identified in the multiple images.
[0011] The groups are then processed using a machine learning method to generate a predictive indicator that indicates the likelihood that the group contains a disease for the individual. The machine learning method thereby effectively predicts whether a disease is present in a group of regions of interest. In this manner, a machine learning method, such as a classifier, generates a predictive indicator by processing groups of regions of interest that contain the same potential disease.
[0012] In the context of the present invention, a predictive indicator is any data that changes in response to changes in the predicted likelihood (determined by machine learning methods) that a group of diseases contains the disease. Predictive indicators include binary, categorical, or numeric data. Binary data indicates a prediction as to whether a group will contain the disease (e.g., "0" indicates a predicted no and "1" indicates a predicted yes, or vice versa). Categorical data indicates a likelihood category (e.g., "possible," "very likely," "unlikely," etc.) that a group contains the disease. Numerical data indicates a numerical probability that a group of interest contains the disease, for example, on a scale of 0 to 1 or 0 to 100.
[0013] The proposed approach reduces the false positive rate of identifying an individual's disease (e.g., from still images alone) by using multiple ultrasound images, e.g., ultrasound images taken over a period of time, thereby using additional contextual information to identify the individual's disease. By first identifying the region of interest, the amount of processing performed to identify the disease is reduced compared to, for example, running a machine learning process on the entire multiple ultrasound images without further processing. The proposed approach thereby provides a mechanism for performing high-quality and highly accurate disease identification with low computational effort.
[0014] Preferably, each ultrasound image is a two-dimensional ultrasound image. Because ultrasound systems for generating 2D ultrasound images have widespread utility and adoption and are increasingly being used in resource-limited, e.g., low-power, environments (e.g., battery-powered or in areas with unreliable power), the present invention is particularly advantageous when used to process two-dimensional ultrasound images. Therefore, reducing the computational burden (minimizing power usage requirements) while still obtaining an accurate indicator of potential disease for such systems would be particularly beneficial.
[0015] In some embodiments, the processing system is configured to process each group of regions of interest by performing a process that includes, for each group of regions of interest, generating a sequence of regions of interest using the regions of interest in the group of regions of interest, and processing the sequence of regions using machine learning methods to predict whether the sequence of regions includes a disease in the individual.
[0016] The sequence of regions of interest effectively forms a single data structure that contains the data from all regions of interest in the group.
[0017] The processing system is configured to generate a sequence of regions by performing steps including stacking the regions of interest. If the regions of interest are extracted from two-dimensional images, e.g., two-dimensional ultrasound images, this process effectively forms a pseudo-3D volume.
[0018] The processing system is configured to identify regions of interest in the ultrasound images by performing a process that includes processing each ultrasound image using a second machine learning method to identify the regions of interest. Thus, at least two machine learning methods are used to predict the absence of disease. This approach increases the likelihood of identifying potential disease and facilitates a fully automated mechanism for identifying potential disease, as well as employing existing and well-developed mechanisms for identifying potential disease in ultrasound images to improve reliability.
[0019] In some examples, a third machine learning method is used to group the regions of interest, for example, using characteristics and / or other metadata of each region of interest (e.g., location, size, shape, etc.) Of course, even if a machine learning method is not used to identify the regions of interest, the machine learning method is still used to group the regions of interest.
[0020] In some preferred examples, the plurality of ultrasound images comprises a sequence of ultrasound images. The sequence of ultrasound images is preferably time-sequential, such that later ultrasound images in the sequence are acquired by the imaging system later in time than earlier ultrasound images. This effectively results in time-dependent information being incorporated into the group of regions of interest. The present disclosure recognizes that the use of time-based information increases the chance of accurately identifying the presence or absence of one or more pathologies.
[0021] Preferably, the multiple ultrasound images are acquired by the same ultrasound imaging system, for example, using the same ultrasound imaging probe. More preferably, the multiple images are acquired while the ultrasound imaging probe (used to acquire the ultrasound images) is stationary. This increases the ease and accuracy of identifying groups of linked regions of interest.
[0022] In some embodiments, the order of ultrasound images in a sequence depends on the time at which each ultrasound image was captured. For example, a sequence of ultrasound images includes sequential frames of an ultrasound video. The context provided by a temporally ordered sequence of ultrasound images (e.g., an ultrasound video) allows time-based information to be taken into account for machine learning methods when predicting the presence or absence of one or more pathologies in a group of regions of interest. The present disclosure identifies that this information is particularly advantageous for reducing false positive detection rates of one or more pathologies in ultrasound images.
[0023] In some examples, for each group of regions of interest, each region of interest is derived from an ultrasound image that is sequentially adjacent to an ultrasound image of another region of interest in the same group of regions of interest.
[0024] Preferably, each region of interest is no greater than 0.4 times the size of the ultrasound image, for example no greater than 0.25 times the size of the ultrasound image, The smaller the size of the regions of interest, the greater the effect of reducing the computational complexity of processing a group of regions of interest (rather than the entire ultrasound image).
[0025] In some examples, the processing system is configured to display a visually perceptible output on a display device in response to each predictive indicator, such that information responsive to the predictive indicator may be output to a user (e.g., a clinician).
[0026] If the at least one predictive indicator indicates that at least one group of regions of interest is likely to contain a lesion, the processing system is configured to: identify a location of the identified lesion on a displayed ultrasound image of the patient; and superimpose a visually perceptible output on the displayed ultrasound image in response to the identified location of the lesion.
[0027] The processing system is configured to group the regions of interest by performing a process including: determining a similarity measure between different regions of interest from different ultrasound images, the similarity measure representing a size of overlap between a relative area occupied by one region of interest in the ultrasound image and a relative area occupied by another region of interest in another ultrasound image; and grouping the two different regions of interest into the same group in response to the similarity measure between the two different regions of interest exceeding a predetermined threshold.
[0028] Thus, the similarity measure is based on the size of the overlap, e.g., an Intersection over Union (IoU) measure, between regions of interest from different ultrasound images. This facilitates a simple mechanism for identifying whether the same potential lesion has been identified, for example, when the same potential lesion is likely to be in the same relative location in each ultrasound image. This approach is particularly advantageous when multiple ultrasound images include frames from an ultrasound video, as this embodiment increases the likelihood that a potential lesion will be in the same position between different frames of the ultrasound video.
[0029] The processing system is configured to group the regions of interest by performing a process that includes: identifying metadata of the regions of interest, the metadata providing information about the size, location, reliability, and / or occurrence of the regions of interest; and grouping the regions of interest based on the metadata of the regions of interest. Thus, additional information is used to determine similarities between different regions of interest. In some embodiments, machine learning methods are used to group the regions of interest, for example, based on the metadata of the regions of interest.
[0030] A computer-implemented method for predicting the likelihood that an individual has one or more diseases is also provided.
[0031] The computer-implemented method includes: acquiring a plurality of ultrasound images of an individual; identifying regions of interest in the plurality of ultrasound images, each region of interest being a portion of one of the plurality of ultrasound images that represents an area of the individual having a potential disease; grouping regions of interest from different ultrasound images of the plurality of ultrasound images based on similarities of the regions of interest; and processing each group of regions of interest using machine learning methods to generate a predictive indicator that indicates the likelihood that the group of regions of interest contains a disease in the individual.
[0032] A processing system is also proposed, comprising a memory and a processor coupled to the memory and configured to execute such a computer-implemented method.
[0033] Also proposed is a computer program product comprising computer program code means which, when executed on a computing device having a processing system, causes the processing system to perform all of the steps of any computer-implemented method described herein.
[0034] One skilled in the art can readily adapt any of the processing systems described herein to perform any of the methods described herein, and vice versa.
[0035] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiments described hereinafter.
[0036] For a better understanding of the present invention, and to show more clearly how the same may be carried into effect, reference will now be made, by way of example only, to the accompanying drawings in which: [Brief explanation of the drawings]
[0037] [Figure 1] FIG. 1 illustrates an ultrasound system for use in one embodiment. [Figure 2] FIG. 1 illustrates a workflow for contextual understanding of an embodiment. [Figure 3] FIG. 1 illustrates a method according to one embodiment. [Figure 4] FIG. 1 illustrates a method according to one embodiment. [Figure 5] FIG. 1 illustrates a method for grouping regions of interest. [Figure 6] FIG. 1 illustrates a processing system according to one embodiment. [Figure 7] FIG. 10 is a diagram illustrating the effects of the embodiment. [Figure 8] FIG. 1 illustrates a processing system. DETAILED DESCRIPTION OF THE INVENTION
[0038] The present invention will be described with reference to the drawings.
[0039] It should be understood that the detailed description and specific examples, while indicating exemplary embodiments of the devices, systems, and methods, are intended for purposes of illustration only and are not intended to limit the scope of the invention. These and other features, aspects, and advantages of the devices, systems, and methods of the present invention will become better understood from the following description, appended claims, and accompanying drawings. It should be understood that the figures are merely schematic and not drawn to scale. It should also be understood that the same reference numerals are used throughout the several views to denote the same or similar parts.
[0040] The present invention provides a mechanism for determining the probability / likelihood of the presence or absence of one or more diseases in an individual by processing multiple ultrasound images of the individual. The ultrasound images are processed to identify regions of interest within each image, each region of interest representing a portion of the ultrasound image suspected of exhibiting a disease. Similar regions of interest are grouped, and groups of similar regions are then processed using machine learning methods to generate a prediction of whether they are likely to contain the disease of interest. Thus, a two-step process for identifying possible diseases is performed.
[0041] Embodiments of the present invention are based on the recognition that important information is lost or absent when using only a single ultrasound image to identify potential diseases, but that processing multiple ultrasound images simultaneously significantly increases the processing load of the image analysis process. Instead, a new approach to assessing the likelihood of the presence or absence of potential diseases is proposed, in which groups of regions of interest from different ultrasound images are formed and processed together to predict the likelihood that a disease is present.
[0042] The embodiments can be used in any suitable ultrasound analysis system, for example, in ultrasound analysis for cancer screening and diagnosis of superficial organs. The proposed embodiments have the potential to be used in any suitable clinical environment, for example, during medical examinations for cancer screening or in hospital ultrasound departments for cancer screening and diagnosis. The present invention is applicable to all ultrasound imaging systems with advanced disease identification software for online or offline use, i.e., processing data from memory or directly generated by the ultrasound system.
[0043] The general operation of an exemplary ultrasound system will first be described with reference to Figure 1. The present invention uses ultrasound images produced by such an ultrasound system, although other techniques and systems for producing ultrasound images will be apparent to those skilled in the art.
[0044] The system includes an array transducer probe 4 having a transducer array 6 for transmitting ultrasound waves and receiving echo information. The transducer array 6 may include a CMUT transducer, a piezoelectric transducer formed from a material such as PZT or PVDF, or any other suitable transducer technology. In this example, the transducer array 6 is a two-dimensional array of transducers 8 capable of scanning either a 2D plane or a three-dimensional volume of interest. In another example, the transducer array is a 1D array.
[0045] The transducer array 6 is coupled to a microbeamformer 12 that controls the reception of signals by the transducer elements. The microbeamformer is capable of at least partial beamforming of signals received by subarrays of transducers (commonly referred to as "groups" or "patches") such as those described in U.S. Patent Nos. 5,997,479 (Savord et al.), 6,013,032 (Savord), and 6,623,432 (Powers et al.).
[0046] It should be noted that the microbeamformer is entirely optional. Additionally, the system includes a transmit / receive (T / R) switch 16 to which the microbeamformer 12 can be coupled and which switches the array between transmit and receive modes, protecting the main beamformer 20 from high-energy transmit signals when the microbeamformer is not in use and the transducer array is operated directly by the main system beamformer. The transmission of ultrasound beams from the transducer array 6 is directed by a transducer controller 18 coupled to the microbeamformer by the T / R switch 16 and the main transmit beamformer (not shown), which can receive input from a user operating a user interface or control panel 38. The controller 18 can include transmit circuitry arranged to drive the transducer elements of the array 6 (either directly or via the microbeamformer) during transmit mode.
[0047] In a typical row-by-row imaging sequence, the beamforming system in the probe operates as follows: During transmit, the beamformer (either a microbeamformer or a main system beamformer, depending on the implementation) activates the transducer array, or a subaperture of the transducer array. A subaperture is a one-dimensional line of transducers or a two-dimensional patch of transducers within a larger array. In transmit mode, the focusing and steering of the ultrasound beam generated by the array, or a subaperture of the array, is controlled as described below.
[0048] When backscattered echo signals are received from the object, the received signals undergo receive beamforming (as described below) to align the received signals, and if subapertures are used, the subapertures are then shifted, for example, by one transducer element. The shifted subaperture is then activated, and the process is repeated until all of the transducer elements of the transducer array have been activated.
[0049] For each line (or subaperture), the total receive signal used to form the associated line of the final ultrasound image is the sum of the voltage signals measured by the transducer elements of the given subaperture during the receive period. Following the beamforming process, the resulting line signals are commonly referred to as radio frequency (RF) data. Each line signal (RF data set) generated by the various subapertures then undergoes additional processing to generate the lines of the final ultrasound image. Changes in the amplitude of the line signal over time contribute to changes in brightness of the ultrasound image with depth, where high amplitude peaks will correspond to bright pixels (or collections of pixels) in the final image. Peaks appearing near the beginning of the line signal will represent echoes from shallow structures, while peaks appearing progressively later in the line signal will represent echoes from structures at increasing depths within the object.
[0050] One of the functions controlled by the transducer controller 18 is the direction in which the beam is steered and focused. The beam may be steered straight in front of the transducer array (perpendicular to the transducer array) or at a different angle for a wider field of view. The steering and focusing of the transmit beam is controlled as a function of transducer element activation time.
[0051] Two methods can be distinguished in general ultrasound data acquisition: plane wave imaging and "beam-guided" imaging. The two methods are distinguished by the presence of beamforming in the transmit ("beam-guided" imaging) and / or receive modes (plane wave imaging and "beam-guided" imaging).
[0052] Looking first at the focusing function, by activating all transducer elements simultaneously, the transducer array generates a plane wave that diverges as it passes through the object. In this case, the ultrasound beam remains unfocused. By introducing a position-dependent time delay to the transducer activation, it is possible to focus the wavefront of the beam at a desired point, called a focal zone. A focal zone is defined as the point where the lateral beam width is less than half the transmit beam width. In this way, the lateral resolution of the final ultrasound image is improved.
[0053] For example, if a time delay is used to sequentially activate the transducer elements, starting with the outermost element and ending with the central element of the transducer array, a focal band is formed at a given distance from the probe according to the central element. The distance of the focal band from the probe will vary depending on the time delay between each subsequent successive transducer element activation. After the beam passes through the focal band, the beam will begin to diverge, forming a far-field imaging region. Note that for focal bands located near the transducer array, the ultrasound beam will rapidly diverge in the far field, resulting in beamwidth artifacts in the final image. Typically, the near field, located between the transducer array and the focal band, shows little detail due to the large overlap in the ultrasound beams. Therefore, changing the location of the focal band can lead to significant changes in the quality of the final image.
[0054] Note that in transmit mode, only one focal point is defined unless the ultrasound image is divided into multiple focal zones (each with a different transmit focal point).
[0055] Additionally, when receiving echo signals from within the object, it is possible to perform the reverse of the above process to perform receive focusing. In other words, the input signals are received by the transducer elements, subject to an electronic time delay before being passed to the system for signal processing. The simplest example of this is called delay-and-sum beamforming. It is possible to dynamically adjust the receive focusing of the transducer array as a function of time.
[0056] Looking now at the function of beam steering, through the precise application of time delays to the transducer elements, it is possible to impart a desired angle for the ultrasound beam as it exits the transducer array. For example, by activating a transducer on a first side of the transducer array followed by the remaining transducers in the sequence ending on the opposite side of the array, the wavefront of the beam will be bent toward the second side. The size of the steering angle relative to the normal to the transducer array depends on the size of the time delay between subsequent transducer element activations.
[0057] Additionally, it is possible to focus the steered beam, so that the total time delay applied to each transducer element is the sum of both the focusing and steered time delays, in which case the transducer array is referred to as a stepped array.
[0058] For CMUT transducers that require a DC bias voltage for their activation, the transducer controller 18 may be coupled to control the DC bias control 45 of the transducer array, which sets the DC bias voltage applied to the CMUT transducer elements.
[0059] For each transducer element of the transducer array, an analog ultrasound signal, typically referred to as channel data, enters the system via a receive channel. In the receive channel, a partially beamformed signal is created from the channel data by the microbeamformer 12 and then passed to the main receive beamformer 20, where the partially beamformed signals from the individual transducer patches are combined into a fully beamformed signal, referred to as radio frequency (RF) data. The beamforming performed at each stage may be performed as described above or may include additional functions. For example, the main beamformer 20 may have 128 channels, each receiving partially beamformed signals from tens or hundreds of transducer element patches. In this way, signals received by thousands of transducers in the transducer array can efficiently contribute to a single beamformed signal.
[0060] The beamformed receive signals are coupled to a signal processor 22, which can process the received echo signals in various ways, such as bandpass filtering, decimation, I and Q component separation, and harmonic signal separation, which operates on separate linear and nonlinear signals to enable identification of nonlinear (higher harmonics of the fundamental frequency) echo signals returned from tissue and microbubbles. The signal processor also performs additional signal enhancements such as speckle reduction, signal combining, and noise removal. The bandpass filter in the signal processor can be a tracking filter, whose passband moves from higher to lower frequency bands as echo signals are received from increasing depths, thereby removing noise at higher frequencies from greater depths that are normally devoid of anatomical information.
[0061] The transmit and receive beamformers can be implemented in different hardware and have different functions. Of course, the receiver beamformer is designed to take into account the characteristics of the transmit beamformer. In FIG. 1, for clarity, only the receiver beamformers 12, 20 are shown. In a complete system, there would also be a transmit chain with a transmit microbeamformer and a main transmit beamformer.
[0062] The function of the microbeamformer 12 is to provide an initial combination of signals to reduce the number of analog signal paths, which is typically performed in the analog domain.
[0063] Final beamforming occurs in the main beamformer 20, typically after digitization.
[0064] The transmit and receive channels use the same transducer array 6 with a fixed frequency band. However, the bandwidth occupied by the transmit pulses can vary depending on the transmit beamforming used. The receive channel can capture the entire transducer bandwidth (which is the traditional approach), or by using bandpass processing, it can extract only the bandwidth containing the desired information (e.g., harmonics of the main harmonic).
[0065] The RF signals are then coupled to a B-mode (i.e., intensity mode, or 2D imaging mode) processor 26 and a Doppler processor 28. The B-mode processor 26 performs amplitude detection on the received ultrasound signals for imaging of internal structures, such as organ tissue and blood vessels. In line-by-line imaging, each line (beam) is represented by an associated RF signal, and the amplitude of the associated RF signal is used to generate an intensity value that is assigned to a pixel in the B-mode image. The exact location of a pixel in the image is determined by the location of the associated amplitude measurement along the RF signal and the number of lines (beams) of the RF signal. B-mode images of such structures are formed in harmonic or fundamental imaging modes, or in a combination of both, as described in U.S. Pat. No. 6,283,919 (Roundhill et al.) and U.S. Pat. No. 6,458,083 (Jago et al.). The Doppler processor 28 processes temporally distinct signals resulting from tissue motion and blood flow for detection of moving material, such as the flow of blood cells, in the image field. The Doppler processor 28 typically includes a wall filter with parameters set to pass or reject echoes returned from selected types of material within the body.
[0066] The structural and motion signals produced by the B-mode and Doppler processors are coupled to the scan converter 32 and multiplanar reformatter 44. The scan converter 32 arranges the echo signals in the spatial relationship in which they were received in the desired image format. In other words, the scan converter acts to convert the RF data from a cylindrical coordinate system to a Cartesian coordinate system suitable for displaying ultrasound images on the image display 40. For B-mode imaging, the brightness of a pixel at a given coordinate is proportional to the amplitude of the RF signal received from that location. For example, the scan converter arranges the echo signals in a two-dimensional (2D) sector format or a pyramidal three-dimensional (3D) image. The scan converter can overlay the B-mode structural image with colors corresponding to the motion at points within the image field, where Doppler estimated velocity produces a given color. The combined B-mode structural image and color Doppler image depict tissue motion and blood flow within the structural image field. The multiplanar reformatter will convert echoes received from points in a common plane in a volumetric region of the body into an ultrasound image of that plane, as described in U.S. Patent No. 6,443,896 (Detmer). The volume renderer 42 converts the echo signals of the 3D data set into a projected 3D image as seen from a given reference point, as described in U.S. Patent No. 6,530,885 (Entrekin et al.).
[0067] The 2D or 3D images are coupled from the scan converter 32, multiplanar reformatter 44, and volume renderer 42 to the image processor 30 for further enhancement, buffering, and temporary storage for display on the image display 40. The imaging processor is adapted to remove certain imaging artifacts from the final ultrasound image: e.g., caused by strong attenuators or refraction, acoustic shadowing e.g., caused by weak attenuators, back enhancement e.g., where highly reflective tissue interfaces are located in close proximity, reflection artifacts, etc. In addition, the image processor is adapted to handle certain speckle reduction functions to improve the contrast of the final ultrasound image.
[0068] In addition to being used for imaging, the blood flow values produced by Doppler processor 28 and the tissue structure information produced by B-mode processor 26 are coupled to a quantification processor 34. The quantification processor produces measurements of different flow conditions, such as blood volume flow rate, as well as structural measurements, such as organ size and gestational age. The quantification processor receives input from a user control panel 38, for example, points within the image anatomy where measurements should be taken.
[0069] Output data from the quantification processor is coupled to a graphics processor 36 for reproduction of measurement graphics and values in an image on a display 40 and for audio output from the display device 40. The graphics processor 36 can also generate image overlays for display with the ultrasound images. These image overlays can include standard identifying information such as the patient name, the date and time of the image, and imaging parameters. For these purposes, the graphics processor receives input, such as the patient name, from a user interface 38. The user interface is also coupled to a transmit controller 18 to control the generation of ultrasound signals from the transducer array 6 and, therefore, the images produced by the transducer array and ultrasound system. The transmit control function of the controller 18 is only one of the functions performed. The controller 18 also takes into account the mode of operation (given by the user) and the corresponding required transmitter configuration and bandpass configuration in the receiver analog-to-digital converter. The controller 18 can be a state machine with fixed states.
[0070] The user interface is also coupled to a multiplanar reformatter 44 for plane selection and control of multiple multiplanar reformatted (MPR) images used to perform quantified measurements in the image field of the MPR images.
[0071] The present disclosure relates to a process for analyzing multiple ultrasound images, such as those generated by the aforementioned ultrasound images, which process is performed by a processor of the ultrasound images, such as graphics processor 36, image processor 30, and / or a separate / dedicated processing system (not shown).
[0072] 2 shows a schematic workflow 200 for understanding the approach taken by an embodiment of the present invention. The workflow represents a process performed by a processing system.
[0073] The workflow involves processing multiple ultrasound images 210, all of the same resolution. The multiple ultrasound images are ultrasound images of an individual, particularly ultrasound images of the same anatomical area of the individual (and are preferably acquired from the same or similar viewpoints, e.g., with an imaging probe that does not move significantly). Preferably, each ultrasound image is a two-dimensional image, although this is not required (e.g., 3D ultrasound images could be used).
[0074] An ultrasound image is any image acquired using an ultrasound imaging process, for example, using an ultrasound imaging system such as those described above. In some examples, the "initial" ultrasound image undergoes additional processing (e.g., using one or more filters and / or one or more layers of a neural network) before being used as the ultrasound image for purposes of the method / workflow. Thus, the ultrasound image is a feature space image or an image space ultrasound image.
[0075] In a particularly preferred example, the plurality of ultrasound images comprises a sequence of ultrasound images, e.g., a sequence of ultrasound images captured by the same ultrasound imaging probe at different times. Preferably, the order of the ultrasound images in the sequence depends on the time at which each ultrasound image was captured. For example, the sequence comprises sequential frames of an ultrasound video.
[0076] Each of the plurality of ultrasound images is processed in process 220 to identify a region of interest within each ultrasound image. A region of interest is a portion or segment of an ultrasound image (i.e., not the entire ultrasound image) that represents an area of an individual that has / contains a potential pathology.
[0077] The lesion may include one or more of a tumor, a growth, an abscess, a nodule, a swelling, a lump, an ulcer, or any other suitable feature resulting from an injury, disease, or disease in an individual. Mechanisms for identifying regions of interest in ultrasound images are well known to those skilled in the art, and may use, for example, machine learning methods, edge detection algorithms, image segmentation techniques, etc.
[0078] The regions of interest 230 are then subjected to a grouping or clustering process 240, in which similar regions of interest are grouped together. In other words, the regions of interest are grouped (e.g., into one or more groups of regions of interest) based on their similarity. The similarity of the regions of interest is defined based on the similarity of the content of the regions of interest. In particular, regions of interest are grouped together if they are predicted to identify the same potential disease in an individual.
[0079] Each group of areas of interest 250, of which only one example is shown, is then processed in a process 260 using machine learning methods to generate a predictive indicator 270 that indicates the likelihood that the group of areas of interest contains a disease in an individual.
[0080] The prediction indicator 270 includes binary, categorical, or numeric data that represents the likelihood that the group contains a disease. As one example, the prediction indicator is a binary indicator (e.g., "0" or "1") that indicates a prediction of whether the group of regions of interest contains a disease. As another example, the prediction indicator is a probability (i.e., a numeric indicator) that the group of regions of interest contains a disease. The numeric indicator is a scale of 0 to 1, 0 to 10, 1 to 10, 0 to 100, or 1 to 100 (although other examples may be used). As yet another example, the prediction indicator is a categorical indicator that indicates a categorical indication of the likelihood that the group of regions of interest contains a disease (e.g., "likely," "unlikely," "very likely," "neither likely nor unlikely," etc.).
[0081] The proposed technique thereby performs a two-step process for predicting the likelihood that a disease is present in an individual using multiple ultrasound images. Each ultrasound image is processed individually to identify regions of interest, which are areas / portions of the ultrasound image that potentially depict disease (i.e., regions that are candidates for containing disease). Similar regions of interest are then grouped and then processed using machine learning methods to generate a predictive indicator responsive to the likelihood that the group indicates disease. Processing using the groups effectively confirms or rejects the suggestion that each region of interest in the group indicates disease (this suggestion being made by processing each individual ultrasound image).
[0082] The proposed approach avoids the need to perform complex, high-intensity processing of all ultrasound images simultaneously by processing each individual ultrasound image separately before subsequently processing a group of ultrasound image portions. The inventors have recognized that the proposed approach can significantly reduce the false positive rate while maintaining high precision / accuracy in disease likelihood prediction.
[0083] Now that an overview of the workflow has been described, a more complete example of how the exemplary workflow can be performed is provided below.
[0084] 3 illustrates a method 300 according to one embodiment. The method 300 is performed by a processing system according to one embodiment.
[0085] The method 300 includes acquiring 310 a plurality of ultrasound images of an individual. The plurality of ultrasound images are as described above with reference to Figure 2. For ease of explanation, the illustrated ultrasound images are image space ultrasound images (i.e., each pixel directly represents a portion of the individual's anatomy).
[0086] In particular, the ultrasound images are preferably two-dimensional ultrasound images, although 3D ultrasound images are also possible. Preferably, the plurality of ultrasound images comprises a sequence of ultrasound images, e.g., where one ultrasound image is captured sequentially after another ultrasound image or later at a later time than the other ultrasound images. In some examples, the plurality of ultrasound images comprises (sequential) frames from an ultrasound video.
[0087] The ultrasound image is then processed in step 320, which includes identifying one or more regions of interest in the ultrasound image, each region of interest being a portion of the ultrasound image that represents an area of potential disease in the individual. A region of interest is therefore a portion of a larger ultrasound image, which is itself an ultrasound image.
[0088] The process for identifying regions of interest in an ultrasound image will be readily apparent to one skilled in the art and involves using machine learning methods (e.g., neural networks or naive Bayes classifiers) to identify any portion of the ultrasound image that contains or represents a potential pathology in an individual. Typically, a region of interest is defined by identifying (the coordinates of) a rectangle or volume that defines the outer boundary of the area or volume that contains the potential pathology in the individual. Each region of interest is thus a specific portion of the ultrasound image.
[0089] Preferably, each region of interest is no more than 0.4 times the size of the ultrasound image, e.g., no more than 0.25 times the size of the ultrasound image. The smaller the regions of interest, the greater the reduction in computational complexity (compared to the entire ultrasound image) of processing a group of regions of interest.
[0090] The regions of interest are processed to remove or delete any overlapping regions of interest (e.g., regions of interest within the same ultrasound image that identify the same potential disease). This is performed, for example, by processing each ultrasound image to identify any regions of interest within that ultrasound image that overlap with each other by more than a predetermined amount (e.g., have an IoU value greater than a predetermined value) and deleting one or more of the overlapping regions of interest. The deleted region of interest is the one associated with the lowest confidence / probability that it identifies the potential disease. Such confidence / probability values may be generated if the regions of interest are identified using machine learning techniques, although other techniques also generate such confidence / probability measures.
[0091] The method 300 then performs step 330 of grouping or clustering the identified regions of interest to form groups of regions of interest. This grouping is performed based on the similarity of the identified regions of interest to each other. Each group is adapted / configured to include regions of interest that depict / have the same potential disease.
[0092] Step 330 is performed by linking or associating regions of interest from different ultrasound images. Linked regions of interest are regions of interest believed to contain the same potential pathology. Step 330 then creates or forms groups of regions of interest by grouping or clustering the linked regions of interest into a single group, for example, if they meet certain requirements.
[0093] Preferably, each group contains only one region of interest from any given ultrasound image (ie, each ultrasound image can contribute up to one region of interest to a group of ultrasound images).
[0094] In a particularly preferred example where the plurality of ultrasound images comprises a sequence of ultrasound images, each group of regions of interest includes only regions of interest from sequentially adjacent ultrasound images, in other words, each region of interest (within a particular group of regions of interest) comes from an ultrasound image (within the sequence of ultrasound images) that is sequentially adjacent to an ultrasound image that includes another region of interest (within a particular group of regions of interest).
[0095] This approach means that groups of regions of interest can only be constructed by comparing regions of interest from ultrasound images that are sequentially adjacent to one another (i.e., immediately before or after them) to effectively construct a sequence of regions of interest, which provides additional information useful for subsequent analysis of the groups of regions of interest, yielding groups of regions of interest that are sequentially related to one another (e.g., temporally related).
[0096] This approach has the advantage of providing groups of regions of interest that are sequential to one another, providing useful contextual information for subsequent analysis (e.g., because sequence information provides an important indicator as to whether disease is present or not), and reducing the number of comparisons that need to be made (because only regions of interest in sequentially adjacent ultrasound images need to be compared to one another).
[0097] In other words, step 330 is performed by linking or associating regions of interest in sequentially adjacent ultrasound images before creating groups of regions of interest by grouping the linked regions of interest into a single group.
[0098] In a first scenario, step 330 involves determining a similarity measure between different regions of interest based on the relative overlap between the areas occupied by each region of interest in their respective ultrasound images, and if this similarity measure exceeds some predetermined threshold, the regions of interest are grouped (or linked / associated with each other).
[0099] For example, the size of the overlap between the area occupied by one region of interest in an ultrasound image and the area occupied by another region in another ultrasound image is used as a similarity measure. The size of the overlap is determined using an IoU evaluation metric. This procedure can establish links between regions of interest across different ultrasound images.
[0100] By way of example only, if the region of interest of a first ultrasound image is represented by coordinates (X1, Y1, X2, Y X Consider a scenario in which the region of interest in one ultrasound image is a rectangle defined by coordinates (X3, Y3, X4, Y4) and the region of interest in the second ultrasound image is a rectangle defined by coordinates (X3, Y3, X4, Y4). Because the coordinate information defines the coverage or extent of each region of interest, this coordinate information can be used to identify the size of the relative overlap between the two regions of interest, which can be used to calculate an IoU measure.
[0101] In some examples, if the IoU measure is greater than some predetermined value, preferably 0.4 or greater, for example 0.5 or greater, for example 0.6 or greater, the regions of interest are considered to be sufficiently similar (i.e., suitable for placement in the same group).
[0102] In a second scenario, step 330 involves grouping regions of interest based on metadata associated with each region of interest.
[0103] This metadata may, for example, describe characteristics of the region of interest. Example metadata for a region of interest may include the location of the region of interest (e.g., the location of the center of the region of interest), the geometry (e.g., the shape) and / or size of the region of interest, a confidence value for the region of interest (e.g., a value representing the confidence or probability that the region of interest contains a potential disease—when provided by a machine learning method), and / or a feature map for the region of interest.
[0104] A feature map of a region of interest is a map generated by applying one or more filters, such as convolutional filters, to the region of interest. Specifically, a feature map is a map generated by applying one or more layers of a neural network to the region of interest. The meaning of the term feature map is well known in the field of machine learning.
[0105] In some examples, a classifier, e.g., a machine learning method, processes each region of interest and classifies the region of interest into one of a plurality of classifications, and a group of regions of interest, e.g., includes only regions of interest that have the same classification.
[0106] Preferably, a combination of the above techniques is used, specifically, for multiple images formed in a sequence of multiple images, each group of regions of interest includes only regions of interest from ultrasound images that are sequentially adjacent to one or two other ultrasound images that provide one or more other regions of interest in the group of regions of interest and that overlap each other by more than a predetermined amount.
[0107] The method 300 includes a step 340 of using machine learning methods to process the group of regions of interest to generate a predictive indicator of the likelihood that the group of regions of interest contains a disease condition in an individual.
[0108] In particularly preferred examples, the output of the machine learning method is data, e.g., in the form of binary, categorical and / or numerical indicators, that vary in response to the predicted likelihood that a group of regions contains disease.
[0109] This process uses multiple, similar regions of interest to predict whether they contain or depict a disease in an individual. This provides additional, contextually relevant information for generating a predictive indicator compared to, for example, simply processing only a single ultrasound image. This approach thereby provides a more accurate mechanism for determining the likelihood of disease presence (in each group of regions of interest) because additional contextual information is provided.
[0110] By using groups of regions of interest, the proposed approach also does not require processing groups of full-sized ultrasound images, which can be computationally expensive and / or require significant additional storage space.
[0111] It is further noted that the proposed approach also increases the accuracy of identifying disease in an individual, as it uses a two-step process to identify or predict the likely presence of disease.
[0112] Step 340 includes, for each group of regions of interest, processing the regions of interest to generate a sequence of regions of interest. The sequence is time-series, such that the regions of interest are ordered sequentially based on the time the first ultrasound image (for each region of interest) was first captured. The sequence of regions is then processed using machine learning methods to generate a predictive indicator.
[0113] The order of the sequence of regions of interest, if acquired from a sequence of ultrasound images, corresponds to the order of the sequence of ultrasound images from which the regions of interest are acquired.
[0114] This sequence of regions of interest is called a "tube" or "tublet," and effectively represents a sequence (portion of an image) depicting the same potential lesion. For example, a "tube" represents a video of the potential lesion (rather than a video of the entire area imaged by the ultrasound imaging probe) that is sized / shaped to fit around the potential lesion.
[0115] In some examples, the sequence of regions of interest includes a stack of regions of interest within a group of regions of interest, e.g., if each region of interest is a two-dimensional image, the sequence of regions of interest includes a 3D volume representing a simple stack of the two-dimensional images on top of each other.
[0116] In other words, a group of regions of interest is effectively treated as a "volume" that can be classified or processed using machine learning methods, i.e., a sequence of regions of interest is a single data structure containing the stacked regions of interest, i.e., all regions of interest combined.
[0117] The input to the machine learning method is the group of regions themselves, the sequence of regions of interest, and / or other data derived from the group of regions of interest, for example, one or more feature maps derived by processing the group of regions of interest (e.g., using one or more filters or layers of a neural network).
[0118] A machine learning algorithm is any self-training algorithm that processes input data to produce or predict output data. For purposes of step 340, the input data includes data derived from a group of regions of interest, and the output data includes a predictive indicator of the likelihood that the group of regions of interest contains a disease in an individual.
[0119] Suitable machine learning algorithms for use in the present invention will be apparent to those skilled in the art. Examples of suitable machine learning algorithms include decision tree algorithms and artificial neural networks (e.g., CNN, RNNS, and / or LSTM). Other machine learning algorithms, such as logistic regression, support vector machines, or naive Bayesian models, are suitable alternatives.
[0120] The structure of an artificial neural network (or simply, a neural network) is inspired by the human brain. A neural network is made up of layers, each layer containing multiple neurons. Each neuron contains a mathematical operation. Specifically, each neuron contains a different weighted combination of a single type of transformation (e.g., the same type of transformation, such as sigmoid, but with different weightings). In the process of processing input data, each neuron's mathematical operation is performed on the input data to produce a numerical output, and the output of each layer in the neural network is sequentially fed to the next layer. The final layer provides the output.
[0121] Methods for training machine learning algorithms are well known. Typically, such methods involve obtaining a training dataset including training input data entries and corresponding training output data entries (commonly labeled "ground truth" data). An initialized machine learning algorithm is applied to each input data entry to generate a predicted output data entry. The error between the predicted output data entry and the corresponding training output data entry is used to correct the machine learning algorithm. This process can be repeated until error convergence occurs, and the predicted output data entries are sufficiently similar to the training output data entries (e.g., within ±1%). This is commonly known as a supervised learning technique.
[0122] For example, when a machine learning algorithm is formed from a neural network, the mathematical operations (weights) of each neuron are modified until the error converges. Known methods for modifying neural networks include gradient descent, backpropagation algorithms, etc.
[0123] The training input data entries correspond to example data derived from groups of regions of interest (e.g., the groups themselves), and the training output data entries correspond to example predictions (in the form of binary, categorical, or numerical data) regarding the likelihood of disease presence in the groups.
[0124] The method 300 further includes displaying 350 a visually perceptible output on a display device in response to the predictive indicator.
[0125] Thus, step 350 includes controlling a display device to provide a visually perceptible output responsive to the predictive indicator (for each group of regions), i.e., to indicate the predicted likelihood that the individual will have the disease. The indicator may, for example, provide a predicted probability that the individual will have the disease and / or a binary indicator (e.g., indicating "disease present" or "disease free").
[0126] As another example, if the predictive indicators are numerical values indicating a predicted probability and one of the predictive indicators has a value that exceeds a predetermined value, then step 350 includes controlling a display device to indicate that disease is present.
[0127] In some examples, the visually perceptible output indicates which (if any) of the groups of regions of interest are predicted to contain or depict disease, for example, in the form of a probability and / or binary indicator for each group of regions or for only those regions that are considered likely to contain disease (e.g., have a likelihood above some predetermined threshold).
[0128] In some examples, in response to determining that at least one of the group of regions of interest contains a lesion, for example, if the predicted likelihood exceeds some threshold, step 350 includes identifying a location of the identified lesion on a displayed ultrasound image of the patient and overlaying a visually perceptible output (e.g., an annotation, a rectangle, etc.) on the displayed ultrasound image in response to the identified location of the lesion. The location of the identified lesion is the location of the region of interest (of the group of regions of interest predicted to contain the lesion) in the displayed ultrasound image. Specifically, the visually perceptible output overlays the location of the lesion, e.g., at the location of the region of interest predicted to contain the lesion, on the ultrasound image. The ultrasound image is one of a plurality of ultrasound images that includes a region of interest that formed part of the group of regions of interest in which the presence of the lesion was identified.
[0129] In some examples, all regions of interest are identified (e.g., using respective markers), and regions of interest within the group of regions of interest that are predicted to contain disease (e.g., have a probability greater than some predetermined value) are highlighted or otherwise enhanced, e.g., with a particular color, pattern, etc. This facilitates increased ease of identification of potential areas where disease may be found (e.g., for diagnostic purposes), while also providing information regarding the automated prediction of the likelihood that the area contains disease.
[0130] In some examples, information regarding the predicted likelihood that a region contains a disease of interest is displayed, for example, using shading or transparency adjustment techniques.
[0131] In some examples, the display device is configured to sequentially display each ultrasound image of the plurality of ultrasound images (e.g., play an ultrasound video). Step 350 includes, in response to determining (e.g., by processing the predictive indicator) that at least one of the group of regions of interest is likely to contain a lesion, controlling the display device to provide a visually perceptible output of the position of the region of interest (from the group of regions of interest) associated with the currently displayed ultrasound image with respect to the currently displayed ultrasound image. In this manner, the relative position of the lesion associated with the group of regions of interest can be tracked or displayed sequentially.
[0132] 2 and 3 are used to illustrate a relatively simple embodiment in which each of the plurality of ultrasound images comprises an image space ultrasound image.
[0133] However, as mentioned above, in some examples, each of the plurality of ultrasound images comprises a feature space ultrasound image, i.e., a feature space image, which is a "traditional" ultrasound image ("image space image") that has undergone further processing, such as using one or more filters and / or layers of a neural network to generate the feature space image.
[0134] In some instances, these feature space images are processed in the same manner as the ultrasound images described above, however, in other instances, some additional processing is performed to identify (groups of) regions of interest.
[0135] FIG. 4 illustrates a process in which multiple ultrasound images include multiple feature space images 410, each derived from a respective image space (ultrasound) image 420 (eg, as a result of a feature extraction process 430).
[0136] Thus, the step of acquiring a plurality of ultrasound images includes acquiring a plurality of feature space ultrasound images, which is performed by processing the plurality of image space ultrasound images in a feature extraction process 430 to generate a feature space ultrasound image.
[0137] The process of identifying a region of interest in the ultrasound image includes independent processing of the feature space ultrasound image to identify a region of interest in the feature space ultrasound image.
[0138] In an alternative embodiment, this process includes a process 435 of identifying regions of interest in the image space ultrasound image 420. This is performed using any of the mechanisms described above, for example, using machine learning methods. Regions of interest in the feature space ultrasound image are then determined based on the regions of interest in the image space ultrasound image, as shown schematically using dotted lines. In particular, spatial relationships between different areas of the feature space ultrasound image and the image space ultrasound image are known and can therefore be used to identify relevant regions of interest in the feature space ultrasound image.
[0139] Thus, the step of identifying a region of interest in an ultrasound image includes identifying the region of interest by first identifying the region of interest in an initial ultrasound image (used to derive the multiple ultrasound images) before identifying the region of interest in the multiple ultrasound images based on the identified region of interest in the initial ultrasound image.
[0140] The regions of interest within the multiple (feature space) ultrasound images are then grouped in process 440 to generate one or more groups 460 of regions of interest.
[0141] Each group of regions of interest 450 is processed using machine learning methods to generate a predictive indicator 470 of the likelihood that the group of regions of interest contains disease. This can be done using any of the techniques previously described.
[0142] FIG. 5 illustrates an exemplary process for grouping regions of interest.
[0143] Specifically, FIG. 5 illustrates a process for grouping regions of interest from consecutive ultrasound images within a sequence of ultrasound images (eg, sequential frames of an ultrasound video).
[0144] 5 shows three ultrasound images, all two-dimensional and with the same resolution. Specifically, there is a first ultrasound image 510, a second ultrasound image 520, and a third ultrasound image 530. The ultrasound images form a (temporal) sequence of ultrasound images, e.g., representing frames of an ultrasound video.
[0145] Each ultrasound image has been processed, for example using machine learning methods, to identify regions of interest, which are shown schematically in Figure 5 using rectangles.
[0146] The regions of interest are then grouped by linking or associating regions of interest in adjacent frames if a measure of their similarity (in adjacent frames) exceeds some predetermined threshold. The measure of similarity is the IoU value, which is calculated by dividing the size of the overlap between the regions (i.e., the size of the overlap if the regions were located in the same ultrasound image) by the size of the union between the regions.
[0147] A sequence of regions of interest is then acquired by selecting linked regions of interest in different ultrasound images. Specifically, the acquired sequence of regions of interest is a sequence that includes up to one region of interest from each ultrasound image, where each region of interest in the sequence is linked or associated with another region of interest in the sequence. The order of the regions of interest in the sequence corresponds to the order of the ultrasound images in the sequence in which they were acquired.
[0148] The regions that were placed in the sequence of regions of interest are then deleted (preserving the sequence of regions of interest) and the process is repeated, as shown diagrammatically using the arrows in Figure 5.
[0149] 6 illustrates a processing system 600 according to one embodiment. The processing system is described in the context of an overall ultrasound imaging system 60.
[0150] The processing system 600 is configured to acquire multiple ultrasound images of the individual, for example, from a memory 610 and / or an ultrasound scanner 620 configured to generate ultrasound images.
[0151] The processing system 600 is also configured to identify regions of interest in the ultrasound image, each region of interest being a portion of the ultrasound image that represents an area of potential disease in the individual.
[0152] The processing system 600 is also configured to group regions of interest from different ultrasound images based on similarities between the regions of interest and process each group of regions of interest using machine learning methods to predict whether the group of regions of interest contains a disease in an individual.
[0153] The processing system 600 is also configured to display a visually perceptible output on a display device in response to each predictive indicator.
[0154] As previously mentioned, processing system 600 is suitably adapted, mutatis mutandis, to perform any of the methods described herein, and one of ordinary skill in the art will readily be able to suitably adapt processing system 600.
[0155] FIG. 7 shows the effectiveness of the proposed approach for predicting whether an individual has one or more diseases.
[0156] A first graph 710 shows a recall-precision curve (recall on the x-axis and precision on the y-axis) for accurately identifying lesions in two-dimensional ultrasound images using known two-dimensional neural network analysis techniques. This graph was generated using benchmark data. Specifically, the first graph 710 shows the recall-precision curve resulting from the use of a lesion detection algorithm using Faster RCNN.
[0157] The second graph 720 shows a recall precision curve for accurately identifying lesions in multiple ultrasound images by implementing the techniques disclosed herein. The second graph 720 was generated using the same benchmark data used to generate the first graph 710.
[0158] The first and second graphs 710, 720 are to the same scale (eg, 0.0 to 1.0 on the x-axis and 0.0 to 1.0 on the y-axis).
[0159] In the proposed approach, higher precision values can be obtained for higher levels of recall, and vice versa. In other words, our inventive improvement shows that the use of additional context provided in multiple ultrasound images in the disclosed manner (e.g., using ultrasound video), where context is not present in simple still images, can reduce false positive and / or missed detections of disease.
[0160] 8 shows a schematic diagram of a processing system 600 according to an embodiment of the present disclosure. As shown, the processing system 600 includes a (data) processor 860, a memory 864, and a communication module 868. These elements are in direct or indirect communication with each other, for example, via one or more buses.
[0161] Processor 860 includes a central processing unit (CPU), a digital signal processor (DSP), an ASIC, a controller, an FPGA, another hardware device, a firmware device, or any combination thereof configured to perform the operations described herein. Processor 860 may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration. In some embodiments, the processor is a distributed processing system, e.g., in the form of a set of distributed processors.
[0162] The memory 864 may include cache memory (e.g., cache memory of the processor 860), random access memory (RAM), magnetoresistive RAM (MRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, solid-state memory devices, hard disk drives, other forms of volatile and non-volatile memory, or a combination of different types of memory. In one embodiment, the memory 864 includes a non-transitory computer-readable medium. The non-transitory computer-readable medium stores instructions. For example, the memory 864 or the non-transitory computer-readable medium may have program code recorded thereon, the program code including instructions for causing the processing system 600, or one or more components of the processing system 600, particularly the processor 860, to perform the operations described herein. For example, the processing system 600 may perform the operations of the method 700. The instructions 866 may also be referred to as code or program code. The terms "instructions" and "code" should be interpreted broadly to include any type of computer-readable statement. For example, the terms "instructions" and "code" refer to one or more programs, routines, subroutines, functions, procedures, etc. "Instructions" and "code" include a single computer-readable statement or multiple computer-readable statements. Memory 864 with code recorded thereon is referred to as a computer program product.
[0163] Communications module 868 may include any electronic and / or logical circuitry for facilitating direct or indirect communication of data between processing system 600, a penetration device, and / or a user interface (or other additional device). In that regard, communications module 868 may be an input / output (I / O) device. In some cases, communications module 868 facilitates direct or indirect communication between processing circuit 600 and / or various elements of the system (FIG. 6).
[0164] It will be appreciated that the disclosed methods are preferably computer-implemented methods, and as such, the concept of a computer program comprising computer program code for performing any of the described methods when said computer program is run on a processing system, e.g., a computer or a set of distributed processors, is also proposed.
[0165] Different portions, lines, or blocks of code of a computer program according to one embodiment are executed by a processing system or computer to perform any of the methods described herein. In some alternative implementations, functions depicted in block diagrams or flow diagrams may occur out of the order depicted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending on the functionality involved.
[0166] The present disclosure proposes a computer program (product) comprising instructions that, when the program is executed by a computer or processing system, cause the computer or processing system to perform (the steps of) any of the methods described herein. The computer program (product) may be stored on a non-transitory computer-readable medium.
[0167] Likewise, a computer-readable (storage) medium is also proposed, which comprises instructions that, when executed by a computer or processing system, cause the computer or processing system to perform (the steps of) any of the methods described herein. Also proposed is a computer-readable data carrier having stored thereon such a computer program (product). Also proposed is a data carrier signal embodying such a computer program (product).
[0168] Variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure, and the appended claims. In the claims, the word "comprises" does not exclude other elements or steps, and singular elements do not exclude the plural. The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.
[0169] The disclosed methods are preferably computer-implemented methods and are performed by a suitable processing system, and any processing system described herein is suitably adapted to perform any method described herein, just as any method described herein is adapted to perform a process performed by any processing system described herein.
[0170] Also proposed is a computer program product comprising computer program code means which, when executed on a computing device having a processing system, causes the processing system to perform all of the steps of any of the methods described herein.
[0171] A single processor or other unit may fulfill the functions of several items recited in the claims. The computer program may be stored / distributed on a suitable medium, for example an optical storage medium or a solid-state medium provided together with or as part of other hardware pieces, but may also be distributed in other ways, for example via the Internet or other wired or wireless telecommunications system.
[0172] The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be advantageously used. When the term "adapted" is used in the claims or the description, it is intended to be equivalent to the term "configured." Any reference signs in the claims should not be construed as limiting the scope.
Claims
1. 1. A processing system for predicting the likelihood that an individual has one or more diseases, comprising: acquiring a plurality of ultrasound images of the individual, the ultrasound images comprising a time series of ultrasound images; identifying regions of interest within the plurality of ultrasound images, each region of interest being a portion of one of the plurality of ultrasound images that represents an area of the individual having a potential disease; grouping regions of interest from different ones of the ultrasound images based on similarities of the regions of interest; processing each group of regions of interest using machine learning methods to generate a predictive indicator of the likelihood that said group of regions of interest contains a disease condition for said individual; A processing system that performs the above.
2. The processing system of claim 1 , wherein each ultrasound image is a two-dimensional ultrasound image.
3. The processing system, for each group of regions of interest, generating a sequence of regions of interest using the regions of interest of the group of regions of interest; processing said sequence of regions of interest using machine learning methods to generate a predictive indicator of the likelihood that said sequence of regions of interest contains a disease in said individual; 3. The processing system of claim 1, wherein each group of regions of interest is processed by performing a process comprising:
4. The processing system of claim 3 , wherein the processing system generates the sequence of regions of interest by performing a step including stacking the regions of interest.
5. 5. The processing system of claim 1, wherein the processing system identifies regions of interest in the ultrasound images by performing a process that includes processing each ultrasound image using a second machine learning method to identify regions of interest.
6. The processing system of claim 1 , wherein the plurality of ultrasound images comprises a video of ultrasound images.
7. The processing system of claim 6 , wherein the order of the ultrasound images in the sequence depends on the time each ultrasound image was captured.
8. 8. The processing system of claim 6 or 7, wherein in each group of regions of interest, each region of interest is derived from an ultrasound image that is sequentially adjacent to an ultrasound image of another region of interest in the same group of regions of interest.
9. The processing system of claim 1 , wherein each region of interest is no larger than 0.25 times the size of the ultrasound image.
10. The processing system of claim 1 , wherein the processing system displays a visually perceptible output on a display device in response to each predictive indicator.
11. If the processing system determines that at least one predictive indicator indicates that at least one group of regions of interest is likely to contain disease, Identifying the location of the identified lesion with respect to a displayed ultrasound image of the patient; and superimposing a visually perceptible output on the displayed ultrasound image in response to the identified location of the lesion. The processing system of claim 10.
12. the processing system comprising: determining a similarity measure between different regions of interest from different ultrasound images, the similarity measure representing the size of the overlap between the relative area occupied by one region of interest in an ultrasound image and the relative area occupied by another region of interest in another ultrasound image; grouping two different regions of interest into the same group in response to said similarity measure between the two different regions of interest exceeding a predetermined threshold; 12. The processing system of claim 1, wherein the processing system groups regions of interest by performing a process comprising:
13. an ultrasound scanner for generating a plurality of ultrasound images of an individual, including a time series of ultrasound images; 12. A processing system according to any one of claims 1 to 11 for predicting the likelihood that the individual has one or more diseases based on the plurality of ultrasound images; 1. An ultrasound imaging system comprising:
14. 1. A computer-implemented method for predicting the likelihood that an individual has one or more diseases, the computer-implemented method comprising: acquiring a plurality of ultrasound images of the individual, the ultrasound images comprising a time series of ultrasound images; identifying regions of interest within the plurality of ultrasound images, each region of interest being a portion of one of the plurality of ultrasound images that represents an area of the individual having a potential disease; grouping regions of interest from different ones of the ultrasound images based on similarities of the regions of interest; processing each group of regions of interest using machine learning methods to generate a predictive indicator of the likelihood that said group of regions of interest contains a disease condition for said individual; 20. A computer-implemented method comprising:
15. 15. A computer program comprising computer program code means which, when executed on a computing device having a processing system, causes said processing system to perform all of the steps of the computer-implemented method according to claim 14.
Citation Information
Patent Citations
Apparatus and method for lesion detection
JP2015154918A
Computer-aided detection using multiple images from different views of a region of interest to improve detection accuracy
JP2019530490A
Computer aided diagnosis (CAD) apparatus and method
US20160148376A1
Methods and systems for motion detection and compensation in medical images
US20200121294A1
Biopsy prediction and guidance with ultrasound imaging and associated devices, systems, and methods
WO2020002620A1