System and method for synthesizing fetal heartbeat audio signature
A processing system generates an audio signal from non-Doppler ultrasound images to address the challenge of blind sweep procedures, improving user confidence and emotional connection by providing audible fetal heartbeat feedback.
Patent Information
- Application Number
- PCT/EP2025/079840
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-10-25
- Filing Date
- 2025-10-16
- Publication Date
- 2026-04-30
AI Technical Summary
Blind sweep ultrasound procedures for fetal monitoring are challenging for non-expert operators due to the inability to visually confirm the procedure's correctness, lacking direct feedback on fetal heartbeat, and there is a desire to improve user confidence and emotional connection for pregnant individuals.
A processing system that generates an audio signal representative of fetal heartbeat characteristics from non-Doppler mode ultrasound images, using machine learning algorithms to estimate and modulate heartbeat features, providing audible feedback during procedures.
Enhances operator confidence and emotional connection by allowing non-experts to verify correct procedure execution and provides reassurance through audible fetal heartbeat, reducing the risk of alarming pregnant individuals with abnormal features.
Smart Images

Figure EP2025079840_30042026_PF_FP_ABST
Abstract
Description
[0001] SYSTEM AND METHOD FOR SYNTHESIZING FETAL HEARTBEAT AUDIO SIGNATURE
[0002] FIELD OF THE INVENTION
[0003] The invention relates to the field of ultrasound imaging, and in particular, to fetal ultrasound imaging.
[0004] BACKGROUND OF THE INVENTION
[0005] Blind sweep ultrasound procedures for fetal monitoring consist of a number of sweeps over the abdomen of a pregnant individual, guided only by external anatomic landmarks. Blind sweep workflows are commonly used for fetal ultrasound examinations performed by non-experts, as blind sweeps can be carried out with only a few hours of training. This enables increased fetal monitoring in places where fully-trained sonographers are scarce, thus improving safety during pregnancy in these places.
[0006] There is an ongoing desire to improve blind sweep ultrasound procedures.
[0007] SUMMARY OF THE INVENTION
[0008] The invention is defined by the claims.
[0009] According to examples in accordance with an aspect of the invention, there is provided a processing system for synthesizing an audible fetal heartbeat sound, the processing system being configured to: receive a sequence of non-Doppler mode ultrasound images of a fetus, wherein a heart of the fetus is at least partially within a field of view of at least a subset of the sequence; and process the sequence of non-Doppler mode ultrasound images to generate an audio signal representative of one or more characteristics of a heartbeat of the fetus.
[0010] One difficulty with performing blind sweep ultrasound procedures is that nonexpert operators are unsure as to whether they are performing the procedure correctly, as they are unable to view live ultrasound images. The ability to directly listen to the heartbeat of the fetus, which is not currently possible with B-mode fetal heart rate determination techniques used in blind sweep workflows, would provide inexperienced operators with an indication that they are performing the procedure correctly, thus improving user confidence. Similarly, the audio signal may improve operator confidence for any ultrasound imaging procedure that does not currently provide audio of the fetal heartbeat, particularly when the imaging procedure is performed by a less experienced operator. Pregnant individuals also often wish to hear the fetal heartbeat, both for reassurance that the fetus is healthy and to help build an emotional connection with the fetus.
[0011] In some examples, the processing system is configured to generate the audio signal by: processing the sequence of non-Doppler mode ultrasound images to estimate the one or more characteristics of the heartbeat of the fetus; and processing the estimated one or more characteristics to generate the audio signal representative of the one or more characteristics.
[0012] In some examples, the processing system is configured to process the estimated one or more characteristics to generate the audio signal by: obtaining an audio recording of an example heartbeat; and modulating the audio recording of the example heartbeat according to the estimated one or more characteristics to generate the audio signal.
[0013] An audio recording of an example heartbeat may, for example, be obtained from a medical database. The audio recording may be derived from a Doppler-based fetal ultrasound procedure.
[0014] In some examples, the one or more characteristics comprise a heart rate of the fetus. In other words, the audio signal may have a frequency modulated to match the heart rate of the fetus.
[0015] In some examples, the one or more characteristics comprise an amplitude of the heartbeat of the fetus.
[0016] In some examples, the one or more characteristics comprise a heart rate variability for the fetus.
[0017] In some examples, the processing system is configured to generate the audio signal by processing the sequence of non-Doppler mode ultrasound images using a machinelearning algorithm.
[0018] In some examples, the machine-learning algorithm has been trained by receiving (or using a training algorithm configured to receive) an array of training inputs and known outputs, each training input corresponding to a respective known output, wherein each training input comprises a sequence of non-Doppler mode fetal ultrasound images, and each known output comprises an audio signal derived from a sequence of Doppler mode ultrasound images acquired in a same ultrasound imaging procedure as the corresponding training input.
[0019] Alternatively, a machine-learning algorithm may be trained to output a synthesized sequence of Doppler mode ultrasound images based on an input of a sequence of non-Doppler mode ultrasound images. The synthesized sequence of Doppler mode ultrasound images may then be used to generate the audio signal.
[0020] In some examples, the sequence of non-Doppler mode ultrasound images is a sequence of B-mode ultrasound images.
[0021] Ultrasound images acquired by non-expert operators, and particularly those acquired during blind sweep procedures, are often B-mode images.
[0022] In some examples, the sequence of non-Doppler mode ultrasound images is a sequence of M-mode ultrasound images.
[0023] M-mode ultrasound images may enable the generated audio signal to more accurately represent the heartbeat of the fetus.
[0024] In some examples, the sequence of non-Doppler mode ultrasound images was acquired during a blind sweep imaging procedure.
[0025] An audio signal representative of a fetal heartbeat may be particularly valuable in blind sweep imaging procedures, which are typically performed by non-expert operators and during which it is not possible to view live ultrasound images. An audio signal representative of the fetal heartbeat, generated based on images acquired during the blind sweep, indicates to a non-expert operator that they have performed the procedure correctly.
[0026] In some examples, the processing system is further configured to process the sequence of non-Doppler mode ultrasound images to identify, if present, any abnormal features of the heartbeat of the fetus; and the audio signal is generated only in response to failing to identify any abnormal features of the heartbeat of the fetus.
[0027] In other words, the audio signal may not be generated if one or more abnormal features are identified in the sequence of non-Doppler mode ultrasound images. This reduces a risk of alarming the pregnant individual, who may be concerned by hearing abnormal features in an audio signal of the fetal heartbeat.
[0028] In some examples, the processing system is further configured to: synchronize the sequence of non-Doppler mode ultrasound images and the generated audio signal; and control a user interface to output the synchronized sequence of non-Doppler mode ultrasound images and generated audio signal.
[0029] According to examples in accordance with another aspect of the invention, there is provided a computer-implemented method for synthesizing a fetal heartbeat, the computer-implemented method comprising: receiving a sequence of non-Doppler mode ultrasound images of a fetus, wherein a heart of the fetus is at least partially within a field of view of at least a subset of the sequence; and processing the sequence of non-Doppler mode ultrasound images to generate an audio signal representative of one or more characteristics of a heartbeat of the fetus.
[0030] There is also provided a computer program product comprising computer program code means which, when executed on a computing device having a processing system, cause the processing system to perform all of the steps of the method according to claim 14.
[0031] These and other aspects of the invention will be apparent from and elucidated with reference to the embodiment s) described hereinafter.
[0032] BRIEF DESCRIPTION OF THE DRAWINGS
[0033] For a better understanding of the invention, and to show more clearly how it may be carried into effect, reference will now be made, by way of example only, to the accompanying drawings, in which:
[0034] Figure 1 illustrates an exemplary ultrasound system;
[0035] Figure 2 illustrates an ultrasound system, according to an embodiment of the invention; and
[0036] Figure 3 illustrates a computer-implemented method for synthesizing a fetal heartbeat.
[0037] DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The invention will be described with reference to the Figures.
[0039] It should be understood that the detailed description and specific examples, while indicating exemplary embodiments of the apparatus, systems and methods, are intended for purposes of illustration only and are not intended to limit the scope of the invention. These and other features, aspects, and advantages of the apparatus, systems and methods of the present invention will become better understood from the following description, appended claims, and accompanying drawings. It should be understood that the Figures are merely schematic and are not drawn to scale. It should also be understood that the same reference numerals are used throughout the Figures to indicate the same or similar parts.
[0040] The invention provides a system and method for synthesizing a fetal heartbeat. A sequence of ultrasound images of a fetus, acquired in a non-Doppler imaging mode and capturing a heart of the fetus, is received and processed to generate an audio signal representative of the heartbeat of the fetus.
[0041] Embodiments are at least partly based on the realization that characteristics of a fetal heartbeat may be extracted from non-Doppler ultrasound images of a fetal heart, and that these characteristics may be used to generate an audio signal representative of the fetal heartbeat. The inventors have recognized that an audio signal representative of the fetal heartbeat may provide a useful indication to less experienced operators that they have performed an ultrasound procedure correctly, as well as providing reassurance to expectant parents and facilitating an emotional connection to their fetus.
[0042] Illustrative embodiments may, for example, be employed in ultrasound systems, and, in particular, ultrasound systems used by inexperienced operators.
[0043] The general operation of an exemplary ultrasound system will first be described, with reference to Figure 1, and with emphasis on the signal processing function of the system since this invention relates to the processing of the signals measured by the transducer array.
[0044] The system comprises an array transducer probe 4 which has a transducer array 6 for transmitting ultrasound waves and receiving echo information. The transducer array 6 may comprise CMUT transducers; piezoelectric transducers, formed of materials such as PZT or PVDF; or any other suitable transducer technology. In this example, the transducer array 6 is a two-dimensional array of transducers 8 capable of scanning either a 2D plane or a three dimensional volume of a region of interest. In another example, the transducer array may be a ID array.
[0045] The transducer array 6 is coupled to a microbeamformer 12 which controls reception of signals by the transducer elements. Microbeamformers are capable of at least partial beamforming of the signals received by sub-arrays, generally referred to as "groups" or "patches", of transducers as described in US Patents 5,997,479 (Savord et al.), 6,013,032 (Savord), and 6,623,432 (Powers et al.).
[0046] It should be noted that the microbeamformer is entirely optional. Further, the system includes a transmit / receive (T / R) switch 16, which the microbeamformer 12 can be coupled to and which switches the array between transmission and reception modes, and protects the main beamformer 20 from high energy transmit signals in the case where a microbeamformer is not used and the transducer array is operated directly by the main system beamformer. The transmission of ultrasound beams from the transducer array 6 is directed by a transducer controller 18 coupled to the microbeamformer by the T / R switch 16 and a main transmission beamformer (not shown), which can receive input from the user's operation of the user interface or control panel 38. The controller 18 can include transmission circuitry arranged to drive the transducer elements of the array 6 (either directly or via a microbeamformer) during the transmission mode. In a typical line-by-line imaging sequence, the beamforming system within the probe may operate as follows. During transmission, the beamformer (which may be the microbeamformer or the main system beamformer depending upon the implementation) activates the transducer array, or a sub-aperture of the transducer array. The sub-aperture may be a one dimensional line of transducers or a two dimensional patch of transducers within the larger array. In transmit mode, the focusing and steering of the ultrasound beam generated by the array, or a sub-aperture of the array, are controlled as described below.
[0047] Upon receiving the backscattered echo signals from the subject, the received signals undergo receive beamforming (as described below), in order to align the received signals, and, in the case where a sub-aperture is being used, the sub-aperture is then shifted, for example by one transducer element. The shifted sub-aperture is then activated and the process repeated until all of the transducer elements of the transducer array have been activated.
[0048] For each line (or sub-aperture), the total received signal, used to form an associated line of the final ultrasound image, will be a sum of the voltage signals measured by the transducer elements of the given sub-aperture during the receive period. The resulting line signals, following the beamforming process below, are typically referred to as radio frequency (RF) data. Each line signal (RF data set) generated by the various sub-apertures then undergoes additional processing to generate the lines of the final ultrasound image. The change in amplitude of the line signal with time will contribute to the change in brightness of the ultrasound image with depth, wherein a high amplitude peak will correspond to a bright pixel (or collection of pixels) in the final image. A peak appearing near the beginning of the line signal will represent an echo from a shallow structure, whereas peaks appearing progressively later in the line signal will represent echoes from structures at increasing depths within the subject.
[0049] One of the functions controlled by the transducer controller 18 is the direction in which beams are steered and focused. Beams may be steered straight ahead from (orthogonal to) the transducer array, or at different angles for a wider field of view. The steering and focusing of the transmit beam may be controlled as a function of transducer element actuation time.
[0050] Two methods can be distinguished in general ultrasound data acquisition: plane wave imaging and “beam steered” imaging. The two methods are distinguished by a presence of the beamforming in the transmission (“beam steered” imaging) and / or reception modes (plane wave imaging and “beam steered” imaging). Looking first to the focusing function, by activating all of the transducer elements at the same time, the transducer array generates a plane wave that diverges as it travels through the subject. In this case, the beam of ultrasonic waves remains unfocused. By introducing a position dependent time delay to the activation of the transducers, it is possible to cause the wave front of the beam to converge at a desired point, referred to as the focal zone. The focal zone is defined as the point at which the lateral beam width is less than half the transmit beam width. In this way, the lateral resolution of the final ultrasound image is improved.
[0051] For example, if the time delay causes the transducer elements to activate in a series, beginning with the outermost elements and finishing at the central element(s) of the transducer array, a focal zone would be formed at a given distance away from the probe, in line with the central element(s). The distance of the focal zone from the probe will vary depending on the time delay between each subsequent round of transducer element activations. After the beam passes the focal zone, it will begin to diverge, forming the far field imaging region. It should be noted that for focal zones located close to the transducer array, the ultrasound beam will diverge quickly in the far field leading to beam width artifacts in the final image. Typically, the near field, located between the transducer array and the focal zone, shows little detail due to the large overlap in ultrasound beams. Thus, varying the location of the focal zone can lead to significant changes in the quality of the final image.
[0052] It should be noted that, in transmit mode, only one focus may be defined unless the ultrasound image is divided into multiple focal zones (each of which may have a different transmit focus).
[0053] In addition, upon receiving the echo signals from within the subject, it is possible to perform the inverse of the above described process in order to perform receive focusing. In other words, the incoming signals may be received by the transducer elements and subject to an electronic time delay before being passed into the system for signal processing. The simplest example of this is referred to as delay-and-sum beamforming. It is possible to dynamically adjust the receive focusing of the transducer array as a function of time.
[0054] Looking now to the function of beam steering, through the correct application of time delays to the transducer elements it is possible to impart a desired angle on the ultrasound beam as it leaves the transducer array. For example, by activating a transducer on a first side of the transducer array followed by the remaining transducers in a sequence ending at the opposite side of the array, the wave front of the beam will be angled toward the second side. The size of the steering angle relative to the normal of the transducer array is dependent on the size of the time delay between subsequent transducer element activations.
[0055] Further, it is possible to focus a steered beam, wherein the total time delay applied to each transducer element is a sum of both the focusing and steering time delays. In this case, the transducer array is referred to as a phased array.
[0056] In case of the CMUT transducers, which require a DC bias voltage for their activation, the transducer controller 18 can be coupled to control a DC bias control 45 for the transducer array. The DC bias control 45 sets DC bias voltage(s) that are applied to the CMUT transducer elements.
[0057] For each transducer element of the transducer array, analog ultrasound signals, typically referred to as channel data, enter the system by way of the reception channel. In the reception channel, partially beamformed signals are produced from the channel data by the microbeamformer 12 and are then passed to a main receive beamformer 20 where the partially beamformed signals from individual patches of transducers are combined into a fully beamformed signal, referred to as radio frequency (RF) data. The beamforming performed at each stage may be carried out as described above, or may include additional functions. For example, the main beamformer 20 may have 128 channels, each of which receives a partially beamformed signal from a patch of dozens or hundreds of transducer elements. In this way, the signals received by thousands of transducers of a transducer array can contribute efficiently to a single beamformed signal.
[0058] The beamformed reception signals are coupled to a signal processor 22. The signal processor 22 can process the received echo signals in various ways, such as: band-pass filtering; decimation; I and Q component separation; and harmonic signal separation, which acts to separate linear and nonlinear signals so as to enable the identification of nonlinear (higher harmonics of the fundamental frequency) echo signals returned from tissue and microbubbles. The signal processor may also perform additional signal enhancement such as speckle reduction, signal compounding, and noise elimination. The band-pass filter in the signal processor can be a tracking filter, with its pass band sliding from a higher frequency band to a lower frequency band as echo signals are received from increasing depths, thereby rejecting noise at higher frequencies from greater depths that is typically devoid of anatomical information.
[0059] The beamformers for transmission and for reception are implemented in different hardware and can have different functions. Of course, the receiver beamformer is designed to take into account the characteristics of the transmission beamformer. In Figure 1 only the receiver beamformers 12, 20 are shown, for simplicity. In the complete system, there will also be a transmission chain with a transmission micro beamformer, and a main transmission beamformer.
[0060] The function of the micro beamformer 12 is to provide an initial combination of signals in order to decrease the number of analog signal paths. This is typically performed in the analog domain.
[0061] The final beamforming is done in the main beamformer 20 and is typically after digitization.
[0062] The transmission and reception channels use the same transducer array 6 which has a fixed frequency band. However, the bandwidth that the transmission pulses occupy can vary depending on the transmission beamforming used. The reception channel can capture the whole transducer bandwidth (which is the classic approach) or, by using bandpass processing, it can extract only the bandwidth that contains the desired information (e.g. the harmonics of the main harmonic).
[0063] The RF signals may then be coupled to a B mode (i.e. brightness mode, or 2D imaging mode) processor 26 and a Doppler processor 28. The B mode processor 26 performs amplitude detection on the received ultrasound signal for the imaging of structures in the body, such as organ tissue and blood vessels. In the case of line-by-line imaging, each line (beam) is represented by an associated RF signal, the amplitude of which is used to generate a brightness value to be assigned to a pixel in the B mode image. The exact location of the pixel within the image is determined by the location of the associated amplitude measurement along the RF signal and the line (beam) number of the RF signal. B mode images of such structures may be formed in the harmonic or fundamental image mode, or a combination of both as described in US Pat. 6,283,919 (Roundhill et al.) and US Pat. 6,458,083 (Jago et al.) The Doppler processor 28 processes temporally distinct signals arising from tissue movement and blood flow for the detection of moving substances, such as the flow of blood cells in the image field. The Doppler processor 28 typically includes a wall filter with parameters set to pass or reject echoes returned from selected types of materials in the body.
[0064] The structural and motion signals produced by the B mode and Doppler processors are coupled to a scan converter 32 and a multi-planar reformatter 44. The scan converter 32 arranges the echo signals in the spatial relationship from which they were received in a desired image format. In other words, the scan converter acts to convert the RF data from a cylindrical coordinate system to a Cartesian coordinate system appropriate for displaying an ultrasound image on an image display 40. In the case of B mode imaging, the brightness of pixel at a given coordinate is proportional to the amplitude of the RF signal received from that location. For instance, the scan converter may arrange the echo signal into a two dimensional (2D) sector-shaped format, or a pyramidal three dimensional (3D) image. The scan converter can overlay a B mode structural image with colors corresponding to motion at points in the image field, where the Doppler-estimated velocities to produce a given color. The combined B mode structural image and color Doppler image depicts the motion of tissue and blood flow within the structural image field. The multi-planar reformatter will convert echoes that are received from points in a common plane in a volumetric region of the body into an ultrasound image of that plane, as described in US Pat. 6,443,896 (Detmer). A volume Tenderer 42 converts the echo signals of a 3D data set into a projected 3D image as viewed from a given reference point as described in US Pat. 6,530,885 (Entrekin et al.).
[0065] The 2D or 3D images are coupled from the scan converter 32, multi-planar reformatter 44, and volume Tenderer 42 to an image processor 30 for further enhancement, buffering and temporary storage for display on an image display 40. The imaging processor may be adapted to remove certain imaging artifacts from the final ultrasound image, such as: acoustic shadowing, for example caused by a strong attenuator or refraction; posterior enhancement, for example caused by a weak attenuator; reverberation artifacts, for example where highly reflective tissue interfaces are located in close proximity; and so on. In addition, the image processor may be adapted to handle certain speckle reduction functions, in order to improve the contrast of the final ultrasound image.
[0066] In addition to being used for imaging, the blood flow values produced by the Doppler processor 28 and tissue structure information produced by the B mode processor 26 are coupled to a quantification processor 34. The quantification processor produces measures of different flow conditions such as the volume rate of blood flow in addition to structural measurements such as the sizes of organs and gestational age. The quantification processor may receive input from the user control panel 38, such as the point in the anatomy of an image where a measurement is to be made.
[0067] Output data from the quantification processor is coupled to a graphics processor 36 for the reproduction of measurement graphics and values with the image on the display 40, and for audio output from the display device 40. The graphics processor 36 can also generate graphic overlays for display with the ultrasound images. These graphic overlays can contain standard identifying information such as patient name, date and time of the image, imaging parameters, and the like. For these purposes the graphics processor receives input from the user interface 38, such as patient name. The user interface is also coupled to the transmit controller 18 to control the generation of ultrasound signals from the transducer array 6 and hence the images produced by the transducer array and the ultrasound system. The transmit control function of the controller 18 is only one of the functions performed. The controller 18 also takes account of the mode of operation (given by the user) and the corresponding required transmitter configuration and band-pass configuration in the receiver analog to digital converter. The controller 18 can be a state machine with fixed states.
[0068] The user interface is also coupled to the multi-planar reformatter 44 for selection and control of the planes of multiple multi-planar reformatted (MPR) images which may be used to perform quantified measures in the image field of the MPR images.
[0069] Figure 2 illustrates an ultrasound system 200, according to an embodiment of the invention. The ultrasound system comprises an ultrasound imaging device 210 and a processing system 220. The processing system 220 is, itself, an embodiment of the invention.
[0070] The ultrasound imaging device 210 is capable of acquiring ultrasound imaging data in at least one non-Doppler-based imaging mode. For instance, the ultrasound imaging device may be capable of acquiring B-mode ultrasound imaging data and / or M-mode ultrasound imaging data. The ultrasound imaging device 210 may be a suitable conventional ultrasound imaging device, as described above with reference to Figure 1.
[0071] The ultrasound imaging device 210 is configured to acquire a sequence of nonDoppler mode ultrasound images 215 of a fetus 230, wherein a heart of the fetus is visible within a field of view of at least a subset of the sequence. The sequence of non-Doppler mode ultrasound images may be acquired in any imaging mode other than a Doppler-based mode; for example, the sequence of non-Doppler mode ultrasound images may be a sequence of B-mode ultrasound images or a sequence of M-mode ultrasound images. In some examples, the sequence of non-Doppler mode ultrasound images may be acquired during a blind sweep imaging procedure. During a blind sweep imaging procedure, an operator of the ultrasound imaging device follows a set path, typically starting and / or ending outside the uterus.
[0072] The processing system 220 is configured to receive the sequence of non-Doppler mode ultrasound images 215 acquired by the ultrasound imaging device 210. In Figure 2, the processing system receives the sequence of non-Doppler mode ultrasound images directly from the ultrasound imaging device; alternatively, the processing system may receive the sequence of non-Doppler mode ultrasound images via one or more other devices (e.g. a memory unit storing the sequence of non-Doppler mode ultrasound images).
[0073] Having received the sequence of non-Doppler mode ultrasound images 215, the processing system 220 is configured to process the sequence of non-Doppler mode ultrasound images to generate an audio signal 225 representative of one or more characteristics of a heartbeat of the fetus 230. Methods for generating the audio signal are described in more detail below.
[0074] The processing system 220 may be configured to generate the audio signal 225 during the ultrasound imaging procedure in which the sequence of non-Doppler mode ultrasound images is acquired. This allows the generated audio signal to be heard by a pregnant individual 240 carrying the fetus 230 (e.g. at the end of the ultrasound imaging procedure). For instance, the processing system receive each non-Doppler mode ultrasound images in the sequence as each image is acquired, and process the sequence to generate the audio signal as soon as the processing system has received enough non-Doppler mode images to estimate the heart rate of the fetus.
[0075] In some cases, the processing system 220 may be unable to generate the audio signal 225 because the heart of the fetus 230 is not visible in the sequence of non-Doppler mode ultrasound images 215 (or is not visible in enough images in the sequence to generate an accurate and reliable audio signal; as the skilled person will appreciate, the number of images in which the heart of the fetus must be visible will depend on the particular algorithm(s) used to generate the audio signal). The processing system 220 may, in response to failing to detect a heart in the sequence of non-Doppler mode ultrasound images (or detecting a heart but failing to detect cardiac activity), be configured to control a user interface to provide an output indicating that cardiac activity has not been detected, prompting an operator of the ultrasound imaging device 210 to acquire a new sequence of non-Doppler mode ultrasound images capturing the heart of the fetus, which may be received and processed by the processing system to generate the audio signal.
[0076] In some examples, the processing system 220 may be configured to process the sequence of non-Doppler mode ultrasound images 215 to identify, if present, abnormal features of the heartbeat of the fetus 230, and to generate the audio signal 225 representative of the one or more characteristics of the heartbeat of the fetus only in response to failing to identify any abnormal features of the heartbeat of the fetus. In other words, in some examples, an audio signal is not generated if one or more abnormal features are identified in the sequence of non-Doppler mode ultrasound images. In this way, the pregnant individual 240 is not alarmed by hearing an audio signal having abnormal features.
[0077] In some examples, in response to identifying one or more abnormal features, the processing system may control a user interface to provide an output indicating that one or more abnormal features have been identified (and, optionally, providing information regarding the identified abnormal feature(s), such as a type of abnormal feature). The output indicating that one or more abnormal features have been identified may be a type of output that is not perceptible to the pregnant individual 240 (e.g. the output may be a visual display, such as an image and / or textual display, displayed on a display device that is not visible to the pregnant individual), or may be provided after the ultrasound imaging procedure has concluded (so that the pregnant individual is no longer present).
[0078] The processing system 220 may identify, as an abnormal feature, any of failure to detect a heartbeat, an abnormally fast heartbeat, an abnormally slow heartbeat, and / or an irregular heartbeat. An abnormally fast or slow heartbeat may be identified based on clinically acceptable parameters; in some examples, thresholds for identifying abnormally fast and slow heartbeats may be defined based on a user input from a clinician (e.g. a midwife). For example, a lower threshold may be in the range 100-130 beats per minute, and an upper threshold may be in the range 180-200 beats per minute (where a heartbeat is identified as abnormally fast in response to exceeding the upper threshold and abnormally slow in response to falling below the lower threshold). Further abnormal features will be readily apparent to the skilled person.
[0079] The identification of any abnormal features may be performed following estimation of one or more characteristics of the heartbeat of the fetus for generating the audio signal, in examples in which the audio signal is generated by first estimating the one or more characteristics (see below). The estimated one or more characteristics may be used to determine whether any abnormal features are present. In examples in which a machine-learning algorithm is used to generate the audio signal (see below), the processing system may first process the sequence of non-Doppler ultrasound images 215 to identify whether any abnormal features are present, and may provide the sequence of non-Doppler ultrasound images as input to the machine-learning algorithm only in response to failing to identify any abnormal features.
[0080] Various methods for generating the audio signal 225 representative of the one or more characteristics of a heartbeat of the fetus 230 are envisaged. In some examples, the processing system 220 is configured to process the sequence of non-Doppler mode ultrasound images to estimate the one or more characteristics of the heartbeat of the fetus, and to process the estimated one or more characteristics to generate the audio signal representative of the one or more characteristics.
[0081] The one or more characteristics may, for example, comprise one or more of a heart rate of the fetus 230, an amplitude of the heartbeat of the fetus, a heart rate variability for the fetus, a frequency domain analysis for the heartbeat of the fetus, one or more time-varying characteristics of the heartbeat of the fetus and / or a shape of a waveform representing the heartbeat of the fetus.
[0082] The estimation of the one or more characteristics may depend on the imaging mode used to acquire the sequence of non-Doppler mode ultrasound images 215. For instance, if the sequence is a sequence of M-mode ultrasound images, the one or more characteristics may be estimated directly from the M-mode signal. For instance, a heart rate of the fetus may be estimated by extracting peaks using a discrete wavelet transform and determining the number of peaks over a particular time interval or by using a suitable machine-learning algorithm. A heart rate variability may be determined by measuring the intervals between adjacent peaks.
[0083] If the sequence 215 is a sequence of B-mode ultrasound images, the sequence may be processed using a segmentation algorithm to segment the heart of the fetus 230. Suitable segmentation algorithms may, for example, include machine-learning-based segmentation algorithms, such as a You Only Look Once (YOLO) neural network trained to segment a fetal heart in B-mode ultrasound images.
[0084] In some examples, the results of the segmentation algorithm may be used to produce a B-mode loop of the heart of the fetus, and the one or more characteristics of the heartbeat of the fetus may be determined from the B-mode loop using any suitable image processing technique (e.g. using an optical flow algorithm). Suitable techniques for determining the one or more characteristics of the heartbeat of the fetus will be readily apparent to the skilled person (see, for example: Wang et al. (2021), “Optical Flow Networks for Heartbeat Estimation in 4D Ultrasound Images”, Proceedings of the 2021 7th International Conference on Computing and Artificial Intelligence (ICCAI '21), 127-131; Ouzir et al. (2019), “Robust Optical Flow Estimation in Cardiac Ultrasound Images Using a Sparse Representation”, IEEE Trans Med Imaging, 38(3):7 1— 752; Luo and Liu (2011), “Optical Flow Computation Based Medical Motion Estimation of Cardiac Ultrasound Imaging”, 2011 5th International Conference on Bioinformatics and Biomedical Engineering, 1-4; and Duan et al. (2005), “Dynamic Cardiac Information From Optical Flow Using Four Dimensional Ultrasound”, 2005 IEEE Engineering in Medicine and Biology 27th Annual Conference, 4465-4468).
[0085] In some examples, the results of the segmentation algorithm may be used to generate one or more M-mode signals, from which the one or more characteristics may be estimated. For instance, an M-mode signal passing through the center of a bounding box for the segmented heart in each image in which the heart is detected may be generated, or a plurality of M-mode signals may be generated, each passing through the bounding box for the segmented heart, and a mean or median across the plurality of M-mode signals may be determined for each of the one or more characteristics. The one or more characteristics may be estimated using any suitable techniques for estimating characteristics of a fetal heartbeat from an M-mode signal (see above).
[0086] In some examples, the processing system 220 may be configured to obtain an audio recording of an example heartbeat (e.g. an example fetal heartbeat). The processing system may then be configured to process the estimated one or more characteristics to generate the audio signal by modulating the obtained audio recording according to the estimated one or more characteristics. In this way, a more realistic-sounding audio signal may be generated.
[0087] An audio recording of an example heartbeat may, for example, be obtained from a medical database or a library of heartbeat sounds. In the case of an example fetal heartbeat, the audio recording may have originally been acquired using Doppler mode ultrasound imaging; audio recordings of other example heartbeats may have been acquired using any suitable technique. The processing system may obtain the audio recording directly from a medical database or library, or from a memory unit (not shown in Figure 2) connected to the processing system (e.g.. a suitable audio recording may have previously been obtained from a medical database / library and stored in the memory unit). In some examples, the processing system may use the same audio recording of an example heartbeat each time the processing system is used to generate an audio signal of a fetal heartbeat; alternatively, the processing system may select one of a plurality of audio recordings, according to the estimated one or more characteristics (e.g. the processing system may select a most similar audio recording from the plurality of audio recordings).
[0088] As the skilled person will readily appreciate, the modulation of the obtained audio recording will depend on the one or more characteristics estimated. For instance, if the estimated one or more characteristics includes a heart rate for the fetus 230, the frequency of the obtained audio recording may be modulated to match the estimated heart rate.
[0089] In some examples, the processing system 220 may be configured to generate the audio signal 225 by processing the sequence of non-Doppler mode ultrasound images 215 using a machine-learning algorithm.
[0090] A machine-learning algorithm is any self-training algorithm that processes input data in order to produce or predict output data. Here, the input data comprises the sequence of non-Doppler mode ultrasound images 215. The output data may comprise a synthesized audio signal representative of one or more characteristics of the heartbeat of the fetus 230. Alternatively, the output data may comprise a synthesized sequence of Doppler mode ultrasound images, which may then be processed (e.g. using existing techniques for generating an audio signal using Doppler mode ultrasound images) to generate the audio signal 225.
[0091] Suitable machine-learning algorithms for being employed in the present invention will be apparent to the skilled person. Examples of suitable machine-learning algorithms include decision tree algorithms and artificial neural networks. Other machinelearning algorithms such as logistic regression, support vector machines or Naive Bayesian models are suitable alternatives.
[0092] The structure of an artificial neural network (or, simply, neural network) is inspired by the human brain. Neural networks are comprised of layers, each layer comprising a plurality of neurons. Each neuron comprises a mathematical operation. In particular, each neuron may comprise a different weighted combination of a single type of transformation (e.g. the same type of transformation, sigmoid etc. but with different weightings). In the process of processing input data, the mathematical operation of each neuron is performed on the input data to produce a numerical output, and the outputs of each layer in the neural network are fed into the next layer sequentially. The final layer provides the output.
[0093] A decision tree algorithm processes input data through a tree of nodes. In the tree of nodes, each successive node splits into two or more further nodes until reaching a terminal or end node. When performing the decision tree algorithm using the tree of nodes, at each node, a decision is made as to which further node to move to next based on the input data. The end node defines the outcome of the decision tree algorithm, and therefore the machinelearning algorithm.
[0094] Methods of training a machine-learning algorithm are well known. Typically, such methods comprise obtaining a training dataset, comprising training input data entries and corresponding training output data entries.
[0095] For some machine-learning algorithms, such as a neural network, training is performed by applying an initialized machine-learning algorithm to each input data entry to generate predicted output data entries. An error between the predicted output data entries and corresponding training output data entries is used to modify the machine-learning algorithm. This process can be repeated until the error converges, and the predicted output data entries are sufficiently similar (e.g. ±1%) to the training output data entries. This is commonly known as a supervised learning technique.
[0096] For example, where the machine-learning algorithm is formed from a neural network, (weightings of) the mathematical operation of each neuron may be modified until the error converges. Known methods of modifying a neural network include gradient descent, backpropagation algorithms and so on.
[0097] Other approaches for training machine-learning algorithms (e.g., decision trees) are known in the art. For instance, decision trees are often trained using a decision tree builder or learning techniques, such as those set out by Suthaharan, Shan, and Shan Suthaharan. "Decision tree learning." Machine Learning Models and Algorithms for Big Data Classification: Thinking with Examples for Effective Learning (2016): 237-269 or Ruggieri, Salvatore. "Yadt: Yet another decision tree builder." 16th IEEE International Conference on Tools with Artificial Intelligence. IEEE, 2004.
[0098] The training input data entries correspond to example sequences of non-Doppler mode ultrasound images. The training output data entries may correspond to audio signals, each audio signal having been derived from a sequence of Doppler mode ultrasound images acquired in a same ultrasound imaging procedure as a corresponding sequence of non-Doppler mode ultrasound images in the training input data entries. Alternatively, the training output data entries may correspond to the sequences of Doppler mode ultrasound images, each acquired in a same ultrasound imaging procedure as a corresponding sequence of non-Doppler mode ultrasound images in the training input data entries (i.e. in order to train a machinelearning algorithm to output a synthesized sequence of Doppler mode ultrasound images, which may then be used to generate the audio signal). In some examples, a sequence of non-Doppler mode ultrasound images in the training input data entries and a corresponding sequence of Doppler mode ultrasound images in the training output data entries (or used to derive a training output data entry) may have been acquired simultaneously, . However, this may not always be possible, and in other examples, the sequence of non-Doppler mode ultrasound images and corresponding sequence of Doppler mode ultrasound images may have been acquired in close temporal proximity (e.g. one after the other), preferably during a time at which the fetus is at rest in order to reduce heart rate variation.
[0099] In yet another alternative, the training input data entries may have been obtained from fetuses in later stages of pregnancy, and the training output data entries may correspond to audio recordings acquired by a stethoscope, each audio recording having been acquired during an ultrasound imaging procedure in which a corresponding sequence of non-Doppler mode ultrasound images in the training input data entries was acquired.
[0100] In some examples, separate machine-learning algorithms may be trained for different imaging mode inputs (for instance, a first machine-learning algorithm may be trained using sequences of B-mode ultrasound images, while a second machine-learning algorithm may be trained using sequences of M-mode ultrasound images). Alternatively, a single machine-learning algorithm may be trained for more than one imaging mode inputs (e.g. a single machine-learning algorithm may be trained using both sequences of B-mode ultrasound images and sequences of M-mode ultrasound images).
[0101] Having generated the audio signal 225, the processing system 220 may be further configured to synchronize the sequence of non-Doppler mode ultrasound images 215 and the generated audio signal, and to control a user interface 250 to output the synchronized sequence of non-Doppler mode ultrasound images and generated audio signal. The synchronized sequence of non-Doppler mode ultrasound images output by the user interface may only include images in which the fetal heart is visible (i.e. the output sequence may consist of a subset of the acquired non-Doppler mode ultrasound images). In other words, an audiovisual clip of the heart of the fetus may be output by the user interface. For instance, the user interface may output an audio-visual loop of the beating fetal heart.
[0102] The user interface 250 may be any user interface comprising a display device and at least one speaker. For instance, the user interface may be a computer monitor comprising at least one speaker (or a computer monitor and at least one separate speaker), a display device and at least one speaker provided on an ultrasound console or a mobile device (e.g. a smart phone or tablet).
[0103] In some examples, the processing system 220 may be further configured to control the user interface 250 to output an indication that the generated audio signal 225 is synthetic (for example, as a notification before outputting the synchronized sequence of non-Doppler mode ultrasound images and generated audio signal, or as textual display provided alongside the sequence of non-Doppler mode ultrasound images). This may reduce a likelihood of misinterpretation of the audio signal.
[0104] In Figure 2, the pregnant individual 240 is carrying a single fetus. In some situations, the pregnant individual may be carrying more than one fetus, and, therefore, the sequence of non-Doppler mode ultrasound images may capture more than one fetal heart. Whether the processing system 220 is able to generate an audio signal for each fetus in the case of a multiple pregnancy will depend on the techniques used by the processing system to generate the audio signal.
[0105] For instance, in examples in which the processing system is configured to estimate one or more characteristics of the heartbeat of the fetus and to process the estimated one or more characteristics to generate the audio signal, some techniques for estimating the one or more characteristics may allow multiple fetus hearts to be differentiated from one another in the sequence of non-Doppler mode ultrasound images, and an audio signal may be generated for each fetus. Other techniques for estimating the one or more characteristics may only allow the detection of one of the fetal heartbeats; in these cases, an audio signal representative of one or more characteristics of a heartbeat of one of the fetuses may be generated.
[0106] Where a machine-learning algorithm is used to generate the audio signal, how the machine-learning algorithm processes a sequence of non-Doppler mode ultrasound images that capture more than one fetal heart will depend on the training of the machine-learning algorithm. For instance, the machine-learning algorithm may be trained to generate a separate audio signal for each fetal heart captured in the sequence of non-Doppler mode ultrasound images by including, in the training input data entries, sequences of non-Doppler mode ultrasound images that capture more than one fetal heart, and including, in the training output data entries, a corresponding training output data entry for each fetal heart.
[0107] In examples in which the sequence of non-Doppler mode ultrasound images captures more than one fetal heart and an audio signal is generated for each fetal heart, the processing system may, for example, be configured to receive a user input indicating which audio signal to output.
[0108] Figure 3 illustrates a computer-implemented method 300 for synthesizing a fetal heartbeat, according to an embodiment of the invention.
[0109] The computer-implemented method 300 begins at step 310, at which a sequence of non-Doppler mode ultrasound images of a fetus is received. A heart of the fetus is at least partially within a field of view of at least a subset of the sequence. The sequence of non-Doppler mode ultrasound images may have been acquired in any imaging mode other than Doppler-based imaging modes; for instance, the sequence may be a sequence of B-mode ultrasound images or a sequence of M-mode ultrasound images. In some examples, the sequence of non-Doppler mode ultrasound images may have been acquired during a blind sweep imaging procedure.
[0110] At step 320, the sequence of non-Doppler mode ultrasound images is processed to generate an audio signal representative of one or more characteristics of a heartbeat of the fetus. The audio signal may, for example, be generated using any of the techniques described above.
[0111] In some examples, the computer-implemented method 300 may further comprise a step 330 of synchronizing the sequence of non-Doppler mode ultrasound images and the generated audio signal, and a step 340 of controlling a user interface to output the synchronized sequence of non-Doppler mode ultrasound images and generated audio signal.
[0112] In some examples, the computer-implemented method 300 may further comprise a step (not shown in Figure 3) of processing the sequence of non-Doppler mode ultrasound images to identify, if present, any abnormal features of the heartbeat of the fetus. The audio signal may then be generated only in response to failing to identify any abnormal features of the heartbeat of the fetus.
[0113] It will be understood that the disclosed methods are computer-implemented methods. As such, there is also proposed a concept of a computer program comprising code means for implementing any described method when said program is run on a processing system.
[0114] The skilled person would be readily capable of developing a processing system for carrying out any herein described method. Thus, each step of a flow chart may represent a different action performed by a processing system, and may be performed by a respective module of the processing system.
[0115] As discussed above, the system makes use of a processing system to perform the data processing. The processing system can be implemented in numerous ways, with software and / or hardware, to perform the various functions required. The processing system typically employs one or more microprocessors that may be programmed using software (e.g. microcode) to perform the required functions. The processing system may be implemented as a combination of dedicated hardware to perform some functions and one or more programmed microprocessors and associated circuitry to perform other functions.
[0116] Examples of circuitry that may be employed in various embodiments of the present disclosure include, but are not limited to, conventional microprocessors, application specific integrated circuits (ASICs), and field-programmable gate arrays (FPGAs).
[0117] In various implementations, the processing system may be associated with one or more storage media such as volatile and non-volatile computer memory such as RAM, PROM, EPROM, and EEPROM. The storage media may be encoded with one or more programs that, when executed on one or more processing systems and / or controllers, perform the required functions. Various storage media may be fixed within a processing system or controller may be transportable, such that the one or more programs stored thereon can be loaded into a processing system.
[0118] Variations to the disclosed embodiments can be understood and effected by those skilled in the art in practicing the claimed invention, from a study of the drawings, the disclosure and the appended claims. In the claims, the word “comprising” does not exclude other elements or steps, and the indefinite article “a” or “an” does not exclude a plurality.
[0119] Functions implemented by a processing system may be implemented by a single processing system or by multiple separate processing units which may together be considered to constitute a “processor” . Such processing units may in some cases be remote from each other and communicate with each other in a wired or wireless manner.
[0120] The mere fact that certain measures are recited in mutually different dependent claims does not indicate that a combination of these measures cannot be used to advantage.
[0121] A computer program may be stored / distributed on a suitable medium, such as an optical storage medium or a solid-state medium supplied together with or as part of other hardware, but may also be distributed in other forms, such as via the Internet or other wired or wireless telecommunication systems.
[0122] If the term “adapted to” is used in the claims or description, it is noted the term “adapted to” is intended to be equivalent to the term “configured to”. If the term “arrangement” is used in the claims or description, it is noted the term “arrangement” is intended to be equivalent to the term “system”, and vice versa.
[0123] Any reference signs in the claims should not be construed as limiting the scope.
Claims
CLAIMS:
1. A processing system (220) for synthesizing a fetal heartbeat, the processing system being configured to:receive a sequence of non-Doppler mode ultrasound images (215) of a fetus (230), wherein a heart of the fetus is at least partially within a field of view of at least a subset of the sequence; andprocess the sequence of non-Doppler mode ultrasound images to generate an audio signal (225) representative of one or more characteristics of a heartbeat of the fetus.
2. The processing system (220) of claim 1, wherein the processing system is configured to generate the audio signal (225) by:processing the sequence of non-Doppler mode ultrasound images (215) to estimate the one or more characteristics of the heartbeat of the fetus (230); and processing the estimated one or more characteristics to generate the audio signal representative of the one or more characteristics.
3. The processing system (220) of claim 2, wherein the processing system is configured to process the estimated one or more characteristics to generate the audio signal (225) by:obtaining an audio recording of an example heartbeat; andmodulating the audio recording of the example heartbeat according to the estimated one or more characteristics to generate the audio signal.
4. The processing system (220) of any of claims 1 to 3, wherein the one or more characteristics comprise a heart rate of the fetus (230).
5. The processing system (220) of any of claims 1 to 4, wherein the one or more characteristics comprise an amplitude of the heartbeat of the fetus (230).
6. The processing system (220) of any of claims 1 to 5, wherein the one or more characteristics comprise a heart rate variability for the fetus (230).
7. The processing system (220) of claim 1, wherein the processing system is configured to generate the audio signal (225) by processing the sequence of non-Doppler mode ultrasound images (215) using a machine-learning algorithm.
8. The processing system (220) of claim 7, wherein the machine-learning algorithm has been trained using a training algorithm configured to receive an array of training inputs and known outputs, each training input corresponding to a respective known output, wherein each training input comprises a sequence of non-Doppler mode fetal ultrasound images, and each known output comprises an audio signal derived from a sequence of Doppler mode ultrasound images acquired in a same ultrasound imaging procedure as the corresponding training input.
9. The processing system (220) of any of claims 1 to 8, wherein the sequence of non-Doppler mode ultrasound images (215) is a sequence of B-mode ultrasound images.
10. The processing system (220) of any of claims 1 to 8, wherein the sequence of non-Doppler mode ultrasound images (215) is a sequence of M-mode ultrasound images.
11. The processing system (220) of any of claims 1 to 10, wherein the sequence of non-Doppler mode ultrasound images (215) was acquired during a blind sweep imaging procedure.
12. The processing system (220) of any of claims 1 to 11, wherein:the processing system is further configured to process the sequence of non-Doppler mode ultrasound images (215) to identify, if present, any abnormal features of the heartbeat of the fetus (230); andthe audio signal (225) is generated only in response to failing to identify any abnormal features of the heartbeat of the fetus.
13. The processing system (220) of any of claims 1 to 12, wherein the processing system is further configured to:synchronize the sequence of non-Doppler mode ultrasound images (215) and the generated audio signal (225); andcontrol a user interface (250) to output the synchronized sequence of nonDoppler mode ultrasound images and generated audio signal.
14. A computer-implemented method (300) for synthesizing a fetal heartbeat, the computer-implemented method comprising:receiving a sequence of non-Doppler mode ultrasound images (215) of a fetus (230), wherein a heart of the fetus is at least partially within a field of view of at least a subset of the sequence; andprocessing the sequence of non-Doppler mode ultrasound images to generate an audio signal (225) representative of one or more characteristics of a heartbeat of the fetus.
15. A computer program product comprising computer program code means which, when executed on a computing device having a processing system, cause the processing system to perform all of the steps of the method (300) according to claim 14.
Citation Information
Patent Citations
Phased array acoustic systems with intra-group processors
US5997479A
Beamforming methods and apparatus for three-dimensional ultrasound imaging using two-dimensional transducer array
US6013032A
Ultrasonic diagnostic imaging with blended tissue harmonic signals
US6283919B1
Method for creating multiplanar ultrasonic images of a three dimensional object
US6443896B1
Ultrasonic harmonic imaging with adaptive image formation
US6458083B1