Method for generating sequences of synthetic ultrasound images of organs

The method iteratively applies a video diffusion model to generate long sequences of synthetic ultrasound images, addressing data scarcity and hardware limitations, achieving high-quality, consistent sequences for neural network training.

JP2026136085APending Publication Date: 2026-08-25DASSAULT SYSTEMES SA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2026016555
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-13
Filing Date
2026-02-04
Publication Date
2026-08-25

AI Technical Summary

Technical Problem

The availability of sufficient ultrasound image data for training neural networks is limited due to privacy concerns, operator dependence, equipment inconsistencies, and the scarcity of pediatric cardiac data, and existing video diffusion models are computationally intensive, processing only a limited number of frames.

Method used

A method using a video diffusion model trained on sequences of two-dimensional organ representations with semantic labels, iteratively applying the model to generate long sequences of synthetic ultrasound images through noise blending, allowing for the generation of high-quality, consistent sequences.

Benefits of technology

The method generates realistic, long-frame synthetic ultrasound sequences that maintain organ appearance and background consistency, overcoming hardware limitations and providing sufficient data for neural network training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026136085000001_ABST
    Figure 2026136085000001_ABST
Patent Text Reader

Abstract

Improvement of solutions for generating ultrasound images of organs. [Solution] A method for generating a sequence of synthetic ultrasound images of organs includes a video diffusion model that takes a sequence of multiple 2D representations of organs, each assigned a semantic label, and a corresponding sequence of multiple noise images as input, and is trained to generate a corresponding denoised sequence of ultrasound images that respects the semantic labels. The method provides multiple consecutive sequences of multiple 2D representations and multiple consecutive sequences of multiple noise images, iteratively applies the video diffusion model to each pair consisting of one consecutive sequence and one corresponding sequence of noise images, and for each denoising step of the reverse processing of the video diffusion model, a portion of the denoised synthetic ultrasound image generated for the denoising step in the previous iteration is used to replace a portion of any pair of images for the denoising step.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer programs and systems, and more specifically, to a method, system and program for generating a sequence of synthetic ultrasound images of organs.

Background Art

[0002] Ultrasonic imaging is widely used and important in the medical field, and ultrasonic images of organs are used, for example, to obtain medical data (e.g., indirect physiological measurements related to these organs). Modern tools in the medical field include software solutions for processing such images. For example, a neural network may be used to process ultrasonic images of the heart, for example, to obtain cardiac segmentation and / or to detect abnormalities.

Summary of the Invention

Problems to be Solved by the Invention

[0003] However, despite the increasing importance in the medical field, the availability of ultrasonic image data remains an issue. For example, the availability of sufficient amounts of such data for training neural networks is often limited for the following reasons.

[0004] · Privacy concerns: strict patient privacy regulations for anonymizing patient-specific information.

[0005] · Operator dependence: the acquisition of high-quality images and related anatomical and functional measurements depends heavily on the operator's expertise.

[0006] · Limited standardization: differences in equipment, settings, and protocols between facilities and vendors can lead to inconsistencies in the collected data and the resulting clinical management.

[0007] • Scarcity of pediatric cardiac data: The availability of pediatric cardiac data is particularly limited, making it difficult to develop a comprehensive model for this patient group.

[0008] Furthermore, some organs require observation over time (i.e., long sequences of ultrasound frames / images of organs need to be observed over time), necessitating ultrasound image data that captures the entire required observation time. In other words, a single image is insufficient; a series of images is required. This applies not only to organs characterized by repeating periods over time (e.g., the heart and arteries) but also to non-periodic organs (e.g., those used for observing needle insertion or stent / valve implantation). Neural network models, such as video diffusion models, exist to generate composite sequences (or frames) of images, recreating one period of video data. However, in the case of organ ultrasound imaging, the data in question is computationally intensive, and due to the limitations of existing GPUs, these models can only process a limited number of images within a sequence, a number considerably less than the number of frames required to accurately model long frame sequences.

[0009] In this context, improvements to the solution are needed for generating ultrasound images of organs. [Means for solving the problem]

[0010] Therefore, a method for generating a sequence of synthetic ultrasound images of an organ, performed by a computer, is provided. The method provides a trained video diffusion model. The video diffusion model is trained to take as input a sequence of multiple two-dimensional representations of the organ and a corresponding sequence of multiple noise images. Each two-dimensional representation is assigned a plurality of semantic labels. The video diffusion model is trained to generate a corresponding denoised sequence of multiple synthetic ultrasound images of the organ, respecting the semantic labels. The method further comprises the step of providing a plurality of consecutive sequences of multiple two-dimensional representations of the organ, each assigned a plurality of semantic labels, and a corresponding sequence of noise images. The method further comprises the step of iteratively applying the video diffusion model to each pair consisting of a sequence of multiple two-dimensional representations of the organ, each assigned a plurality of semantic labels, and a corresponding sequence of noise images. The process of iteratively applying the video diffusion model comprises, when applying the video diffusion model to any pair, the step of using a portion of the denoised synthetic ultrasound image generated in the previous iteration for the denoising step in each denoising step of the reverse processing of the video diffusion model, in the denoising step, by replacing a portion of the image of the any pair for the denoising step. The method may hereafter be referred to as the "noise blending method" or simply as "noise blending".

[0011] The above method may include one or more of the following:

[0012] A portion of the denoised synthetic ultrasound image generated during the previous iteration is present in the denoised synthetic ultrasound image generated during the previous iteration, and a portion of any pair of images is present in the leading image. The total number of images in the aforementioned multiple consecutive sequences is less than 40, and the number of images in each sequence is 24 or less. The number of images in each sequence is 16. The number of images in a portion of the denoised synthetic ultrasound image generated in the previous iteration is the number in the tail portion of the denoised synthetic ultrasound image generated in the previous iteration, the number of images in a portion of any pair of images is the number in the leading portion of the image, and is between one-quarter and one-half of the number of images in each sequence, and / or The aforementioned organs are the heart, liver, kidneys, carotid arteries, aorta, or coronary arteries.

[0013] Furthermore, the present invention provides a database (stored, for example, on a non-temporary recording medium) characterized by comprising multiple sequences of synthetic ultrasound images of organs generated by the above method.

[0014] The present invention also provides a method for using the database, which is performed by a computer. The method of use comprises the step of training a neural network based on the database. The neural network is trained for organ segmentation based on a single sequence of multiple ultrasound images of the organ, organ tracking based on a single sequence of multiple ultrasound images of the organ, organ detection based on a single sequence of multiple ultrasound images of the organ, registration tasks based on a single sequence of multiple ultrasound images of the organ, or classification tasks based on a single sequence of multiple ultrasound images of the organ.

[0015] Furthermore, a neural network that can be obtained by the aforementioned method of use is provided.

[0016] Furthermore, the present invention provides a computer program that includes instructions for carrying out the above-mentioned method and / or the above-mentioned method of use.

[0017] Furthermore, the present invention provides a device comprising a data storage medium on which the computer program and / or the neural network and / or the database are recorded.

[0018] The device may form a temporary or non-temporary computer-readable medium, or may be a temporary or non-temporary computer-readable medium, for example, in SaaS (Software as a Service), other servers, cloud-based platforms, etc. Further, the device may include a processor connected to a data storage medium. Therefore, the device may form an entire computer system or a part thereof (for example, the device is a subsystem of the entire system, etc.). The system may further include a graphical user interface connected to the processor.

Brief Description of Drawings

[0019] [Figure 1] It is a diagram showing a video diffusion model composed of two processes, a forward process and a reverse process, in the noise blending method. [Figure 2] It is a flowchart showing the noise blending method. [Figure 3] It is a diagram showing a latent diffusion model. [Figure 4] It is a diagram showing how the video diffusion model is repeatedly applied. [Figure 5] It is a diagram showing frame replacement in the noise blending method. [Figure 6] It is a diagram comparing the y-t cross-sections of synthetic ultrasound images of a specific patient generated by the noise blending method and the sequence-based conditioning method. [Figure 7] It is a diagram showing the Dice score and IoU in terms of frames. [Figure 8] It is a diagram showing an apical three-chamber 2D representation with semantic labels. [Figure 9] It is a diagram showing two apical three-chamber images generated by the pixel diffusion model and the latent diffusion model. [Figure 10] It is another diagram showing two apical three-chamber images generated by the pixel diffusion model and the latent diffusion model. [Figure 11]A diagram showing the process of obtaining an intravascular ultrasound sequence and characterizing and measuring coronary artery calcification. [Figure 12] A block diagram showing an example of the system.

Best Mode for Carrying Out the Invention

[0020] Non-limiting examples will be described below with reference to the accompanying drawings.

[0021] A computer-executed method for generating a sequence of synthetic ultrasound images of an organ is proposed. The method of the present disclosure provides a trained video diffusion model. This video diffusion model is trained to receive, as inputs, one sequence of a plurality of two-dimensional representations of an organ and one sequence of corresponding plurality of noisy images. Each two-dimensional representation is assigned a plurality of semantic labels. Such a video diffusion model is trained to generate one sequence of denoised synthetic ultrasound images of the corresponding organ that respects the semantic labels. The method of the present disclosure further comprises the step of providing a plurality of consecutive sequences of the plurality of two-dimensional representations of the organ, each assigned a plurality of semantic labels, and a plurality of consecutive sequences of corresponding noisy images. The method of the present disclosure further comprises the step of repeatedly applying the video diffusion model to each pair composed of one consecutive sequence of the plurality of two-dimensional representations of the organ, each assigned a plurality of semantic labels, and one corresponding sequence of one noisy image. The step of repeatedly applying the video diffusion model comprises, when applying the video diffusion model to any pair, using, in the replacement of a part of the images of the arbitrary pair for the denoising process, a part of the denoised synthetic ultrasound image generated for the denoising process in the previous iteration in each denoising process of the reverse processing of the video diffusion model.

[0022] This provides an improved solution for generating synthetic ultrasound images of an organ.

[0023] In practice, the method of this disclosure uses a pre-trained video spread model to generate a synthetic ultrasound image based on noise images, semantic labels associated with organs, and labeling of 2D representations of those organs. Therefore, the synthetic ultrasound image generated by the method of this disclosure is more realistic because its generation is guided by the semantic labels associated with the target organ. On the other hand, video spread models have a limited number of images or frames that they can process simultaneously to generate a sequence. Typically, these models are trained on data of sequences of organ frames, in which case the number of frames in each sequence is small, typically 16, or 24 if very efficient hardware is available. This is because the target data, i.e., video data, is high-dimensional and computationally intensive, even when considered as images / frames that make up a sequence. For this reason, existing hardware, typically GPUs, cannot train video spread models on such computationally intensive data unless the number of frames is limited to, for example, 16, or rarely 24. Nevertheless, in organs such as the heart, liver, kidneys, carotid arteries, aorta, and coronary arteries, for example, these organs have periodic characteristics that repeat over time, and more than 16 or 24 frames are already required for one cycle to occur, or for other long-term observations (such as the aforementioned needle insertion or stent / valve implantation), it is necessary to generate a composite sequence of frames longer than 16 or 24 frames.

[0024] To achieve this, the method of the present disclosure assumes a long sequence that is of a size that can be processed by the video spread model and is divided into a series of sequences corresponding to a noise image, each of which is a 2D representation of a labeled organ, and iteratively applies the trained video spread model to each individual sequence (a labeled 2D representation corresponding to a noise image) that makes up the long sequence. Next, in each iteration, i.e., when the model is applied to an individual sequence, for that iteration, in each denoising step in the reverse processing of the video spread model, the method of the present disclosure replaces a portion of the image of the individual sequence currently being processed (i.e., the sequence to which the model is currently applied) (e.g., the beginning of the sequence) with an image of the individual sequence that has already been (partially) denoised in the denoising step in the previous iteration (e.g., the end of the sequence). In other words, as is known from the video diffusion model itself, the video diffusion model takes as input a single individual sequence of multiple 2D representations of an organ and a corresponding noise image, adds the noise image stepwise to the corresponding 2D representation image (forward processing), and then iteratively denoises the result (in reverse processing), in which denoising is performed in each iteration of denoising (each denoising step) to further denoise all images of the sequence until the sequence is completely denoised. What is done in the method of this disclosure is that when applying the model to each individual sequence (i.e., naturally the second and subsequent sequences), in each denoising step, a (partially) denoised sequence obtained from the denoising step of the reverse processing of applying the model to the previous individual sequence is partially used to replace a portion of the image of the current sequence being denoised in that denoising step. This is referred to in this disclosure as "noise blending".

[0025] This noise blending allows for the iterative generation of consecutive sequences while maintaining consistency and smooth transitions between them, because the noise blending conditions the generation of each intermediate noise frame of a new sequence during denoising (reverse processing) on ​​the intermediate noise frames of the previously generated sequence. By generating smooth and consistent sequences over time without scene cuts, it becomes possible to synthetically reproduce real-world clinical ultrasound data. In this context, consistency refers to maintaining the appearance of the object and background color throughout the sequence, and avoiding disturbances. Therefore, the method of this disclosure can generate realistic, high-quality, long-frame synthetic sequences of organs (long enough to cover one cycle of an organ) even when using a video spread model that is limited in frame count by hardware-specific constraints (e.g., existing GPUs).

[0026] The method of this disclosure is for generating a sequence of synthetic ultrasound images of an organ. In other words, the method of this disclosure takes as input a plurality of consecutive sequences of plurality of two-dimensional representations of an organ, each with a semantic label, and a corresponding plurality of consecutive sequences of noise images, and then generates a corresponding long sequence of synthetic denoised ultrasound images of the organ that respect the labeling (through iterative application of a video diffusion model). This sequence is said to be "long" because it corresponds to the concatenation of a plurality of consecutive sequences, also referred to as "individual sequences." These individual sequences are "short" in that they have a large number of frames (the same number in all sequences) that are small enough to be processed by the video diffusion model. A "synthetic ultrasound image" refers to an image that is synthesized (computer-generated) rather than acquired by an ultrasound imaging device, but that reproduces the features of an ultrasound image as accurately as possible. In this disclosure, since the synthetic image is generated by a video diffusion model (processed by the method of this disclosure in the noise blending step as described above), the level of accuracy at which the synthetic image resembles an actual ultrasound image corresponds to the level of accuracy provided by the video diffusion model. "Denoised" means that the image is free of noise.

[0027] The method of this disclosure comprises providing a trained video spreading model. Therefore, the method of this disclosure provides (receives) a video spreading model that has already been trained as input.

[0028] The video spread model is well known. As shown in Figure 1, the video spread model consists of two processes: forward and reverse. The forward process consists of stepwise noise addition to the input sequence over multiple steps, for example, T=1000 steps, until pure Gaussian noise is obtained. The reverse process consists of stepwise denoising using a trained neural network over the same number of steps, ultimately generating a noise-free sequence.

[0029] In this embodiment, assuming a fixed number of steps T (e.g., T=1000) and a fixed number of frames f (e.g., f=16), the model performs the following processing. It receives as input: 1) a sequence of f 2D representations of an organ, each with multiple semantic labels, and 2) a sequence of f corresponding noise images (i.e., one noise image for each 2D representation in the sequence). In practice, these sequences may be arranged as tensors, as described later. The video diffusion model then iteratively removes noise from the noise images (reverse processing) (in T iterations) until it obtains a composite ultrasound image (a corresponding denoised sequence of multiple composite ultrasound images of an organ that respects multiple semantic labels) as a result of denoising the noise images so that the noise images correspond to the 2D representations of the organ, under constraints on the 2D representations and their semantic labels.

[0030] Forward processing is used only during training. Subsequently, only inverse processing is used during inference (especially when using the model in the manner of this disclosure). Specifically, during training, the model encounters multiple training examples, each consisting of 1) a sequence of multiple 2D representations labeled with the same type of organ as the sequences the model encounters during inference / use, and 2) a corresponding sequence of multiple measured ultrasound images of the organ corresponding to those 2D representations. During training, upon encountering a training example, the model first applies deterministic forward processing to iteratively add noise to the actual ultrasound images given as training examples (over T steps). Subsequently, the model applies inverse processing, which is a neural network that removes noise from the results of the forward processing. This training consists of adjusting the weights / parameters of the neural network so that the denoised image corresponds to the actual ultrasound image in terms of texture and respects the multiple labels of the multiple 2D representations given as training examples.

[0031] As shown in Figure 1, in computer execution, the model assumes that the conditional data distribution is Assuming we follow TIFF2026136085000002.tif7170, the likelihood The goal is to maximize TIFF2026136085000003.tif7170. Here, y represents the true sequence of the 2D representation labeled with ground truth, and s0 represents the sequence of the synthetic ultrasound image to be generated. Forward processing TIFF2026136085000004.tif7170 is a true image with Gaussian noise. TIFF2026136085000005.tif7170 schedule difference This is a step-by-step process that follows TIFF2026136085000006.tif7170. TIFF2026136085000007.tif7170

[0032] To satisfy TIFF2026136085000008.tif7170, You may choose a high value for TIFF2026136085000009.tif7170. Then, the model's neural network (reverse processing) aims to learn the reverse processing. TIFF2026136085000010.tif11170

[0033] t=T, ..., 1. Gaussian noise TIFF2026136085000011.tif7170 represents a valid ultrasound image (as well as the corresponding ground truth image). The image is denoised iteratively until TIFF2026136085000012.tif7170 is obtained.

[0034] Appropriate known video diffusion model architectures may be used. In computer execution, the neural network portion (reverse processing) uses a modified version of the semantic diffusion model (SDM) presented in the reference Semantic image synthesis via diffusion models. arXiv preprint arXiv:2207.00050 (2022) (incorporated herein by reference). In fact, SDM was originally designed for image generation. However, in these executions, its architecture is applied for sequence generation by adding a temporal attention block after each spatial attention block, as presented in the reference Ho et al., Video Diffusion Models, NeurIPS 2022 (incorporated herein by reference), and as shown in Figure 2, which illustrates the overall model architecture in these executions. The temporal attention block is used to address the temporal consistency and long-range dependency between the start and end frames in a 16-frame sequence. As shown in Figure 2, semantic labels of the conditioned 2D representation are injected into the decoder. This may be done using the known method of spatial adaptive normalization ("SPADE"). TIFF2026136085000013.tif8170

[0035] Here, TIFF2026136085000014.tif7170 contains the input / output features of SPADE. TIFF2026136085000015.tif7170 is group normalized. TIFF2026136085000016.tif7170 represents the spatially adaptive weights and biases learned from the semantic labels, respectively.

[0036] During inference, at each step, noise is estimated from the input noisy sequence by a semantic label-conditional denoising network. This network extracts embeddings from the semantic labels and injects them into the noisy data by applying learnable transformations, with spatially adaptive weights and biases learned from the semantic layout. Using the estimated noise, a less noisy sequence is generated by formulating posterior probabilities. This iterative process ensures the generation of realistic outputs that are faithfully derived from the semantic labels.

[0037] This model may directly process the pixel space of an ultrasound image (Pixel Video Diffusion Model (PVDM)). Alternatively, as shown in Figure 3, it may process the encoding of the pixel image into the latent space after applying a pre-trained variational autoencoder (Latent Video Diffusion Model (LVDM)). The pixel diffusion model is a diffusion model that operates directly in the high-dimensional pixel space of an image. As mentioned above, the method of this disclosure includes adding noise to the pixel values ​​in steps during forward processing and removing it during reverse processing. While the method of this disclosure is accurate, the large size of the pixel data results in a high computational load and consumes a large amount of memory. The latent diffusion model is an alternative to this memory-intensive method. This new method is a diffusion model that operates in a low-dimensional latent space and is usually derived from a compressed representation of the data using a pre-trained variational autoencoder (VAE). Since the final image is reconstructed from the denoised latent representation using a decoder of the same VAE, the latent diffusion model significantly reduces computational requirements while maintaining high fidelity by performing diffusion in this reduced space. For example, if the original data resolution is 256x256, the pixel spread model will produce an image / sequence of the same resolution. However, in the latent spread model, the VAE compresses the data to a lower resolution, 64x64 in this invention. The spread processing is then performed on this lower resolution empty, generating a 64x64 latent representation. The VAE's decoder maps the generated latent data to a high resolution of 256x256.

[0038] As stated above, the video spreading model is provided as input to the method of this disclosure in a pre-trained state. Therefore, providing the video spreading model may involve retrieving / obtaining the model for training from memory, a server, or a database (e.g., remote) where the model is stored. Alternatively, the method of this disclosure may involve training the model by any suitable known training method.

[0039] In this disclosure, a two-dimensional representation of an organ is a two-dimensional image showing the shape of a two-dimensional image of an organ, such as a cross-sectional view of the organ (e.g., a two-dimensional cross-sectional view of the left atrium and left ventricle of the heart). The two-dimensional image is identical in all representations relating to the methods of this disclosure. The two-dimensional image may correspond to an image captured by ultrasound imaging, but only represents the corresponding shape of the organ. In other words, it represents the shape that corresponds to what is observed in ultrasound imaging, but does not include the texture of the ultrasound imaging. The 2D image may be obtained from a mesh representation of the organ, or from a mesh representation of the organ by any appropriate method. The 2D image may consist of a set of 2D pixels that represent the shape of the organ from the viewpoint of the 2D image. The 2D representation may be labeled as follows: each pixel is labeled with a semantic label from a predefined list, which consists of a background label (indicating the absence of an organ) and one label for each region of interest of the organ. Therefore, each pixel is labeled with a label indicating whether or not a part of an organ is represented, and if so, which part is represented. For this reason, the 2D representation is also called a semantic label map. For example, if the organ is the heart, the labels may be "background," "left ventricular blood pool," "left ventricular myocardium," and "atrial blood pool" (if only the left ventricle and atria are of interest). Alternatively, the "background" label may be replaced by not assigning a label. However, these are just examples, and any meaningful semantic labels for an organ can be used. For a particular organ, the same set of semantic labels is used throughout the entire sequence involved in the method and video diffusion model training of this disclosure.

[0040] In this disclosure, a two-dimensional representation is always a sequence of N representations (where N is always the same for all sequences in the methods of this disclosure). "Sequence" means that the N two-dimensional representations correspond to the consistent time evolution of an organ (for example, the sequence represents a part of the organ's cycle), i.e., the first frame of the sequence corresponds to the organ at a particular point in time, the second frame corresponds to the organ one time step after that point in time, the third frame corresponds to the organ two time steps after that point in time, and so on.

[0041] In this specification, all sequences of noise images always correspond to a two-dimensional representation sequence, and the number of frames / images in the noise image sequence is the same as the number of frames / images in the two-dimensional representation sequence. The nature of the noise images and any differences between them are not important; what matters is that by denoising each noise image based on the two-dimensional representation sequence, the noise images become the corresponding ultrasound images of the organs, respecting the labels. However, it may be necessary for all sequences of noise images in this specification to correspond to (e.g., be identical to) the sequences obtained by the video diffusion model after forward processing, and if forward processing generates sequences of noise images corresponding to Gaussian noise, then all sequences of noise images in this specification must similarly correspond to Gaussian noise. For example, sequences of noise images in this specification may be formed by a Gaussian noise tensor. The Gaussian noise tensor is, It may also be in the shape of TIFF2026136085000017.tif7170, and here, TIFF2026136085000018.tif7170 represents the number of channels, number of frames, height, and width, respectively. As mentioned above, when applying the video spread model (more precisely, its inverse), this tensor is iteratively denoised until a sequence of identical shapes is obtained that represents the sequence of organs corresponding to the sequence of 2D representations. The number of denoising steps used during inference is s. When all tensors obtained during the denoising process are grouped together, a tensor with a shape s times larger than the initial shape of the tensor is obtained. That is, The file TIFF2026136085000019.tif7170 is obtained.

[0042] The method of this disclosure also comprises providing (i.e., as input) a plurality of consecutive sequences (i.e., a plurality of individual sequences) of a plurality of two-dimensional representations of an organ, each assigned a plurality of semantic labels (i.e., each individual sequence is of the same type as the sequence trained to receive as input) and a plurality of corresponding sequences of noise images (i.e., one sequence of noise images for each individual sequence). "Consecutive" means that the sequences follow one another to show a consistent temporal evolution of the organ (each image / frame represents an increment of a particular time step). The consecutive sequences represent, for example, one cycle of the organ. Providing consecutive sequences may comprise, for example, providing a single long sequence of two-dimensional representations (and corresponding noise images) that includes a period corresponding to, or at least equivalent to, the time length of one cycle of the organ, and dividing it into individual sequences. The long sequence may be obtained from a relevant dataset (such as the datasets described below).

[0043] In addition to providing input, the method of this disclosure further comprises iteratively applying a video spread model to sequence pairs formed by individual sequences of the provided 2D representations and their corresponding noise image sequences. Iterative application of the model comprises, in each iteration, applying the video spread model to a given sequence pair, using a portion of the denoised synthetic ultrasound image generated in the previous iteration for the denoising step in the denoising step of the reverse processing of the video spread model, in which a portion of the image of the given sequence pair is replaced for that denoising step. In other words, the portion of the denoised synthetic ultrasound image generated in the previous iteration may consist of the tail portion of the denoised synthetic ultrasound image generated in the previous iteration, and the portion of the image of the given sequence pair may consist of the beginning portion of that image. The replacement in this iteration will be described with reference to Figure 4.

[0044] In the first iteration, the first individual sequence of a set of consecutive sequences is considered. The first individual sequence is a sequence of 2D representations of organs with f labels (f=16 in the example shown in Figure 4). TIFF2026136085000020.tif7170 and the corresponding sequence of noise images It consists of TIFF2026136085000021.tif7170. The first iteration is: TIFF2026136085000022.tif7170 should respect the labeling. Applying the inverse processing of the video spread model to TIFF2026136085000023.tif7170 as a condition Stepwise denoising of TIFF2026136085000024.tif7170 and the corresponding ultrasound image sequence This consists of obtaining TIFF2026136085000025.tif7170. TIFF2026136085000026.tif7170 represents all intermediate results obtained iteratively in each denoising step. Thus, the first iteration consists only of applying the model to the first sequence, i.e., no frame replacement ("noise blending") is performed at this point.

[0045] In the second iteration, the second distinct sequence in a set of consecutive sequences is considered. The second distinct sequence is a sequence of 2D representations of the following f labeled organs (in the example shown in Figure 4, f remains at 16). TIFF2026136085000027.tif7170 and the corresponding sequence of noise images. It consists of TIFF2026136085000028.tif7170. The second iteration is, TIFF2026136085000029.tif7170 should respect the labeling. Applying the inverse processing of the video diffusion model to TIFF2026136085000030.tif7170 as a condition Stepwise denoising of TIFF2026136085000031.tif7170 and the corresponding ultrasound image sequence TIFF2026136085000032.tif7170 is composed of obtaining this. TIFF2026136085000033.tif7170 represents all intermediate results obtained iteratively in each denoising step. As shown in Figure 4, at the time of the second iteration, frame replacement is performed (for each i between T and 1) From TIFF2026136085000034.tif7170 In the noise reduction process to obtain TIFF2026136085000035.tif7170, A portion of the frame of TIFF2026136085000036.tif7170 (in Figure 4, the first half, but other options may be considered) Replace with a portion of the frame of TIFF2026136085000037.tif7170 (in Figure 4, the last / second half, but other options may be considered here as well), and then this sequence Further noise reduction processing was performed on TIFF2026136085000038.tif7170, and in the process, a portion of the frame was deformed. By replacing a portion of the frame in TIFF2026136085000039.tif7170 The file TIFF2026136085000040.tif7170 is obtained.

[0046] Although not shown in Figure 4, the following individual sequences Perform the same process for TIFF2026136085000041.tif7170 (for each i between T and 1), From TIFF2026136085000042.tif7170 In each noise reduction step to obtain TIFF2026136085000043.tif7170, A portion of the frame of TIFF2026136085000044.tif7170 (for example, the first half) Replace a portion of the frame in TIFF2026136085000045.tif7170 (for example, the latter half), and then this sequence Further noise reduction processing was performed on TIFF2026136085000046.tif7170, and in the process, a portion of the frame was deformed. By replacing a portion of the TIFF2026136085000047.tif7170 frame The file TIFF2026136085000048.tif7170 is obtained. All sequences in which video diffusion processing is performed consecutively. The following individual sequences will be applied until TIFF2026136085000049.tif7170 is applied. Repeat the same process for TIFF2026136085000050.tif7170.

[0047] A portion of the denoised synthetic ultrasound image generated in the previous iteration may consist of a plurality of denoised synthetic ultrasound images at the end generated in the previous iteration, and a portion of the images in the predetermined sequence pair may consist of a plurality of images at the beginning. In other words, continuing the above description with reference to Figure 4, for all m between 1 and M and i between T and 1, From TIFF2026136085000051.tif7170 In each noise reduction step to obtain TIFF2026136085000052.tif7170, The first few images in TIFF2026136085000053.tif7170 are It will be replaced by multiple images at the end of TIFF2026136085000054.tif7170. For example, The first n images in TIFF2026136085000055.tif7170 are The n trailing images of TIFF2026136085000056.tif7170 are replaced, where n is a number between, for example, f / 4 and f / 2.

[0048] The total number of images in multiple consecutive sequences (i.e., long sequences) The total number of images in the entire TIFF2026136085000057.tif7170 may be greater than 40 images. In other words, the total number of 2D representations in multiple consecutive sequences is greater than 40, and the total number of noise images is equal to the total number of 2D representations, and therefore greater than 40. The number of images in TIFF2026136085000058.tif7170 (i.e., the number of 2D representations equal to the number of noise images) may be, for example, 24 or less, or for example, 16.

[0049] It should be understood that the total number of 2D representations / noise images does not necessarily have to be a multiple of 16. If it is not, for example, if the total is 73 (this is merely an explanatory example), the method of this disclosure may proceed as follows. • Generate frames 1-16, and then, • Generate frames 13-28 and create an overlap of 4 frames (13-16), • Generate frames 25-40, create an overlap of 4 frames (25-28), and then, • Generate frames 37-52, create an overlap of 4 frames (37-40), and then, • Generate frames 49 through 64, create an overlap of 4 frames (49-52), then, • Generate frames 58-73 and create an overlap of 7 frames (58-64).

[0050] The above example of overlap is not unique and can be adapted to any total number of frames, as long as the overlap process is carefully selected to retain all frames. In this example, the overlap process is selected within [16 / 4=4, 16 / 2=8]. The overlap process may be selected by the user (for example, in an initial step of the method of this disclosure) or it may be automatically selected to comply with the above rule of retaining all frames.

[0051] Figure 5 further illustrates frame replacement (noise blending). As is clear from Figure 5, at once To generate a long ultrasonic sequence using a video diffusion model trained to generate only 170 frames, shape Generate the first sequence of TIFF2026136085000060.tif7170. Generated during denoising. All 170 intermediate tensors are preserved. The next step consists of extending the length of the current first sequence. Overlap is used for this purpose. In practice, when generating a new sequence, the first TIFF2026136085000062.tif9170 frames with intermediate noise tensor is the final result of the pre-generated sequence. TIFF2026136085000063.tif7 is replaced with a tensor containing intermediate noise from 170 frames. Also, the first of the newly generated sequences TIFF2026136085000064.tif7170 frames are the last of the pre-generated sequence TIFF2026136085000065.tif7 has the same semantic labeling condition applied to 170 frames. In total, Only 170 new frames are generated each time in TIFF2026136085000066.tif7. These are directly concatenated to the sequence obtained from previous generation.

[0052] This disclosure describes the evaluation of noise blending proposed in this disclosure. For these evaluations, the inventors trained a video diffusion model using a dataset for 2D cardiac echocardiography evaluation. The dataset includes 2D apical four-chamber and apical two-chamber views at end-diastolic and end-systolic, including hemicardial cycle sequences acquired from 500 patients. Each acquisition is accompanied by corresponding segmentation of the left ventricular myocardium, left ventricular blood pool, and left atrium. Since the right ventricle and right atrium are not segmented in the four-chamber view, it was decided to train the model only on the two-chamber view.

[0053] First, the inventors compared the noise blending method with the sequence-based conditioning method. In the latter, new sequences are conditioned directly on previously generated sequences rather than on each noisy intermediate frame. Figure 6 illustrates this comparison. The yt-sections of a composite ultrasound image of a specific patient generated by both methods are compared, where the yt-section is formed by extracting identical vertical lines from each frame of the sequence and stacking these lines horizontally to create a single image. The sequence-based conditioning method results in noticeable scene cuts between each chunk generation, while the noise blending method creates a seamless and smooth transition.

[0054] Next, the inventors of the present invention compared the use of the noise blending method with a pixel diffusion model and with a latent diffusion model.

[0055] To evaluate noise blending methods applying pixel diffusion and latent diffusion models, the inventors generated two datasets using both methods, each consisting of 200 long cardiac ultrasound sequences that collectively represent the entire cardiac cycle. The number of frames varied for each patient in the dataset, ranging from 40 to over 80 frames.

[0056] The purpose of this evaluation is to assess the quality of the invention in the following two ways: • Temporal consistency maintained even when the generation of long sequences is divided into multiple steps. • Fidelity to the semantic labels used to guide the generated sequence.

[0057] Generally, generative AI methods lack rigorous criteria for evaluating the quality of the generated sequences and images. Below, we discuss all criteria that could be considered useful.

[0058] (temporal consistency) Regarding the evaluation points at the beginning, the yt cross-section of the noise blend method shows better temporal consistency compared to the sequence-based conditioning method (see Figure 6). In addition to the yt cross-section, the Fréchet video distance (FVD) and Fréchet start distance (FID) criteria are employed. Both criteria measure the Fréchet distance between the real-world data distribution and the distribution defined by the generative model. FVD is designed for video datasets, and FID is adapted for image datasets. Specifically, both criteria use pre-trained networks, one for video and the other for images, to extract features from real-world and synthetic datasets, enabling meaningful comparisons. Here, FVD is used to compare sequences, and FID is applied to compare individual images, where individual images correspond to frames in the sequence. For both criteria, a lower value indicates better method performance.

[0059] [Table 1]

[0060] A low FID value in LVDM indicates that the image for each frame is closer to the original database distribution. Conversely, a low FVD value in PVDM suggests better temporal consistency.

[0061] (Label fidelity) To evaluate the fidelity of a synthetic sequence to its input labels, a segmentation neural network is trained on a real-world ultrasound sequence and applied to the synthetic sequence. The segmentation results obtained by this neural network are compared to the input labels. A higher score indicates higher fidelity. The network used for this task is nnU-Net, discussed in Isensee, et al. (2018). nnU-Net:Self-adapting framework for u-net-based medical image segmentation. arXiv preprint arXiv:1809.10486 (incorporated herein by reference). This neural network is recognized as one of the best segmentation networks in the field of medical imaging.

[0062] Table 2 uses four metrics to evaluate the quality of segmentation. TIFF2026136085000068.tif14170TIFF2026136085000069.tif14170TIFF2026136085000070.tif14170TIFF2026136085000071.tif14170

[0063] [Table 2]

[0064] The semantic labels consist of three classes: 0 represents the left ventricular blood pool, 1 represents the left ventricular myocardium, and 2 represents the atrial blood pool. Table 2 also includes the mean values ​​for all three classes.

[0065] The numerical results in Table 2 show that PVDM more accurately follows the original semantic labels during generation. In fact, almost all scores obtained from the PVDM dataset show higher values. Figure 7 shows the dice scores and IoU on a frame-by-frame basis. PVDM consistently achieves high scores not only in the average but also in all individual frames.

[0066] Based on the results above, PVDM has been shown to be superior in terms of temporal consistency and fidelity to semantic labels.

[0067] As an additional validation to evaluate the generalization ability of the pixel diffusion and latent diffusion models, although the models were trained only on two-chamber images, we attempted to generate images of an apical three-chamber image. The reason we evaluated how the pixel diffusion and latent diffusion models generalize to unseen images is that the semantic label sequence for the apical three-chamber image is unavailable. Figure 8 shows a semantic-labeled 2D representation of the apical three-chamber image. Figures 9 and 10 show two apical three-chamber images generated by the pixel diffusion and latent diffusion models, respectively.

[0068] Visually, the pixel diffusion model demonstrates the ability to generate previously unseen cardiac cross-sections, while the latent diffusion model struggles to generate the aorta. This may be due to the absence of the aorta during training. This indicates that the pixel diffusion model can generalize effectively, whereas the latent diffusion model cannot.

[0069] More precisely, two apical three-chamber diagram datasets were generated using both methods. Morphological snake segmentation, a type of unsupervised segmentation (discussed in Marquez-Neila et al. (2013). A morphological approach to curvature-based evolution of curves and surfaces. IEEE Transactions on Pattern Analysis and Machine Intelligence, 36(1), 2-17, and incorporated herein by reference), was used. The purpose of this was to determine whether the aorta was detected and to use this as an indicator of the generalization ability of the generative model.

[0070] [Table 3]

[0071] The measurements indicate that the pixel spatial diffusion model can be generalized to previously unseen organs (e.g., the aorta). The latent diffusion model achieved competitive scores only in the left ventricle and left atrium, as these are part of the training dataset. However, the performance indicators for the aorta were significantly lower with the latent diffusion model.

[0072] (Conclusion) This noise blending method demonstrated the ability to generate long cardiac ultrasound sequences guided by semantic labels. As mentioned above, noise blending makes it possible to generate cardiac ultrasound videos that simultaneously satisfy the following two conditions: • Long-form cardiac ultrasound video of any number of frames using a diffusion model. • Sequences induced by semantic labels.

[0073] This evaluation demonstrates the main advantage of the proposed method, namely its ability to generate long cardiac ultrasound videos guided to semantic labels at any number of frames. Generally, as mentioned earlier, video diffusion models are trained on a small number of frames due to computational limitations. By using noise blending, it becomes possible to generate geometrically guided ultrasound videos regardless of the length of each patient's cardiac cycle.

[0074] The advantages shown in the above evaluation can be summarized in the following list. • Long-form video generation: It is possible to generate ultrasound videos of any length while maintaining consistent quality throughout the entire process without any loss of image quality. • Adaptability to individual cardiac cycles: Video generation is customized based on each patient's unique cardiac cycle duration. • Overcoming computational constraints: Generating long video cardiac ultrasounds does not require models trained on a large number of frames. For example, training a model to generate 16 frames is sufficient to generate over 60 frames of video during inference. • Geometrically guided generation: Semantic labels are used to guide video generation, ensuring that the anatomical details and movements of the heart are realistic and clinically accurate.

[0075] The noise blending method proposed here was evaluated against cardiac images (i.e., the organ in this method is the heart) as a background. However, as mentioned above, this method is not limited to this specific organ. For example, this organ may be the heart, liver, kidney, carotid artery, aorta, or coronary artery. This method can be used in general to generate ultrasound sequences for any organ from the labels required for ultrasound image sequences in diagnosis, treatment planning, intervention, and follow-up. The evaluation test presented here was conducted in a cardiac use case (for example, see the reference incorporated herein by reference Ohte, N., Ishizu, T., Izumi, C., Itoh, H., Iwanaga, S., Okura, H., ...&Japanese Circulation Society Joint Working Group. (2022) JCS2021 guideline on the clinical application of echocardiography. Circulation Journal, 86(12), 2045-2119, for a discussion of this particular application case), but this method is used in the carotid artery (see Townsend, RR, Wilkinson, IB, Schiffrin, ELAvolio, AP, Chirinos, JA, Cockcroft, JR, ...&Weber, T. (2015). Recommendations for improving and standardizing vascular research on arterial stiffness: a scientific statement from the American Heart Association. Hypertension, 66(3), See pp. 698-722), the aorta (Writing Committee Members, Isselbacher, EM, Preventza, O., Hamilton Black III, J., Augoustides, JG, Beck, AW, ... & Woo, YJ (2022), incorporated herein by reference).See the 2022 ACC / AHA guideline for the diagnosis and management of aortic disease: a report of the American Heart Association / American College of Cardiology Joint Committee on Clinical Practice Guidelines. Journal of the American College of Cardiology, 80(24), e223-e393; the liver (see Kelly, EM, Feldstein, VA, Parks, M., Hudock, R., Etheridge, D., & Peters, MG (2018). An assessment of the clinical accuracy of ultrasound in diagnosing cirrhosis in the absence of portal hypertension. Gastroenterology & hepatology, 14(6), 367, incorporated herein by reference); the kidney (see Hansen, KL, Nielsen, MB, & Ewertsen, C. (2015). Ultrasonography of the kidney: a pictorial The method can also be applied to other organs / arteries where ultrasound sequences may be involved, such as the coronary arteries (see review.Diagnostics, 6(1), 2), Xia, M.; Yan, W.; Huang, Y.; Guo, Y.; Zhou, G.; Wang, Y. IVUS Image Segmentation Using Superpixel-Wise Fuzzy Clustering and Level Set Evolution. Appl.Sci.2019, 9, 4967. https: / / doi.org / 10.3390 / app9224967, incorporated herein by reference). By adapting the training database to the use case and considering acquisition frequency and probe, it is possible to generate data with high label accuracy using the method presented herein.

[0076] For example, in the carotid artery, motion extracted from the carotid wall provides unique information for evaluating the cardiovascular system. As described in the reference incorporated herein by reference, Rizi, FY, Au, J., Yli-Ollila, H., Golemati, S., Makunaite, M., Orkisz, M., ... & Zahnd, G. (2020). Carotid wall longitudinal motion in ultrasound imaging: an expert consensus review. Ultrasound in Medicine & Biology, 46(10), 2605-2624, longitudinal wall motion corresponds to the displacement of tissue layers in a direction parallel to blood flow during the cardiac cycle. To generate relevant data to address this issue, relevant labels include the carotid lumen, intima, media, adventitia, and atherosclerotic plaque.

[0077] Similarly, as shown in Figure 11, intravascular ultrasound (IVUS) sequences are acquired to characterize and measure coronary artery calcification. Structural segmentation, classification, and tracking are necessary to identify where and how the calcification moves relative to the artery, which can later be used to assess the risk of plaque rupture.

[0078] Liver research presents significant challenges due to the organ's complexity (structure, vascular distribution) and the fact that its movement is induced by respiration. This movement often deforms target structures such as cysts. The method proposed here allows for the generation of a synthetic ultrasound database that considers the respiratory cycle, and enables the induction of movement in simulated data to accommodate various physiological shapes. In this example, labels may represent different types of cysts, blood vessels, and liver tissue.

[0079] Furthermore, a database containing sequences of synthetic ultrasound images of organs generated according to this method is proposed. In other words, this database consists of data fragments, each fragment being a long sequence of synthetic ultrasound images of an organ obtained as a result of the iterative application of a video diffusion model. There may be one such database for each target organ (e.g., heart, liver, kidney, carotid artery, aorta, coronary artery). This database (or one for each organ) may be stored on a temporary or non-temporary computer-readable storage medium.

[0080] Furthermore, we propose a method for using a database run by a computer. This method may be performed independently of the noise blending method described above. Alternatively, both methods may be part of a process run by the same computer. This method involves training a neural network based on the database, which is trained for one of the following tasks: • Organ segmentation based on sequences of ultrasound images of organs, that is, training a neural network to take a sequence as input and output organ segmentation to characteristic parts (e.g., ventricles and atria in the case of the heart), or • Organ tracking based on sequences of ultrasound images of organs, that is, training a neural network to take a sequence as input and output a transformation (movement field) to geometrically align the organs along the sequence, or • Organ detection based on sequences of ultrasound images of organs, i.e., training a neural network to receive a sequence as input and output bounding boxes corresponding to the organ's location for each frame of the sequence, or • A registration task based on sequences of ultrasound images of organs, i.e., training a neural network to take sequences and corresponding sequences from other views as input and output transformations geometrically along the organ, or • The task involves classifying organs based on ultrasound sequences, specifically training a neural network to take a sequence as input and output a classification of the organ or its characteristics (e.g., benign or malignant).

[0081] Neural networks for these tasks are well known, along with their training capabilities. However, this disclosure provides new training data, namely long sequences generated from ultrasound images, to improve the quality of training.

[0082] Furthermore, we propose a neural network obtained according to this method of use, that is, a neural network having weights / parameters equal to those obtained as a result of training according to this method, for example, a neural network directly obtained by training according to this method, with weights and parameters directly set as a result of training according to this method. We also propose a computer-based method for using this neural network for inference to perform any of the above tasks on a sequence of ultrasound images of actually measured organs (measured by an appropriate medical measuring device). Therefore, the method of using the neural network includes the following steps. For example, the process of obtaining a long sequence of ultrasound images of an organ by actually performing the ultrasound imaging process or simply by obtaining a sequence that has already been acquired, and • Apply a neural network to the sequence obtained above, and the neural network will perform the task on which it was trained.

[0083] The use of this neural network may be carried out independently of other methods described herein, or within a process running on the same computer.

[0084] The methods of this disclosure are performed by a computer. This means that the steps (or substantially all steps) of the methods are performed by at least one computer or a similar system. Thus, the steps of the methods can be performed by a computer, or fully automatically or semi-automatically. As an example, the triggers for at least some of the steps of the methods of this disclosure may be performed through user-computer interaction. The required level of user-computer interaction may depend on the expected level of automation and may be balanced with the need to perform user requests. As an example, this level may be defined by the user and / or predefined.

[0085] A typical example of computer execution of the method is to implement the method using a system suited to this purpose. The system may have a processor connected to memory and a graphical user interface (GUI), and the memory contains a computer program containing instructions for implementing the method. The memory may also store a database. The memory may be any hardware suitable for such storage and may consist of several physically distinct parts (e.g., one for the program and possibly one for the database).

[0086] Figure 12 shows an example of such a system. This system is a client computer system, for example, a user's workstation.

[0087] The client computer in this example includes a central processing unit (CPU) 1010 connected to an internal communication bus 1000, and random access memory (RAM) 1070 connected to the same bus. Furthermore, the client computer includes a graphics processing unit (GPU) 1110 associated with video random access memory 1100 connected to the same bus. The video RAM 1100 is also known in this technology as a frame buffer. The mass storage controller 1020 manages access to mass storage devices such as a hard drive 1030. Mass storage devices suitable for concretely embodying computer program instructions and data include all forms of non-volatile memory, and exemplify semiconductor memory devices such as EPROMs, EEPROMs, and flash memory devices, magnetic disks such as internal hard disks and removable disks, and magneto-optical disks. Any of these may be complemented or incorporated by specially designed ASICs (Application-Specific Integrated Circuits). The network adapter 1050 manages access to the network 1060. The client computer may also include a cursor control device and tactile devices 1090 such as a keyboard. A cursor control device is used in a client computer to allow the user to selectively position the cursor at any location on the 1080-resolution display. Furthermore, the cursor control device allows the user to select various instructions and input control signals. The cursor control device comprises several signal generating devices for inputting control signals to the system. Typically, the cursor control device may be a mouse, with the mouse buttons used for signal generation. Alternatively or additionally, the client computer system may include a sensitive pad and / or a sensitive screen.

[0088] A computer program may comprise a set of instructions executable by the computer, which provide means for causing the system to implement the method. The program may be recordable on any data storage medium, including the system's memory. The program may be implemented, for example, as digital electronic circuits, computer hardware, firmware, software, or a combination thereof. The program may be implemented as a device, for example, as a product specifically embodied in a machine-readable storage device for execution by a programmable processor. The steps of the method may be implemented by a programmable processor that executes a program of instructions for performing the functions of the method by performing operations on input data and generating outputs. Thus, the processor may be programmable and connected to receive and transmit data and instructions from a data storage device, at least one input device, and at least one output device. The application program may be executed in a high-level procedural programming language or an object-oriented programming language, and, if necessary, in assembly language or machine language. In any case, the language may be a compiled language or an interpreted language. The program may be a complete installation program or an update program. In any case, the application of the program on the system provides instructions for implementing the method. The computer program may instead be stored and executed on a server in a cloud computing environment, and the server communicates with one or more clients over a network. In such a case, the processing unit executes the instructions contained in the program, thereby causing the method to be implemented on the cloud computing environment.

Claims

1. A method for generating a sequence of synthetic ultrasound images of organs, performed by a computer, A video diffusion model trained to take as input a sequence of multiple two-dimensional representations of the organ, each assigned multiple semantic labels, and a corresponding sequence of multiple noise images, and to generate a corresponding denoised sequence of multiple composite ultrasound images of the organ, respecting the semantic labels, The process of providing a plurality of consecutive sequences of a plurality of two-dimensional representations of the organ, each assigned a semantic label, and a plurality of corresponding consecutive sequences of noise images, The video diffusion model is iteratively applied to each pair consisting of one sequence of consecutive two-dimensional representations of the organ, each assigned a plurality of semantic labels, and a corresponding sequence of noise images; and in the iterative application of the video diffusion model, when applying the video diffusion model to any pair, a step is taken to use a portion of the denoised synthetic ultrasound image generated for the denoising step in the previous iteration in replacing a portion of the image in the arbitrary pair for each denoising step of the reverse processing of the video diffusion model. A method characterized by the following:

2. A portion of the denoised synthetic ultrasound image generated during the previous iteration is present in the denoised synthetic ultrasound image generated during the previous iteration, and a portion of the arbitrary pair of images is present in the leading image. The method according to claim 1, characterized by the features described above.

3. The total number of images in the aforementioned multiple consecutive sequences is less than 40, and the number of images in each sequence is 24 or less. The method according to 1 or 2, characterized by the above.

4. Each sequence contains 16 images. The method according to feature 3.

5. The number of images in a portion of the denoised synthetic ultrasound image generated in the previous iteration is the number in the tail portion of the denoised synthetic ultrasound image generated in the previous iteration, the number of images in a portion of any pair of images is the number in the leading portion of the image, and the number is between one-quarter and one-half of the total number of images in each sequence. The method according to feature 3 or 4.

6. The aforementioned organs are the heart, liver, kidneys, carotid artery, aorta, or coronary artery. The method according to any one of 1 to 5, characterized by the features described herein.

7. The method comprises a plurality of sequences of synthetic ultrasound images of organs generated by any one of claims 1 to 6. A database characterized by the following features.

8. A method for using the database of claim 7, which is performed by a computer, The process includes training a neural network based on the aforementioned database, The aforementioned neural network is Segmentation of the organ based on a single sequence of multiple ultrasound images of the organ, Tracking of the organ based on a single sequence of multiple ultrasound images of the organ, Detection of the organ based on a sequence of multiple ultrasound images of the organ, A registration task based on a single sequence of multiple ultrasound images of the aforementioned organ, or Trained for classification tasks based on a single sequence of multiple ultrasound images of the aforementioned organs. A method for using a database characterized by the following features.

9. A neural network obtainable by the method of claim 8.

10. When executed by a computer system, the system includes an instruction to cause the computer system to carry out the method according to any one of claims 1 to 6 and / or the method according to claim 8. A computer program characterized by the following features.

11. The computer program described in claim 10 and / or the database described in claim 7 and / or the neural network described in claim 9 A computer-readable data storage medium that records data.

12. The system comprises a processor connected to a memory that stores the computer program described in claim 10 and / or the database described in claim 7 and / or the neural network described in claim 9. A computer system characterized by the following features.