Computer system, method of processing ultrasound imaging data and computer program

The computer system enhances ultrasound image transformation by generating classification information for ultrasound data features and using a machine learning model to condition the process, resulting in reliable and accurate rendered images.

JP2026015232APending Publication Date: 2026-01-29CANON MEDICAL SYST CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025111209
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-07-17
Filing Date
2025-07-01
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Existing ultrasound image transformation techniques are prone to failures, resulting in abnormal images.

Method used

A computer system that generates classification information for ultrasound imaging data features and uses an image transformation machine learning model to produce reliable rendered images by conditioning the transformation process with this information.

Benefits of technology

Enables highly reliable image conversion of ultrasound imaging data by reducing the generation of abnormal images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026015232000001_ABST
    Figure 2026015232000001_ABST
Patent Text Reader

Abstract

To enable image conversion of ultrasonic imaging data with high reliability.SOLUTION: Generating classification information for each of a plurality of features in the two dimensional ultrasound image from input data including at least one of the two dimensional ultrasound image and three dimensional ultrasound data corresponding to the two dimensional ultrasound image, and generating a rendering image by providing the two dimensional ultrasound image as an input image and the classification information for each of the plurality of features in the input image to an image transformation machine learning model.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD OF THE INVENTION The embodiments disclosed herein relate generally to ultrasound image processing methods, and more particularly to a computer system, a method for processing ultrasound imaging data, and a computer program product for processing ultrasound images to generate a rendered image. [Background technology]

[0002] Ultrasound images are formed by transmitting high-frequency sound pulses from an ultrasound probe into tissue. These pulses bounce off tissues in the patient's body with different reflection characteristics, reflecting back toward the probe and being detected. Ultrasound scanners construct images based on measurements of these reflected pulses. Three-dimensional (3D) ultrasound data can be acquired using a probe specifically designed to collect 3D data. Alternatively, 3D ultrasound data can be acquired by moving the ultrasound probe and acquiring multiple ultrasound images. For example, tilting the ultrasound probe captures reflected pulses at different probe orientations. Processing the reflected pulses captured at different probe orientations generates a 3D array of multiple voxels representing the imaged structure. After capturing the 3D data, two-dimensional (2D) images are generated from selected angles by applying volume rendering techniques to the 3D data. These 2D images are useful for visualizing the raw information captured by the ultrasound scanner.

[0003] Recent advances have been made in image generation models, such as machine learning models. Machine learning models are trained to generate new images based on specific inputs. Some generative models are specifically trained to generate specific types of images and receive only a random seed (e.g., a vector) as input, from which the output image is generated. Other models can operate in a mode (called txt2img) that uses textual information as a prompt in addition to the seed. The prompt causes the model to generate an image that reflects the text prompt. Some models can also operate in a mode (called img2img) that is given an input image along with a seed and generates an output image that reflects a transformation of the input image. In this case, the model can be referred to as an image transformation model. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] US Patent Application Publication No. 2016 / 0242740 Summary of the Invention [Problem to be solved by the invention]

[0005] One of the problems to be solved by the embodiments disclosed in this specification and the drawings is to enable highly reliable image conversion of ultrasound imaging data. However, the problems to be solved by the embodiments disclosed in this specification and the drawings are not limited to the above problem. Problems corresponding to the effects of each configuration shown in the embodiments described below can also be positioned as other problems. [Means for solving the problem]

[0006] A computer system according to an embodiment, comprising at least one processor and at least one memory including computer-readable instructions, wherein the at least one processor executes the computer-readable instructions to cause the computer system to acquire a two-dimensional ultrasound image, generate classification information for each of a plurality of features in the two-dimensional ultrasound image from input data including at least one of the two-dimensional ultrasound image and three-dimensional ultrasound data corresponding to the two-dimensional ultrasound image, and generate a rendering image by providing the two-dimensional ultrasound image as an input image and the classification information for each of a plurality of features in the input image to an image conversion machine learning model. [Brief explanation of the drawings]

[0007] Several embodiments of the present disclosure will now be described, by way of example only, with reference to the accompanying drawings, in which: [Figure 1A] FIG. 1A is a diagram illustrating an example of a computing device according to an embodiment of the application. [Figure 1B] FIG. 1B is a diagram illustrating an example of a calculation server according to an embodiment of the application. [Figure 1C] FIG. 1C illustrates an example of a computing device communicating with a computing server. [Figure 2] FIG. 2 is a diagram illustrating an example of a system including an ultrasound system that generates ultrasound imaging data and an image processing system that generates a transformed image of the ultrasound imaging data. [Figure 3] FIG. 3 shows an example of a transformation of a fetal ultrasound image along with conditioning information. [Figure 4] FIG. 4 is a diagram illustrating an example of a neural network. [Figure 5A] FIG. 5A illustrates an example of a first portion of a convolutional neural network. [Figure 5B] FIG. 5B illustrates an example of a second portion of a convolutional neural network. [Figure 6A]FIG. 6A illustrates a generative adversarial network (GAN) where conditioning information is input into the network. [Figure 6B] FIG. 6B is a diagram illustrating a simplified example of a generative adversarial network (GAN) network. [Figure 7] FIG. 7 illustrates a diffusion network for performing image transformation. [Figure 8] FIG. 8 illustrates a U-net that denoises images as part of image processing using a diffusion network. [Figure 9A] FIG. 9A illustrates another network that modifies the output of a diffusion network based on conditioning information. [Figure 9B] FIG. 9B shows in more detail another network that modifies the output of the diffusion network based on conditioning information. [Figure 10] FIG. 10 is a diagram illustrating an example of a training method according to the embodiment. [Figure 11] FIG. 11 is a diagram illustrating an example of a method according to an embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0008] It has been proposed to improve the quality of ultrasound images by applying image transformation techniques, however, such image transformation techniques are prone to failure, resulting in abnormal images.

[0009] A computer-implemented method for processing ultrasound imaging data according to a first aspect of the embodiment includes acquiring a two-dimensional ultrasound image, generating classification information for each of a plurality of features in the two-dimensional ultrasound image from input data including at least one of the two-dimensional ultrasound image and three-dimensional ultrasound data corresponding to the two-dimensional ultrasound image, and generating a rendered image by providing the two-dimensional ultrasound image as an input image and the classification information for each of the plurality of features to an image transformation machine learning model. The rendered image may also be referred to as a further image. The rendered image output from the image transformation machine learning model is a realistic rendered image that is clearly different from the two-dimensional ultrasound image (or the rendered image output from a rendering system).

[0010] The inventors have discovered that reliable image transformation of ultrasound imaging data is possible by first generating classification information for different features in the input data. The classification information, for example, represents the classification of different body parts being imaged and can be used to condition the image transformation process to reduce the generation of abnormal images.

[0011] A computer system according to a second aspect includes at least one processor and at least one memory containing computer-readable instructions, the at least one processor executing the computer-readable instructions causing the system to perform the method of the first aspect.

[0012] A computer program according to a third aspect comprises computer readable instructions which, when executed by at least one processor of a computer system, cause the system to perform the method of the first aspect.

[0013] A non-transitory computer-readable medium according to a fourth aspect stores the computer program according to the third aspect.

[0014] The embodiments will now be described in more detail with reference to the accompanying drawings.

[0015] Please refer to Figure 1. Figure 1 illustrates an example of a data processing system 100 in which embodiments may be implemented. Data processing system 100 may be, for example, a server, a terminal or workstation, a personal computer (PC), or other device.

[0016] Data processing system 100 includes an interface 140 through which it transmits and receives signals. Interface 140 may be a wired or wireless interface. For example, interface 140 may include a wired interface for connecting to a wired network (e.g., a local area network and / or the Internet). Alternatively or additionally, interface 140 may include a transceiver device for communicating over a wireless interface. The transceiver device may be implemented, for example, using a radio unit and an associated antenna array. The antenna array may be located inside or outside data processing system 100.

[0017] The data processing system 100 includes at least one data processing unit 115, at least one random access memory 120, at least one hard disk drive 125, and possible other components 130 used to perform software / hardware-assisted system tasks. Tasks performed by the system include controlling, accessing, and communicating with access systems and other communication devices. The at least one random access memory 120 and the hard disk drive 125 communicate with the data processing unit 115 as a data processor. The data processing unit, storage, and other related controls are optionally implemented on a circuit board and / or chipset. A user controls the operation of the data processing system 100 using an interface, such as a keypad 110, or by voice commands. A display 105 may be provided on the data processing system 100 to display visual content to the user. The data processing system 100 may also include a speaker to provide audio content.

[0018] The memory of data processing system 100 (i.e., random access memory 120 and hard disk drive 125) stores computer-readable instructions for data processing unit 115 to perform data processing functions, as referred to herein as performed by data processing system 100. Alternatively, components 130 may include hardware components, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC), for performing the operations performed by data processing system 100 herein. In embodiments, the operations performed by data processing system 100 herein may be performed by a combination of hardware components or a processor executing computer-readable instructions.

[0019] 1A depicts data processing system 100 as a single, integrated device, in other embodiments, data processing system 100 may comprise multiple interconnected devices. Operations performed by data processing system 100 refer to operations performed by processing circuitry (e.g., data processing unit 115, components 130) of data processing system 100 that executes the operations. In particular, operations performed by data processing system 100 refer to operations performed by processing circuitry that executes computer-readable instructions stored in random access memory 120 and hard disk drive 125 of data processing system 100.

[0020] Please refer to FIG. 1B. FIG. 1B illustrates an example of a computer system 150 that can be used to perform the processes described herein. The computer system 150 can be used to train a machine learning model. Additionally or alternatively, the computer system 150 can be used to operate the machine learning model. While the computer system 150 is illustrated as a single, enclosed device, in another embodiment, the computer system 150 is a distributed system that includes multiple data processing devices operating in communication with each other. The computer system 150 can include a server, a back-end system, etc.

[0021] The computer system 150 includes at least one random access memory 160, at least one hard disk drive 170, at least one data processing unit 180, 190, and an input / output interface 195. The random access memory 160 and the hard disk drive 170 store data, including data input to one or more models and results of processing performed during operation of the one or more models. The random access memory 160 and the hard disk drive 170 may store training data applied to train the machine learning models. The random access memory 160 and the hard disk drive 170 further store computer-executable code that is executed by the at least one data processing unit 180, 190 to implement one or more machine learning models. At least one of the data processing units 180, 190 performs one or more of the following operations: processing related to one or more models, training the models, and necessary preprocessing of data used by the models. The computer system 150 receives data items that constitute a training dataset and / or data items that constitute an operational dataset via the input / output interface 195. The computer system 150 also transmits, via the input / output interface 195, the results produced by running the model on the input data.

[0022] 1C, which illustrates a system 161 that includes a data processing system 100 in communication with a computer system 150. In this example, the computer system 150 stores and runs one or more machine learning models. For example, the computer system 150 may be a cloud-based server or a graphics processing unit (GPU), which processes data from the data processing system 100 by inputting the data into one or more machine learning models and returning the processed results to the data processing system 100.

[0023] Please refer to FIG. 2. FIG. 2 provides an overview of the process, showing the different modules included in system 200 and the data items generated by the modules that contribute to generating an output image. System 200 includes ultrasound system 205, which acquires raw ultrasound data, generates 3D data from the raw ultrasound data, and renders the 3D data to generate a 2D image. System 200 also includes image processing system 210, which receives the 2D images from ultrasound system 205 and generates an output image from the 2D images. Image processing system 210 may correspond to any of data processing system 100, computer system 150, and system 161 described above. In this case, the corresponding data processing system 100, computer system 150, and system 161 only processes the output of ultrasound system 205. Alternatively, system 200 as a whole may correspond to any of data processing system 100, computer system 150, and system 161. In this case, the corresponding data processing system 100, computer system 150, and system 161 also performs rendering of the 3D data. Operations performed by modules or components of system 200 (e.g., an operational machine learning model) refer to operations performed by at least one processor of data processing system 100, computer system 150, or system 161 executing computer-readable instructions of a computer program that is maintained for performing the operations in at least one memory of data processing system 100, computer system 150, or system 161. Training of a machine learning model that processes as part of the system may be performed by at least one processor of data processing system 100, computer system 150, or system 161 executing computer-readable instructions, or by a separate system (not shown) that provides a trained machine learning model to the corresponding data processing system 100, computer system 150, or system 161.

[0024] As shown in Figure 2, ultrasound system 205 is part of system 200. Ultrasound system 205 includes a probe 215. Probe 215 can be tilted to capture ultrasound data at different orientations or can directly acquire volumetric data. Probe 215 may also be referred to as a transducer. Ultrasound system 205 includes processing circuitry including a raw data processing module 220 and a rendering system 225.

[0025] The probe 215 acquires ultrasound raw data and provides 3D data of the imaging subject. The probe 215 may be used, for example, to image a fetus in the womb. The probe 215 outputs the ultrasound raw data to the raw data processing module 220, which processes the ultrasound raw data to create volumetric data. The volumetric data includes an array of voxels, with values ​​corresponding to each voxel in the array indicating the presence of different types of tissue or material. The 3D data is provided to a rendering system 225.

[0026] The rendering system 225 generates a 2D image from the 3D data. The generated 2D image is a 2D projection of the 3D data. In addition to the 3D data, the rendering system 225 receives as input certain parameters that enable the creation of the 2D image. The parameters include the position and orientation of a camera relative to the volume represented by the 3D data. The 2D image is created from the viewpoint of this camera. The parameters may also include lighting information for illuminating the 2D image.

[0027] The rendering system 225 can perform volume rendering to generate a 2D image through a number of different methods. The rendering system 225 can perform direct volume rendering by calculating the intensity of different points in the 2D image. As an example, direct volume rendering uses an integral to calculate the intensity I of point x on the viewport received along a direction ω. Specifically, the intensity is given by Equation 1 below:

[0028]

number

[0029] T is a transfer function that integrates the local attenuation along the path between the viewpoint and position s.

[0030]

number

[0031] E represents the emission at a point along the path, and S represents the scattering and reflection contributions.

[0032] In embodiments, the rendering system 225 may add more realistic lighting to a 2D image by applying global illumination to the image, where the rendering system 225 determines the luminance at various points in the image without using some of the global approximation used in traditional direct volume rendering (DVR). When applying global illumination, the scattering function S is evaluated as a function of the recursive integral over the irradiance and the fraction of light reflected in a given direction.

[0033] The output of the rendering system 225 includes a 2D image from a given viewpoint, i.e., camera position and orientation. The rendering system 225 also outputs depth information determined based on the 3D data. The depth information is in the form of a depth map that indicates the depth of surfaces projected within the 2D image. The depth information indicates the position of surfaces within the 2D image. The depth information can also include surface normals to indicate the orientation of the surfaces.

[0034] The ultrasound system 205 provides 2D images generated by the rendering system 225 to the image processing system 210. The ultrasound system 205 also provides 3D data and / or depth information to the image processing system 210.

[0035] The classifier module 230 is implemented by processing circuitry in the image processing system 210. The classifier module 230 determines and outputs classification information based on the 2D images or 3D data output by the rendering system 225. The classifier module 230 includes one or more convolutional neural networks (CNNs) that receive input ultrasound image data (2D images or 3D data) and classify features in the image data. When the classifier module 230 processes 3D data, the classifier module 230 outputs 3D classification information. The classifier module 230 performs additional rendering on this classification information to create a 2D classification map. The 2D classification map corresponds to the 2D image output by the rendering system 225.

[0036] The classification information may include a classification map suitable for overlay onto the 2D image, representing classifications of different parts of the 2D image, such as nose, eyes, mouth, ears, etc. Alternatively, the classification information may be a segmentation map. Alternatively, the classification information may be pose information. If the 2D image and 3D data are image data of a fetus, the classifier module 230 may identify parts in the image data that correspond to body parts of the fetus, such as nose, eyes, mouth, ears, etc.

[0037] When the classification information is pose information, the classifier module 230 identifies different parts of the 2D image, e.g., nose, eyes, mouth, ears, etc., and estimates the positions and orientations of these parts. The pose information associated with each part includes a matrix (e.g., a transformation matrix) that represents the position and orientation of the object. Each object for which the matrix is ​​defined is selected from a predefined set of objects that represent different parts of the image, such as the nose, mouth, etc. That is, the pose information includes an identifier of the type of object in addition to the position and orientation information associated with the object.

[0038] The image processing system 210 also processes an image transformation model 235 that receives the 2D image and the classification information and generates an output image based on such information. The image transformation model 235 may also receive depth information output from the rendering system 225 and generate an output image based on the depth information. The image transformation model 235 may also receive text information input to the image processing system 210 by a user and generate an output image based on the text information. The image transformation model 235 includes one or more machine learning models that generate an image based on an initial input image (in this case, a 2D image obtained from ultrasound data) and predetermined conditioning information such as classification information, text prompts, and depth information. Examples of machine learning models used to determine the output image are described in more detail below.

[0039] While the example in which the raw data processing module 220 outputs a 3D data set based on the ultrasound raw data has been described, in some embodiments, the 3D data may include, for example, a single 3D data set representing the state of the imaging object at a particular time. In another embodiment, the 3D data may include a time-series 3D data set representing the state of the imaging object over time. In the case of data representing the state of the imaging object at a certain time, the rendering system 225 generates a single 2D image based on this state, and the image transformation model 235 generates a single output image based on this state. However, in the case of time-series 3D data, the rendering system 225 generates multiple 2D images, each corresponding to a different time. The classifier module 230 generates a single classification corresponding to each of the multiple 2D images. The multiple 2D images represent a video of the imaging object. The image transformation model 235 generates a transformed video based on the multiple images and the multiple classifications generated by the rendering system 225.

[0040] In another embodiment, the rendering system 225 performs rendering on a single received 3D dataset to generate multiple 2D images, each corresponding to a different viewpoint / camera orientation. The image processing system 210 inputs each rendered image into the image transformation model 235 to obtain a neural radiance field (NeRF) object.

[0041] In embodiments, upon obtaining the output image produced by image transformation model 235, image processing system 210 further processes the obtained output image, for example, by adding further detail to the image through context-aware infill processing, which is performed in accordance with additional text prompts provided by the user.

[0042] Please refer to FIG. 3. FIG. 3 illustrates an example of another data involved in the process shown in FIG. 2. In this case, the ultrasound image data is image data of a fetus. FIG. 3 illustrates an example of an input image 310 output from the rendering system 225. FIG. 3 also illustrates a text prompt that may be optionally input to the image processing system 210 by user input. The text prompt describes a predetermined characteristic of the fetus that is input by a human user. For example, the text prompt may indicate whether the fetus's eyes are open or closed. The text prompt may also indicate the ethnicity of the fetus.

[0043] Figure 3 shows an example of a classified image 320 (corresponding to the input image 310) output from the classifier module 230. Figure 3 also shows an example of a depth map 330 output from the rendering system 225. The image transformation model 235 takes as input the input image 310 and the classified image 320, and optionally the depth map 330 and a text prompt. The image transformation model 235 outputs a photo-realistic, alternatively rendered image 340. The alternatively rendered image 340 is referred to as the rendered image or the photo-realistically rendered image.

[0044] Once the image transformation model 235 outputs an output image, the image processing system 210 performs a validation check (in the validation module 240) on the image to determine whether the image meets predetermined requirements. The validation module 240 includes a machine learning model that is trained to determine whether the output image meets the validation criteria. The machine learning model is trained based on a set of images classified by a user. Each image in the user-classified image set is labeled as either a good image (i.e., an image that passes the validation check) or a bad image (i.e., an image that does not pass the validation check). The machine learning model is a convolutional neural network (CNN) that outputs a value indicative of an image quality score in response to receiving the output image generated by the image transformation model 235. The image processing system 210 compares the image quality score with a threshold to determine whether the output image matches the image quality score.

[0045] If the output image does not pass the validation check, the image transformation model 235 repeats the process to generate the output image. When regenerating the output image, the image transformation model 235 uses a different seed. The different seed includes a different set of noise that is added to the 2D input image as part of the transformation process performed by the image transformation model 235. Additionally or alternatively, when regenerating the output image, the image transformation model 235 uses a different text prompt provided by the user. Additionally or alternatively, when regenerating the output image, the image transformation model 235 may use another image generated by the rendering system 225 based on the same 3D data. This other image is generated by the rendering system 225 with different lighting or a different viewing (i.e., camera) direction.

[0046] When a new output image is obtained, the validation module 240 performs another check to determine whether the new output image meets the validation requirements. If the output image passes the validation check, the image processing system 210 controls the display 105 to display the output image.

[0047] As described below, the processing performed by the image processing system 210 involves multiple machine learning models. The classifier module 230 includes a CNN for generating classification information for multiple features in an input image. The image transformation model 235 may include one or more generative machine learning models for image transformation. For example, the image transformation model 235 includes a generative adversarial network (GAN) or a diffusion model. The validation module 240 includes a CNN for determining a quality score for the image output from the image transformation model 235. The operation of these various models is described in more detail below.

[0048] FIG. 4 is a schematic diagram of a neural network 400. The neural network 400 includes input nodes 410, hidden nodes 420, and output nodes 430. In practice, the neural network 400 will often have more nodes than shown, with two or more hidden layers. Each input node 410 receives a single value of input data and generates an activation value, or node value, at its output. The activation value, or node value, is generated by applying the input value to an activation function (e.g., a sigmoid). Each input node 410 is connected to each of the hidden nodes 420. A weight matrix defines the connectivity between the input nodes 410 and the hidden nodes 420. The vector of node values ​​output by the input node 410 is scaled by a vector of weights at each input of the hidden nodes 420. Each weight defines the connectivity between one of the input nodes 410 and its associated hidden node 420. In FIG. 4, the weights attached to the input of one of the hidden nodes 420 are denoted w0...w3. The input values ​​of each hidden node 420 are given by the dot product of the corresponding weight vector and the output value of input node 410. The output values ​​of these hidden nodes 420 are determined by applying an activation function to the input values ​​of the hidden nodes 420. The output vectors of the hidden nodes 420 are provided to each output node 430 in the next layer of neural network 400 and are similarly used to generate the output values ​​of the next layer.

[0049] Neural network 400 may be trained using supervised or unsupervised learning. In one embodiment, neural network 400 is trained through supervised learning by determining at least one set of output values ​​based on at least one set of input values ​​included in training data. The output values ​​are compared to known labels in the training data, and an error or loss is calculated (i.e., based on the difference between the output values ​​and the labels). The error or loss is backpropagated through neural network 400 to update the weights so that neural network 400 can learn to better approximate the labels from the input values. In the next cycle, the weights are further updated based on the updated weights and other training data, and a more approximate label is reproduced based on the input values ​​of other training data. In this way, neural network 400 can learn to perform a specific task.

[0050] 5A and 5B, which illustrate an example of the operation of a convolutional neural network that can be used to identify and classify predetermined features within an image. In the illustrated example, the input image 310 is a 2D rendering of a fetus. The convolutional neural network generates a set of output values ​​that indicate the classification of different portions of the input image 310.

[0051] A kernel 510 is applied to specify the convolution of the input image 310 with the kernel 510. An activation function adds nonlinearity to the output of this convolution. The activation function shown in FIG. 5A is a rectified linear unit (RELU) function. That is, if the input is positive, it outputs the input, and if not, it outputs zero. Multiple feature maps are generated from the input image by convolving the input image with different kernels. Each kernel represents a different basic feature, such as vertical or horizontal lines.

[0052] Next, each feature map created by the convolution-activation function undergoes a pooling process to reduce the spatial size of the convolved features. The pooling process involves moving the kernel 510 to the other side of the feature map, sampling pixels, and returning the maximum or average value from each sampled pixel in the feature map. Another convolution process (applying a RELU function) is performed on each pooled feature map using a different kernel to generate another set of feature maps on which pooling is performed again.

[0053] As shown in Figure 5B, the reduced feature map resulting from multiple stages of convolution and pooling is flattened into a one-dimensional array (denoted as a flattening layer). The one-dimensional array becomes a set of input values ​​for a feed-forward neural network. The resulting output values ​​represent classifications of different parts of the input image 310. The classifier module 230 converts the output values ​​into a classification map of the input image 310 or into pose information for the input image 310.

[0054] The convolutional neural network is trained by comparing the output values ​​for different images with the labels of those images and adjusting the weights of the feedforward part of the convolutional neural network.

[0055] In an embodiment, the image translation model 235 includes a generator model of a generative adversarial network (GAN). See Figures 6A and 6B. Figures 6A and 6B show an example of a GAN 600.

[0056] FIG. 6A illustrates two components of the GAN 600. The first component, referred to as the generator 610, generates an image based on a specific input, denoted x. The input may be a random vector or data representing an image. The generator 610 also receives conditional information, denoted c. The conditional information is provided to the generator 610 as an additional input layer. The generator 610 creates an output image G(x|c). In an embodiment, the input x is one of the images created by the rendering system 225, and the conditional information includes the classification information of the image determined by the classifier module 230. The classification information may be in the form of an input set representing a classification map, such as data indicating the classification of pixels in a 2D image. Alternatively, the data may be in the form of pose information, including, for each classified object in the 2D image, an indication of the object type, and the object's position and orientation. After training, the generator 610 outputs a higher-quality image (e.g., the rendered image 340), as described above.

[0057] GAN 600 further includes a second component called discriminator 620. Discriminator 620 is used as part of the training process for generator 610. Discriminator 620 is trained to assign a score to images output from generator 610. The score indicates the degree of match between the generated image and the training image set. In other words, discriminator 620 is trained to distinguish whether an image is real or generated. Discriminator 620 receives as input data G(x|c) representing the input image output from generator 610. Discriminator 620 also receives the same condition information c as received by generator 610 at an additional input layer.

[0058] The generator 610 and the discriminator 620 are trained as part of the same training process, in which the loss function of the discriminator 620 is used to update the model parameters of both the generator 610 and the discriminator 620. This training process is executed by the computer system 150. The generator 610 generated by the training process is capable of performing image transformation of the 2D image output from the rendering system 225 as the image transformation model 235 based on data representing classification information from the classifier module 230 as condition information. The generator 610 may also receive data indicating depth information and / or data indicating a text prompt as condition information.

[0059] 6B shows a simplified example of a generator 610 and a discriminator 620, along with example layers. As shown, the generator 610 receives an input x representing an input image. The generator 610 also receives conditional information c containing classification information for features in the input image x. Both x and c are mapped to hidden layers of the generator 610's neural network by a predetermined activation function. In response to these inputs, the generator 610 generates an output G(x|c) representing an output image.

[0060] The discriminator 620 receives an input G(x|c) representing an input image. The discriminator 620 also receives condition information c containing classification information of features in the input image x. Both x and c are mapped to a hidden layer of the neural network of the discriminator 620 by a predetermined activation function. In response to these inputs, the discriminator 620 generates an output D(G|c) representing the score of the input image G(x|c).

[0061] In an embodiment, the image transformation model 235 includes a diffusion model that performs the transformation of the image.

[0062] Please refer to FIG. 7. FIG. 7 shows how the diffusion model converts an input image (in this example, input image 310) into an output image (in this example, rendered image 340). The diffusion model 700 includes an encoder 710. The encoder 710 encodes the input image 310, converting it from pixel space to a latent space image. The output image 720 from the encoder 710 is the input image 310 in the latent space. The diffusion model 700 also includes a module 730. The module 730 adds noise to the input image 310 to generate a noisy image 740. The noisy image 740 retains some of the information of the original input image 310. The noisy image 740 is input to a denoiser module 750 of the diffusion model 700. The denoiser module 750 includes a CNN that is trained to iteratively denoise the noisy image 740. The image denoising is repeated multiple times until an output image 760 is generated. The output image 760 is provided to a decoder 770 which converts the output image 760 into pixel space and produces a rendered image 340 suitable for display.

[0063] Additionally, conditioning information may be applied to the denoising process to control the conversion process. The conditioning information is provided to an encoder 780, which converts the conditioning information into a set of values ​​suitable for application to the denoiser module 750 via a cross-attention mechanism. For example, the conditioning information may include a text prompt, which is converted by the encoder 780 into a set of numeric values ​​that are then provided to the denoiser module 750 via a cross-attention mechanism.

[0064] The diffusion model 700 learns to generate output images in an adversarial manner by training the discriminator 620 to assign quality scores to the output images of the diffusion model 700. The quality scores indicate the degree to which the output image corresponds to the desired output.

[0065] The denoiser module 750 includes a U-net model. Reference is now made to Figure 8, which shows an example of the U-net model included in the denoiser module 750. The U-net model included in the denoiser module 750 includes a first set of encoding stages consisting of convolution stages 810, 820, 830, and 840 and pooling stages 815, 825, and 835.

[0066] The denoiser module 750 contains a U-net model that receives an input image 740 and applies it to a downsampling convolution stage 810. The downsampling convolution stage 810 convolves the input image 740 to generate a set of feature maps using an activation function. A portion of each feature is preserved and concatenated with other feature maps later in the process. The feature maps resulting from the convolution are then pooled to generate a reduced feature map. The convolution-pooling process is repeated to further downsample the data through each stage 820, 825, 830, 835, and 840.

[0067] In the second part of the U-net model, which is included in the denoiser module 750, the pooling stages 815, 825, and 835 are replaced by upsampling stages 845, 855, and 865. Interposed between the upsampling stages 845, 855, and 865 are convolution stages 850, 860, and 870, respectively, in which convolutions are performed. Each convolution stage 850, 860, and 870 operates on the result of concatenating the output of the previous convolution stage 810, 820, and 830 with the output of the preceding upsampling stage 845, 855, and 865.

[0068] The result of the processing of the U-net model included in the denoiser module 750 is an output image 875 that represents the input image 740 with at least some noise removed. To denoise the input image 740, the U-net is applied multiple times.

[0069] During the training process, the U-net model contained in the denoiser module 750 may be trained to appropriately denoise an image to generate a transformed image by adjusting the filters used during the convolution operation.

[0070] In the above embodiment, conditioning information is applied to the U-net model contained in the denoiser module 750 to adjust the filters of the denoiser module 750. This can be applied for certain conditioning information, such as text, by providing a cross-attention mechanism between the coded conditioning information and the denoiser module 750. For spatial conditioning information, such as classification information, other techniques for applying conditioning information can be employed.

[0071] 9A , which illustrates a diffusion network block 900 that is part of another example image translation model 235. The image translation model 235 includes a denoiser module 750 and an additional section 910 for processing classification information. The diffusion network block 900 is an encoder block with one convolutional pooling stage of the denoiser module 750. The image translation model 235 also includes a copy 920 of the diffusion network block 900. The diffusion network block 900 is trained, and its model parameters are fixed before training the copy 920 of the diffusion network block 900. The additional section 910 also includes convolutional layers 930 and 940.

[0072] The classification information is provided as a 2D image. The convolutional layer 930 performs a 1x1 convolution on the classification information and combines the result with the input x to the diffusion network block 900. The output of the copy 920 of the diffusion network block 900 is input to another convolutional layer 940, which also applies the 1x1 input. The result of this other convolutional layer 940 is combined with the output of another diffusion network block 950 of the denoiser module 750 to produce the output y. The other diffusion network block 950 is a decoder block that includes one convolution and corresponding upsampling stage from the denoiser module 750.

[0073] 9B, which illustrates the denoiser module 750 and a network 960 for processing classification information. The network 960 includes encoder blocks 920a-920c corresponding to the encoder blocks of the denoiser module 750. However, the encoder blocks 920a-920c are trained using training data containing classification information and have different parameters than the encoder blocks of the denoiser module 750. Each encoder block includes a convolution stage and a corresponding downsampling stage. The network 960 includes convolution layers 940a-940c. The convolution layers 940a-940c perform a 1×1 convolution on each input received from the previous encoder block of the network 960. The output of each of the convolution layers 940a-940c is combined with the output of the corresponding decoder block of the denoiser module 750.

[0074] The additional network 960 can be trained with a relatively smaller amount of training data than is required to train the denoiser module 750. The training data for training the additional network 960 includes a set of images generated from the ultrasound data (by the rendering system 225) and corresponding classification information (obtained by the classifier module 230).

[0075] Please refer to Figure 10. Figure 10 shows a process 1000 for training the additional network 960. The process is performed by the computer system 150 or the system 161. In S1010, images obtained from the ultrasound data are input to the trained denoiser module 750. Meanwhile, classification information (e.g., the segmentation map and pose information described above) is input to the additional network 960 for conditioning. This results in a set of output images in S1020.

[0076] At S1030, a quality index is assigned to each image in the set of output images. The quality index may be assigned manually by a user or by providing the image as an input to a discriminator network. The quality index may be a binary index or a score on a quality scale. At S1040, the model parameters of the additional network 960 are updated based on the image quality and used to train the network 960 to generate higher quality images. Updating the network 960 may include updating the convolution filters and / or updating the weights and biases. The training process 1000 proceeds to S1010 for another training iteration. During this process, the parameters of the denoiser module 750 are not updated.

[0077] Please refer to Figure 11. Figure 11 illustrates a computer-implemented method 1100 according to an embodiment. Method 1100 is implemented in a computer system (e.g., data processing system 100, computer system 150, system 161) by at least one processor executing computer-readable instructions contained in a computer program.

[0078] In S1110, the system acquires a 2D ultrasound image.

[0079] In S1120, the system generates classification information from the input data for each of a plurality of features of the 2D ultrasound image.

[0080] In S1130, the system generates a rendered image by providing the 2D ultrasound image as input to an image transformation machine learning model, and as part of generating the rendered image, the system also provides the classification information obtained in S1120 to the image transformation machine learning model.

[0081] In S1140, the system determines whether the rendered image passes the validation check. If not, the operations of S1130 are repeated, but one or more parameters or inputs of the process are changed to generate another rendered image. For example, S1130 may be repeated while adding a different noise set to the 2D ultrasound image, applying a different text prompt, or using another 2D ultrasound image drawn from the same 3D data.

[0082] If the generated rendered image passes the validation check in S1140, the system displays the rendered image on the display in S1150.

[0083] Implementation of the subject matter and operations described herein may be realized by digital electronic circuitry, computer software, firmware, hardware, or one or more combinations thereof, including the structures and structural equivalents disclosed herein. For example, hardware may include a processor, a microprocessor, electronic circuits, electronic components, integrated circuits, etc. Implementation of the subject matter described herein may also be realized by one or more computer programs, i.e., by one or more modules of computer program instructions encoded on a computer storage medium and executed by or controlling the operation of a data processing device. Alternatively or additionally, the program instructions may be encoded in an artificially generated propagated signal. Examples of propagated signals include mechanically generated electrical, optical, or electromagnetic signals. These signals encode information that is transmitted to a suitable receiving device and executed by a data processing device. A computer storage medium may be, or may be included in, the following devices: a computer-readable storage device, a computer-readable storage substrate, a random or serial access memory array or device, or one or more combinations of these. Computer storage media other than a propagated signal may also be a source or destination of computer program instructions encoded in an artificially generated propagated signal. Also, a computer storage medium may be, or be contained in, one or more separate physical components or media (e.g., multiple CDs, disks, or other storage devices).

[0084] In various embodiments, a medical imaging device includes an ultrasound scanner capable of acquiring a 3D volume; a classifier system capable of identifying different structures within the volume (an anatomical classifier (e.g., eyes, mouth, nose, ears, arms, legs, torso, etc.)); a rendering system capable of converting the 3D image into an intermediate 2D image along with depth and object information for each pixel in the image; a diffusion-based image transformation system that takes the 2D image, the depth image, a mask image of the object, and lighting and style information and reconstructs a photorealistic image of the object based thereon; and an image display system. In embodiments, the device further includes an image validation procedure and a feedback loop. In embodiments, a neural radiance field volume is generated and displayed on the scanner. In embodiments, other 3D model representations, such as a polygon mesh, are generated that can be rendered later. In embodiments, the diffusion-based processing is performed on a remote computer. In embodiments, the image display system is a separate web-based computer, allowing the images to be viewed at a later time. In embodiments, a continuously transformed image is generated based on the time series of volumes, which can be viewed as an animation or video. In embodiments, a diffusion infill process is used to fill areas of the image without discernible features.

[0085] According to various embodiments, a computer system includes at least one processor and at least one memory containing computer-readable instructions that execute the computer-readable instructions to cause the system to acquire a two-dimensional ultrasound image, generate classification information for each of a plurality of features in the two-dimensional ultrasound image from input data including at least one of the two-dimensional ultrasound image and three-dimensional ultrasound data corresponding to the two-dimensional ultrasound image, and generate a rendered image by providing the two-dimensional ultrasound image as an input image and the classification information for each of a plurality of features in the input image to an image transformation machine learning model.

[0086] In various embodiments, the classification information includes at least one of a segmentation map and pose information.

[0087] In various embodiments, the two-dimensional ultrasound image is a first two-dimensional ultrasound image and the rendered image is a first rendered image. The at least one processor executes the readable instructions to cause the system to acquire a time series of two-dimensional ultrasound images including the first two-dimensional ultrasound image, acquire classification information for features included in each of the two-dimensional ultrasound images in the time series, and generate a time series of rendered images including the first rendered image by applying the time series of two-dimensional ultrasound images and the classification information for features included in each of the two-dimensional ultrasound images in the time series to the image transformation machine learning model.

[0088] In various embodiments, the image transformation machine learning model comprises a diffusion model.

[0089] In various embodiments, the image transformation machine learning model includes an additional machine learning model that processes the classification information, and the at least one processor executes the computer-readable instructions to cause the system to use the diffusion model to generate the rendered image according to a result of processing the classification information by the additional machine learning model.

[0090] In various embodiments, the diffusion model includes a denoiser network with multiple encoders and multiple decoders, the additional machine learning model includes copies of the multiple encoders, each having different model parameters, and generating the rendered image includes adding outputs of the multiple encoder copies to modify an output of a decoder.

[0091] In various embodiments, the image transformation machine learning model includes a generator model trained as part of a generative adversarial network.

[0092] In various embodiments, the at least one processor executes the set of computer-readable instructions to cause the system to provide the classification information as conditioning information to the image transformation machine learning model.

[0093] In various embodiments, the at least one processor executes the computer-readable instructions to cause the system to acquire the two-dimensional ultrasound image by performing volume rendering on the three-dimensional ultrasound data.

[0094] In various embodiments, the at least one processor executes the set of computer-readable instructions to cause the system to obtain a depth map of the two-dimensional ultrasound image and generate the rendered image by applying the depth map to the image transformation machine learning model.

[0095] In various embodiments, the at least one processor executes the computer-readable instructions to cause the system to generate the rendered image by providing text prompts to the image transformation machine learning model.

[0096] In various embodiments, the at least one processor executes the set of computer-readable instructions to cause the system to perform a validation check by submitting the rendered image to a validation machine learning model that outputs an image quality indicator for the rendered image, and, in response to the rendered image failing the validation check, generate a third image corresponding to the three-dimensional ultrasound data by again submitting the two-dimensional ultrasound image as an input image to the image transformation machine learning model.

[0097] In various embodiments, the image transformation machine learning model is a diffusion model that adds noise groups to the two-dimensional ultrasound image to generate the rendered image, and generating the third image includes re-applying the diffusion model to the two-dimensional ultrasound image as the input image with a different noise group added to the two-dimensional ultrasound image.

[0098] In various embodiments, generating the rendered image is performed by providing a text prompt as conditioning information to the image transformation machine learning model, and generating the third image includes reapplying the image transformation machine learning model to the two-dimensional ultrasound image as the input image, with a different text prompt as conditioning information.

[0099] In various embodiments, generating the third image includes generating another two-dimensional ultrasound image by performing volume rendering on the three-dimensional ultrasound data from a different viewpoint, and generating the third image by providing the another two-dimensional ultrasound image as an input image to the image transformation machine learning model.

[0100] According to various embodiments, a computer-implemented method for processing ultrasound imaging data comprises acquiring a two-dimensional ultrasound image, generating classification information for each of a plurality of features in the two-dimensional ultrasound image from input data including the two-dimensional ultrasound image and / or three-dimensional ultrasound data corresponding to the two-dimensional ultrasound image, and generating a rendering image by providing the two-dimensional ultrasound image as an input image and the classification information for each of the plurality of features to an image transformation machine learning model.

[0101] In various embodiments, a computer program includes computer-readable instructions that, when executed by at least one processor of a computer system, cause the system to acquire a two-dimensional ultrasound image, generate classification information for each of a plurality of features in the two-dimensional ultrasound image from input data including the two-dimensional ultrasound image and / or three-dimensional ultrasound data corresponding to the two-dimensional ultrasound image, and generate a rendered image by providing the two-dimensional ultrasound image as an input image and the classification information for each of the plurality of features to an image transformation machine learning model.

[0102] Although several embodiments have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. The novel method and system may be embodied in various other forms, and various omissions, substitutions, and modifications may be made without departing from the spirit of the invention. These embodiments and modifications are intended to be included within the scope and spirit of the invention, as well as within the scope of the invention and its equivalents as defined in the claims.

Claims

1. 1. A computer system comprising at least one processor and at least one memory containing computer readable instructions, the at least one processor executing the computer readable instructions to cause the computer system to: acquiring a two-dimensional ultrasound image; generating classification information for each of a plurality of features in the two-dimensional ultrasound image from input data including at least one of the two-dimensional ultrasound image and three-dimensional ultrasound data corresponding to the two-dimensional ultrasound image; Execute generating a rendering image by providing the two-dimensional ultrasound image as an input image and the classification information for each of a plurality of features in the input image to an image conversion machine learning model; Computer system.

2. the classification information includes at least one of a segmentation map and pose information; 10. The computer system of claim 1.

3. the two-dimensional ultrasound image is a first two-dimensional ultrasound image; the rendered image is a first rendered image; The at least one processor executes the computer-readable instructions to provide the computer system with: acquiring a time series of two-dimensional ultrasound images including the first two-dimensional ultrasound image; obtaining classification information for features included in each of the time series of two-dimensional ultrasound images; Execute generating a time series of rendering images including the first rendering image by providing the time series of two-dimensional ultrasound images and the classification information for features included in each of the time series of two-dimensional ultrasound images to the image conversion machine learning model; 10. The computer system of claim 1.

4. The image transformation machine learning model includes a diffusion model.

10. The computer system of claim 1.

5. the image transformation machine learning model includes an additional machine learning model for processing the classification information; The at least one processor executes the computer-readable instructions to provide the computer system with: using the diffusion model to generate the rendered image according to a result of processing the classification information by the additional machine learning model; 5. The computer system of claim 4.

6. the diffusion model includes a denoiser network comprising a plurality of encoders and a plurality of decoders; the additional machine learning models include copies of the plurality of encoders, each having different model parameters; generating the rendered image includes adding outputs of the copies of the encoders to modify outputs of the decoders; 6. The computer system of claim 5.

7. The image transformation machine learning model includes a generator model trained as part of a generative adversarial network.

10. The computer system of claim 1.

8. The at least one processor executes the computer-readable instructions to provide the computer system with: providing the classification information as conditioning information to the image transformation machine learning model; 10. The computer system of claim 1.

9. The at least one processor executes the computer-readable instructions to provide the computer system with: performing volume rendering on the three-dimensional ultrasound data to obtain the two-dimensional ultrasound image; 10. The computer system of claim 1.

10. The at least one processor executes the computer-readable instructions to provide the computer system with: obtaining a depth map of the two-dimensional ultrasound image; and generating the rendered image by providing the depth map to the image transformation machine learning model; 10. The computer system of claim 1.

11. The at least one processor executes the computer-readable instructions to provide the computer system with: generating the rendered image by providing a text prompt to the image transformation machine learning model; 10. The computer system of claim 1.

12. The at least one processor executes the computer-readable instructions to provide the computer system with: performing a validation check by submitting the rendered image to a validation machine learning model that outputs an image quality index for the rendered image; and generating a third image corresponding to the three-dimensional ultrasound data by again providing the two-dimensional ultrasound image as an input image to the image transformation machine learning model in response to the rendering image failing the validation check; 2. The computer system of claim 1, wherein the computer system executes the following:

13. The image transformation machine learning model is a diffusion model that generates the rendering image by adding noise groups to the two-dimensional ultrasound image, generating the third image includes reapplying the diffusion model to the two-dimensional ultrasound image as the input image, the two-dimensional ultrasound image having a different noise group added thereto; 13. The computer system of claim 12.

14. generating the rendered image by providing a text prompt as conditioning information to the image transformation machine learning model; generating the third image includes reapplying the image transformation machine learning model to the two-dimensional ultrasound image as the input image with a different text prompt added as conditioning information; 13. The computer system of claim 12.

15. generating the third image generating another two-dimensional ultrasound image by performing volume rendering on the three-dimensional ultrasound data from a different viewpoint; and generating the third image by providing the other two-dimensional ultrasound image as an input image to the image transformation machine learning model; 13. The computer system of claim 12.

16. 1. A computer-implemented method for processing ultrasound imaging data, comprising: acquiring a two-dimensional ultrasound image; generating classification information for each of a plurality of features in the two-dimensional ultrasound image from input data including at least one of the two-dimensional ultrasound image and three-dimensional ultrasound data corresponding to the two-dimensional ultrasound image; Including, generating a rendering image by providing the two-dimensional ultrasound image as an input image and the classification information for each of the plurality of features to an image conversion machine learning model; A method for processing ultrasound imaging data.

17. A computer program comprising computer-readable instructions, the instructions being executed by at least one processor of a computer system to cause the computer system to: acquiring a two-dimensional ultrasound image; generating classification information for each of a plurality of features in the two-dimensional ultrasound image from input data including at least one of the two-dimensional ultrasound image and three-dimensional ultrasound data corresponding to the two-dimensional ultrasound image; Execute generating a rendering image by providing the two-dimensional ultrasound image as an input image and the classification information for each of the plurality of features to an image conversion machine learning model; Computer program.

Citation Information

Patent Citations

  • Apparatus and method for optimization of ultrasound images

    US20160242740A1