A highly anatomically realistic multi-organ ultrasound CT image dataset construction method
By combining generative artificial intelligence and numerical solver simulation of wave fields, a large-scale multi-organ ultrasound CT image dataset was constructed, which solved the problem of lack of data in ultrasound CT and achieved high-resolution image reconstruction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2026-04-10
AI Technical Summary
Traditional ultrasound CT lacks a standardized tomographic anatomical image database, making it difficult to obtain large-scale human organ datasets, which limits the training data scale of deep learning models and affects the image reconstruction effect.
A small-scale ultrasound CT phantom dataset was generated using a physics-based style transfer method. Generative artificial intelligence was used for data augmentation, and a large-scale multi-organ ultrasound CT image dataset was constructed by combining a numerical solver to simulate the wave field. This included the processes of tissue segmentation, assignment of medium parameters, and simulation of the wave field.
The generated datasets are diverse and anatomically realistic, enabling effective training of neural networks to achieve high-resolution ultrasound CT image reconstruction, thus improving reconstruction speed and quality.
Smart Images

Figure CN121074542B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical image reconstruction, and particularly relates to a multi-organ ultrasound CT image dataset construction method with high anatomical reality. BACKGROUND
[0002] Ultrasound Computed Tomography (ultrasound CT) is a low-cost, non-radiation, high-resolution medical imaging technology. However, its image reconstruction process needs to solve a complex partial differential equation (PDE) constrained optimization inverse problem, i.e., full waveform inversion (FWI), which has challenges such as computational intensity and numerical instability.
[0003] Taking the case of medium parameters as sound velocity, under the steady-state condition, the FWI is expressed as an optimization problem constrained by the following Helmholtz equation:
[0004] The optimization objective is to minimize:
[0005] The constraint condition is:
[0006] where c(x) represents the sound velocity distribution of the medium to be measured in the calculation domain, u k (x) and u k (x f ) represent the wave fields in the calculation domain and the transducer position, respectively, y k represents the observed wave field, x defines the spatial coordinates of the calculation domain, x f represents the transducer position of the ultrasound CT device, ω represents the angular frequency of the wave field, ρ k (x) represents the wave source term in the calculation domain, and k = 1, …, K, K is the number of transducers.
[0007] In recent years, researchers have used deep learning models to solve the above challenges, and the methods mainly fall into two categories. The first category of methods uses neural networks such as neural operators as surrogate models to forward simulate the physical process from tissue sound velocity c to wave field u; the second category of methods directly trains end-to-end neural networks to reconstruct the observed wave field data y into tissue sound velocity c.
[0008] Despite the above differences, both types of methods require a large amount of real training data, i.e., pairs of sound speed-wave field data, which requires accurate organ sound speed modeling and wave field simulation. However, unlike X-ray CT or magnetic resonance imaging (MRI) technology, traditional ultrasound imaging lacks a standardized tomographic anatomic image database. For ultrasound CT, the limitations of the instrument equipment and the inefficiency of the numerical solver make it difficult to collect a large amount of real data, and it is difficult to obtain a large-scale human organ data set. Therefore, it is urgent to create a method to expand the size of the limited data set, so as to effectively train a deep learning model on a large scale. SUMMARY
[0009] In view of the problems of the prior art, the present application provides a highly anatomically realistic multi-organ ultrasound CT image data set construction method.
[0010] The highly anatomically realistic multi-organ ultrasound CT image data set construction method of the present application comprises the following steps:
[0011] 1) Generate a small-scale ultrasound CT phantom data set using a physics-based style transfer method:
[0012] Provide the original image of the organ;
[0013] According to the original image of the organ, use an artificial intelligence segmentation model to segment the organ into multiple types of tissues;
[0014] According to the acoustic characteristics of the tissues, assign appropriate medium parameter values to each tissue, thereby generating a small-scale but anatomically realistic ultrasound CT phantom data set;
[0015] 2) Data augmentation using generative artificial intelligence:
[0016] Use the ultrasound CT phantom data set generated in step 1) to fine-tune the basic generative artificial intelligence model, transfer the anatomical knowledge in the ultrasound CT phantom data set to the basic generative artificial intelligence model, and the fine-tuned basic generative artificial intelligence model generates a large amount of new data according to the ultrasound CT phantom data set generated in step 1); filter out the output results that do not conform to the anatomical principles, and check the medium parameter values, if they exceed the reference range of the medium parameters of the corresponding tissues, use image threshold segmentation method to re-segment the tissues and assign medium parameter values, thereby generating a large number of diversified and physically realistic phantom data, completing the construction of a large-scale phantom data set;
[0017] 3) Use a numerical solver to simulate the corresponding wave field:
[0018] The wave field is solved using a wave equation numerical solver, with the parameters of the accurate ultrasound CT experimental device as the parameters of the numerical simulation, to simulate the output scattered wave field after the interaction of the sound wave with the phantom, to complete a large-scale medium parameter
[0019] The wave field data pairs are constructed to obtain a multi-organ ultrasound CT image data set.
[0020] In step 1), the original image is a two-dimensional image of a digital simulation model slice or a clinical image of another modality than ultrasound CT; for the breast, a two-dimensional image of a digital simulation model slice is provided; for the upper arm and the thigh, a clinical image of another modality than ultrasound CT is provided; the other modality than ultrasound CT is X-ray CT or magnetic resonance imaging (MRI).
[0021] In the segmentation of multiple types of tissues, for the breast, the segmentation is into skin, fat, gland and other tissues; for the upper arm and the thigh, the segmentation is into skin, fat, muscle, bone and other tissues. The medium parameter is the sound speed, the density or the acoustic attenuation. The type of the medium parameter is first selected, and then the medium parameter value at the corresponding position is assigned to different tissues, and the medium parameter value is a two-dimensional distribution, i.e., a medium parameter distribution is obtained; the medium parameter value assignment method: using a linear or polynomial function to map the pixel value of the original image to the reference range of the medium parameter of the human tissue.
[0022] The data volume of a small-scale but anatomically real ultrasound CT phantom data set is about 100 to 1000 phantom images of a single organ. The ultrasound CT phantom data set includes multiple two-dimensional phantom images, and the pixel value of each pixel point of the phantom image corresponds to the medium parameter value at the corresponding position, which is also a two-dimensional matrix.
[0023] In step 2), the number of large-scale phantom data sets is about 5000 to 10000 phantom images of a single organ.
[0024] The input of the basic generative artificial intelligence model is the ultrasound CT phantom data set and the prompt, and the output is a new phantom image. The prompt is the requirement for the output result.
[0025] In step 3), the wave equation is the Helmholtz equation. When the medium parameter is the sound speed, the wave equation is the standard Helmholtz equation; when the medium parameter is the density, the wave equation adds a density gradient term in the Helmholtz equation; when the medium parameter is the acoustic attenuation, the wave equation uses the Helmholtz equation of the complex wave number.
[0026] In the simulation of the scattered wave field by means of the numerical solver, a plurality of sound sources fixed in position are provided in the numerical solver, and each sound source respectively emits a plurality of sound waves of different frequencies to the phantom. The frequency of the sound source is 0.2-1.2MHz, and the interval is 0.01-0.1MHz. Through the numerical solver simulation, the sound waves emitted by the fixed sound source pass through the phantom, and the corresponding scattered wave field output is obtained.
[0027] Advantages of the present application:
[0028] The ultrasound CT data set constructed by the present application has the characteristics of large scale, diversity and high anatomical reality, and the neural network trained on this basis can realize the high-resolution ultrasound CT image reconstruction task. BRIEF DESCRIPTION OF DRAWINGS
[0029] Figure 1 Flowchart of the highly anatomically realistic multi-organ ultrasound CT image data set construction method of the present application;
[0030] Figure 2 Breast ultrasound CT reconstruction result graph obtained for an embodiment of the highly anatomically realistic multi-organ ultrasound CT image data set construction method according to the present application;
[0031] Figure 3 Large arm and thigh ultrasound CT reconstruction result graph obtained for an embodiment of the highly anatomically realistic multi-organ ultrasound CT image data set construction method according to the present application. DETAILED DESCRIPTION
[0032] The present application will be further described below with reference to the accompanying drawings and specific embodiments.
[0033] In this embodiment, a sound speed-wave field data set containing three organs of breast, large arm and thigh is constructed, and a neural network is trained based on the data set to complete the ultrasound CT image reconstruction of the corresponding organs to verify its feasibility. The data set includes 22047 digital organ models, i.e. phantoms, containing 7520 breasts, 7526 large arms and 7001 thighs. For each sound speed phantom, we simulated the wave field of 8 different frequencies from 64 different point sources, resulting in a total of 11288064 sound speed-wave field data pairs; the data set website is https: / / open-waves-usct.github.io / .
[0034] As shown in Figure 1 The highly anatomically realistic multi-organ ultrasound CT image data set construction method of the present embodiment includes the following steps:
[0035] 1) Generate a small-scale ultrasound CT phantom data set using a physics-based style transfer method:
[0036] For the breast, two-dimensional images of digital simulation model slices are provided; for the upper arm and thigh, clinical images of X-ray CT are provided;
[0037] Using an artificial intelligence segmentation model, the breast is segmented into skin, fat, gland and other tissues according to the two-dimensional images of the digital simulation model slices of the breast; the upper arm and thigh are segmented into skin, fat, muscle, bone and other tissues according to the clinical images of X-ray CT of the upper arm and thigh;
[0038] The medium parameter of this embodiment adopts sound velocity, and according to the specific acoustic characteristics of the tissue, the pixel value of the original image is mapped into the reference range of the sound velocity of the human tissue using a linear or polynomial function, and appropriate two-dimensional sound velocity values are allocated to each tissue, thereby generating a small-scale but anatomically real ultrasound CT phantom data set;
[0039] The specific method for generating a small-scale ultrasound CT sound velocity phantom data set of three organs is as follows:
[0040] a) Breast:
[0041] 1-1. Create an anatomically real digital three-dimensional breast phantom using the VICTRE tool of the US-FDA project; according to the different densities, the generated breast is divided into four types of full fat (FAT), fibrous gland (FIB), heterogeneous (HET) and extremely dense (EXD);
[0042] 1-2. Convert the three-dimensional phantom to a two-dimensional phantom by random slicing; in addition, the two-dimensional phantom slices are translated to the center of the region and randomly scaled to increase diversity;
[0043] 1-3. Segment different breast tissues such as skin, fat and gland using the Segment Anything Model visual base model, and assign physically real sound velocities to the corresponding tissue regions, with a small random perturbation for each sound velocity;
[0044] 1-4. Set the medium parameter value around the breast model to the sound velocity of water to replicate the experimental environment;
[0045] b) Upper arm and thigh:
[0046] 1-1. Collect X-ray CT arm scan data of 151 volunteers and X-ray CT leg scan data of another 41 volunteers at Peking University Third Hospital;
[0047] 1-2. Segment the X-ray CT images using the Segment Anything Model visual base model;
[0048] 1-3. Map each segmented region to a sound speed range consistent with physical reality; randomly rotate the samples to enhance the realism of the data;
[0049] 1-4. Set the surrounding medium parameter value to the speed of sound of water;
[0050] 2) Using generative artificial intelligence for data augmentation:
[0051] Using the ultrasound CT phantom dataset generated in step 1), the DreamBooth method is applied to fine-tune the text-to-image pre-training of the basic generative artificial intelligence model StableDiffusion. Fine-tuning involves retraining the model using the ultrasound CT sound velocity phantom dataset from step 1.
[0052] i. For each organ, fine-tune the Stable Diffusion model using the following prompt: "Grayscale image of phantom cross-section of breast / upper arm / thigh";
[0053] ii. Use the same prompts to generate images from the fine-tuned model and filter out unreasonable output results, such as missing or duplicated tissues, based on anatomical principles;
[0054] iii. Check the sound velocity values in the image. If they exceed the reference range, re-segment the tissue using Otsu's image thresholding method and assign corresponding sound velocity values. The approximate reference range for sound velocity is as follows: skin 1600–1700 m / s, fat…
[0055] 1420–1480 m / s, muscles 1550–1650 m / s, glands 1500–1600 m / s, bones 3000–4000 m / s;
[0056] This generates a large amount of diverse and physically realistic phantom data, completing the construction of a large-scale phantom dataset;
[0057] 3) Use a numerical solver to simulate and solve the corresponding wave field:
[0058] The scattered wave field was obtained by solving the Helmholtz equation using a numerical solver based on the convergent Born series (CBS) principle. The standard Helmholtz equation was used, and the parameters of the precise ultrasound CT experimental setup were used as the parameters for the numerical simulation. The parameters were set as follows: 256 transducers were uniformly distributed on a ring with a diameter of approximately 22 cm, and the frequencies were set to 0.25 MHz to 0.6 MHz, with 8 discrete values at equal intervals of 0.05 MHz. In the actual simulation, one of the 256 transducers was used at equal intervals of 4, for a total of 64 transducers. The output scattered wave field after the interaction between the sound wave and the phantom was obtained through simulation, and a large-scale sound velocity-wave field data pair was constructed to obtain a multi-organ ultrasound CT image dataset.
[0059] The reconstruction results of the breast, the upper arm and the thigh using the trained neural network on the experimental data are shown in FIGS. 1 1A, 1 1B and 1 1C, respectively. Figure 2 and Figure 3 as shown in FIGS. 1 1A, 1 1B and 1 1C, respectively.
[0060] Figure 2 The reconstruction results of the breast using four methods, i.e., full waveform inversion (FWI) using a traditional numerical solver, full waveform inversion using a neural network trained by the data set generated by the present application (Neural-FWI), delay and sum method (DAS) and time-of-flight tomography (ToFT), are respectively shown in FIGS. 10A, 10B, 10C and 10D. Among them, the reconstruction result using the neural network trained by the data set generated by the present application has a resolution similar to that of the traditional FWI, but is significantly improved in speed; and the resolution is significantly higher than that of the delay and sum method and the time-of-flight method.
[0061] Figure 3 The reconstruction results of the upper arm and the thigh are shown in FIGS. 12A and 12B. Among them, the first and second columns are the three-dimensional reconstruction results of the upper arm and the thigh, respectively, and the third and fourth columns are the two-dimensional cross-sectional views of the three-dimensional reconstruction results of the upper arm and the thigh, respectively. Neural-FWI is full waveform inversion using a neural network trained by the data set generated by the present application, MRI is magnetic resonance imaging, FWI is full waveform inversion using a traditional numerical solver, and DAS is the reconstruction result using the delay and sum method. The reconstruction result using the neural network trained by the data set generated by the present application achieves high definition close to that of magnetic resonance imaging, and the reconstruction speed is significantly faster than that of the traditional numerical FWI method.
[0062] Finally, it should be noted that the purpose of the disclosed embodiments is to help further understand the present application, but those skilled in the art can understand that various replacements and modifications are possible without departing from the spirit and scope of the present application and the appended claims. Therefore, the present application should not be limited to the disclosed embodiments, and the scope of the present application is defined by the scope of the claims.
Claims
1. A method for constructing a multi-organ ultrasound CT image dataset, characterized in that, The method includes the following steps: 1) A small-scale ultrasound CT phantom dataset was generated using a physics-based style transfer method, which included multiple phantom images. The phantom images reflected the media parameter values assigned to the corresponding locations after tissue segmentation. 2) Using generative artificial intelligence for data augmentation: Using the ultrasound CT phantom dataset generated in step 1), the basic generative artificial intelligence model is fine-tuned. The fine-tuned basic generative artificial intelligence model generates a large-scale new dataset based on the ultrasound CT phantom dataset generated in step 1). Output results that do not conform to anatomical principles are filtered out, and the media parameter values are checked. If they exceed the reference range of the media parameters of the corresponding tissue, the tissue is re-segmented and the media parameter values are reassigned, thereby generating a large amount of diverse and physically realistic phantom data and completing the construction of a large-scale phantom dataset. 3) Use a numerical solver to simulate and solve the corresponding wave field: The wave field was solved using a wave equation numerical solver. The wave equation is the Helmholtz equation. When the medium parameter is the speed of sound, the wave equation is the standard Helmholtz equation. When the medium parameter is density, the wave equation is the Helmholtz equation with added density gradient terms. When the medium parameter is sound attenuation, the wave equation is the Helmholtz equation with complex wave numbers. The parameters of the precise ultrasound CT experimental device were used as the parameters for the numerical simulation. The output scattered wave field after the interaction between the sound wave and the phantom was simulated, and a large-scale medium parameter-wave field data pair was constructed to obtain a multi-organ ultrasound CT image dataset.
2. The method as described in claim 1, characterized in that, In step 1), a small-scale ultrasound CT phantom dataset is generated, including the following steps: Provides raw images of the organs; Based on the original images of the organs, an artificial intelligence segmentation model is used to segment the organs into multiple tissue types; By assigning appropriate media parameter values to each tissue based on its specific acoustic properties, a small-scale but anatomically realistic ultrasound CT phantom dataset was generated.
3. The method as described in claim 2, characterized in that, The original images are two-dimensional images of digitally simulated model slices or other modal clinical images different from ultrasound CT; for the breast, two-dimensional images of digitally simulated model slices are provided; for the upper arm and thigh, other modal clinical images different from ultrasound CT are provided.
4. The method as described in claim 1, characterized in that, In step 1), the data size of the small but anatomically realistic ultrasound CT phantom dataset is 100 to 1000.
5. The method as described in claim 1, characterized in that, In step 2), the number of large-scale phantom datasets is 5,000 to 10,000.
6. The method as described in claim 1, characterized in that, In step 2), the tissue is re-segmented and media parameter values are assigned using an image thresholding segmentation method.
7. The method as described in claim 1, characterized in that, In step 3), a numerical solver is used to simulate the scattered wave field. Multiple sound sources with fixed positions are set in the numerical solver, and each sound source emits multiple sound waves of different frequencies to the phantom.
Citation Information
Patent Citations
Three-dimensional image guide positioning method and system and storage medium
CN113041516A
Ultrasound image segmentation method and apparatus, terminal device, and storage medium
US20230386048A1