A Data Standardization Processing Method for AI Ultrasonic Large Model Testing
By identifying differences in equipment imaging features and assessing image quality, a benchmark center was selected for feature distribution calibration and defect repair. This solved the problems of data distribution offset and quality defects in multi-center evaluation of AI ultrasound large models, achieving the unification and quality improvement of cross-center data, and ensuring the stability and accuracy of the model.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- XIEHE HOSPITAL ATTACHED TO TONGJI MEDICAL COLLEGE HUAZHONG SCI & TECH UNIV
- Filing Date
- 2026-03-02
- Publication Date
- 2026-05-26
AI Technical Summary
In multi-center, cross-device diagnostic performance evaluations of AI ultrasound large models, existing technologies have failed to effectively address the issues of data distribution shifts and evaluation result distortions caused by differences in device imaging characteristics and inconsistent image quality.
By acquiring multi-center ultrasound image data, identifying differences in equipment imaging characteristics, selecting a reference center, using a feature distribution calibration algorithm to eliminate equipment differences, evaluating quality based on image clarity and integrity indicators, and specifically repairing defects, a standardized ultrasound image dataset is established.
It has achieved the unification and quality improvement of cross-center ultrasound image data, reduced the coefficient of variation of model performance indicators, provided a scientific and reliable testing scheme, and ensured the stability of AI ultrasound large model in clinical application.
Smart Images

Figure CN121767352B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of ultrasound data processing technology, and in particular to a data standardization processing method for AI ultrasound large model testing. Background Technology
[0002] In recent years, the application of AI-powered ultrasound large-scale models in clinical diagnosis has advanced rapidly, and their diagnostic efficiency and accuracy have been gradually verified. However, the clinical implementation of AI-powered ultrasound large-scale models requires evaluation of their generalization capabilities across multiple centers and devices. This means that the model must maintain stable diagnostic performance on data collected from different medical centers and different brands and models of ultrasound equipment. This is a key criterion for judging whether the model has clinical practical value.
[0003] Current AI ultrasound large-scale model testing mainly faces the following problems: First, differences in data equipment interfere with the accuracy of evaluation. Ultrasound equipment from different centers differs in imaging principles, probe frequencies, and image processing algorithms, resulting in significant differences in resolution, grayscale contrast, noise levels, and tissue boundary clarity of the acquired ultrasound images. When directly using raw data to test the model, the "data distribution shift" caused by equipment differences can mask the model's generalization ability, leading to distorted evaluation results. Second, inconsistent data quality affects the reliability of evaluation. Test data provided by external centers may have quality defects such as image blurring, missing target areas, and artifact interference due to differences in the operating physician's technique. Existing testing methods often do not standardize data quality, and poor-quality data can lead to incorrect diagnostic results from the model, thus misjudging the model's insufficient generalization ability.
[0004] Patent application CN116721779A discloses a medical data preprocessing technology, including the following steps: removing irrelevant symbols from medical data and correcting errors in the text data; segmenting the text data into different fields using multiple medical word segmenters to obtain segmentation results, and inputting the segmentation results into a secondary word segmenter to obtain the final segmentation result; constructing a medical knowledge graph, and labeling the medical fields in the final segmentation result based on the medical knowledge graph to obtain partially labeled medical data; and using swarm intelligence to label the unlabeled fields in the partially labeled medical data to obtain fully labeled medical data. This technical solution achieves standardized management of medical data and optimizes the data labeling workflow by constructing a data preprocessing module to clean, extract key information, and label medical data from different sources, providing a suitable data foundation for training medical AI models. However, this technical solution mainly focuses on data format unification and labeling, failing to provide an effective standardized processing solution for multi-center ultrasound images due to differences in equipment imaging characteristics and image quality defects. Summary of the Invention
[0005] In view of this, the present invention proposes a data standardization processing method for AI ultrasound large model testing, in order to solve the problem of inconsistent data distribution caused by differences in the imaging characteristics of equipment in multi-center ultrasound image data in the prior art, as well as the problem of diverse types of ultrasound image quality defects and the lack of targeted standardized repair mechanisms.
[0006] The technical solution of this invention is implemented as follows: This invention provides a data standardization processing method for AI ultrasound large model testing, including the following steps:
[0007] S1. Acquire ultrasound image data from multiple external data acquisition centers and record the metadata of each acquisition center, including the ultrasound equipment brand and model, acquisition parameters, and acquisition environment information;
[0008] S2. Based on the metadata, identify the differences in equipment imaging features of different data acquisition centers. Based on the similarity of image feature distribution of each data acquisition center, select at least one reference center from multiple data acquisition centers. Based on the image feature distribution of the reference center, establish a target reference distribution. Use a feature distribution calibration algorithm to calibrate the grayscale features and texture features of the ultrasound image data of non-reference centers to the target reference distribution, eliminate the imaging differences between different equipment brands and models, and obtain the ultrasound image data after equipment calibration.
[0009] S3. Based on image clarity and integrity indicators, perform quality assessment on the ultrasound image data after the device is calibrated, identify ultrasound images with quality defects and their defect types, including image blurring, imaging artifacts and composite defects, and use corresponding image restoration strategies for different defect types to perform quality restoration processing to obtain a standardized ultrasound image dataset.
[0010] S4. Use the standardized ultrasound image dataset for AI ultrasound large model testing.
[0011] Based on the above technical solutions, preferably, in step S1, the ultrasound equipment brand and model include the equipment brand and model, the acquisition parameters include probe frequency, gain value, dynamic range and imaging mode, and the acquisition environment information includes the operator's title, acquisition site, acquisition time, and patient's age and gender.
[0012] Based on the above technical solutions, preferably, step S2 specifically includes:
[0013] S21. Extract features from the ultrasound images of each data acquisition center to obtain the image feature distribution of each data acquisition center. The image feature distribution includes gray-level statistical distribution, texture feature distribution, and frequency domain feature distribution.
[0014] S22. Calculate the image feature distribution distance between each data acquisition center, construct the distribution distance matrix between the centers, calculate the coverage index based on the distribution distance matrix, and select at least one reference center from multiple data acquisition centers according to the coverage index. The coverage index represents the average of the minimum distribution distances of all data acquisition centers to the reference center combination.
[0015] S23. Establish a target reference distribution based on the image feature distribution of the reference center;
[0016] S24. The ultrasound image data of the non-reference center is calibrated to the target reference distribution using a feature distribution calibration algorithm, wherein the feature distribution calibration algorithm includes grayscale distribution alignment and texture feature mapping, to obtain the ultrasound image data after calibration by the device.
[0017] Based on the above technical solutions, preferably, in step S21, the gray-level statistical distribution is obtained by calculating the gray-level mean, gray-level standard deviation, skewness and kurtosis of the image; the texture feature distribution is obtained by extracting the contrast, correlation, energy and homogeneity of the image through the gray-level co-occurrence matrix; and the frequency domain feature distribution is obtained by converting the image to the frequency domain space through Fourier transform and extracting spectral features.
[0018] Based on the above technical solution, preferably, in step S22, the Wasserstein distance is used to calculate the image feature distribution distance between each data acquisition center. The Wasserstein distance measures the transmission cost between the image feature distributions of two centers. The candidate reference center combination with the smallest coverage index is selected as the final reference center. The formula for calculating the coverage index is:
[0019] ;
[0020] in, The number of all data acquisition centers. For candidate benchmark center combinations, The image feature distribution of the k-th data acquisition center. The image feature distribution of the j-th center in the reference center combination. Let the Wasserstein distance be the distance between the k-th center and the j-th center. This represents the minimum Wasserstein distance from the k-th center to all centers in the combination of reference centers.
[0021] Based on the above technical solutions, preferably, in step S24, the gray-level distribution alignment adopts a histogram matching method, which groups the ultrasound images of non-reference centers according to the equipment brand, extracts the gray-level histogram distribution for each equipment brand group, calculates the cumulative distribution function of the gray-level histogram distribution, maps the cumulative distribution function to the cumulative distribution function corresponding to the target reference distribution, establishes a gray-level value mapping relationship, and realizes the gray-level distribution alignment of all images within the same equipment brand group;
[0022] The texture feature mapping adopts a multinomial regression method to establish a regression model between device parameters and texture features. The regression model is used to eliminate residual feature differences caused by device brand and model.
[0023] Based on the above technical solutions, preferably, step S3 specifically includes:
[0024] S31. Perform a quality assessment on the calibrated ultrasound image data of the device. The quality assessment includes sharpness assessment, integrity assessment and artifact-free assessment to obtain a comprehensive quality score.
[0025] S32. Divide the calibrated ultrasound image data of the device into regions, calculate the quality score of each region, identify ultrasound images with a comprehensive quality score lower than the quality threshold, and construct a quality gradient field based on the quality score difference between adjacent regions.
[0026] S33. Determine the defect type based on the quality score of each region block. The defect types include fuzzy defects, artifact defects, composite defects, and no defects.
[0027] S34. Appropriate image inpainting algorithms are used to repair regions with different defect types, and the repair intensity of each region is dynamically adjusted according to the mass gradient field.
[0028] S35. Perform quality verification on each repaired region block, use a partition verification mechanism to judge the repair effect of each region block, retain the successfully repaired region blocks, restore the unrepaired region blocks to their original pixel values, and obtain the standardized ultrasound image dataset.
[0029] Based on the above technical solutions, preferably, in step S31, the sharpness assessment obtains a sharpness score by calculating the average gradient magnitude of the image at multiple scales, the integrity assessment obtains an integrity score by calculating the proportion of the effective field of view area to the total area in the image, the artifact-free assessment obtains an artifact-free score by detecting the number and intensity of abnormal peaks in the frequency domain, and the comprehensive quality score is a weighted sum of the sharpness score, the integrity score, and the artifact-free score.
[0030] Based on the above technical solution, preferably, in step S34, a denoising algorithm is used to process the fuzzy defect region, a filtering algorithm is used to process the artifact defect region, and filtering and denoising are performed sequentially on the composite defect region. The mass gradient field is used to identify boundary regions where the mass gradient amplitude is greater than a preset gradient threshold. The repair intensity is reduced in the boundary regions, and a standard repair intensity is used in the internal regions with uniform mass. The formula for calculating the repair intensity parameter of the fuzzy defect region is as follows:
[0031] ;
[0032] in, The repair strength parameter for the region block in row m and column n is... Basic repair parameters, For clarity weighting coefficients, Score the sharpness of the region block in row m and column n. These are the gradient field weight coefficients. The mass gradient magnitude of the region in the m-th row and n-th column. This is the gradient field sensitivity threshold.
[0033] Based on the above technical solutions, preferably, in step S35, the partition verification mechanism includes:
[0034] Calculate the edge intensity of each region before and after restoration, and obtain the edge intensity change rate. The edge intensity change rate is defined as the ratio of the difference in edge intensity before and after restoration to the edge intensity before restoration. For regions where the edge intensity change rate exceeds a preset threshold, the restoration of the region is determined to be unsuccessful, and the region is restored to its original pixel value before restoration. For regions where the edge intensity change rate does not exceed the preset threshold, the restoration of the region is determined to be successful, and the restored pixel value is retained. After all regions are verified, image fusion is performed.
[0035] The data standardization processing method for AI ultrasound large model testing of the present invention has the following advantages over the prior art:
[0036] (1) Through dynamic benchmark selection, feature distribution calibration, and quality gradient field-guided partition adaptive repair, non-physiological variations introduced by equipment differences and quality defects in the test data are eliminated, and the cross-center variation coefficient of the model performance index is significantly reduced, providing a scientific and reliable test scheme for the clinical application of AI ultrasound large model.
[0037] (2) Regarding the elimination of equipment differences, a dynamic benchmark selection algorithm based on Wasserstein distance is used to select benchmark center combinations based on quantified distribution distance and coverage indicators. Compared with manual experience selection, this method eliminates subjectivity, reduces the average distribution distance of all non-benchmark centers to the benchmark group, and makes the benchmark group more fully cover the data distribution space. Through a hybrid calibration method combining histogram matching and multinomial regression, equipment parameters and image features are mapped to the consensus benchmark feature space, which effectively eliminates the differences in grayscale distribution, texture features, and boundary clarity of images from different brands and models of ultrasound equipment, and achieves the unification of data distribution.
[0038] (3) In terms of quality standardization, a three-dimensional quality scoring system encompassing clarity, integrity, and artifact-free characteristics was constructed to achieve quantitative assessment of ultrasound image quality. The concept of a quality gradient field was introduced, extending the quality score from a scalar field to a vector field with spatial variation information. The spatial gradient of the quality score was analyzed to distinguish between uniformly low-quality regions and boundary regions. A partitioned adaptive repair algorithm based on the quality gradient field was used to specifically repair different defect types in different regions. In boundary regions with drastic quality changes, the repair intensity was automatically reduced using an exponential decay function, while sufficient repair was implemented in uniformly quality internal regions, achieving a balance between image quality improvement and protection of critical diagnostic information. The success rate of repairing mildly poor-quality data was improved, and the retention rate of lesion boundary clarity was enhanced, avoiding the problem of the repair algorithm destroying boundary features. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0040] Figure 1 This is a flowchart of the data standardization processing method for AI ultrasound large model testing according to the present invention;
[0041] Figure 2 This is a flowchart illustrating the dynamic benchmark selection algorithm of the present invention;
[0042] Figure 3 This is a flowchart illustrating the quality assessment method of the present invention;
[0043] Figure 4 This is a schematic diagram of the partition adaptive repair process of the present invention. Detailed Implementation
[0044] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0045] like Figure 1 As shown, this invention provides a data standardization processing method for AI ultrasound large model testing, including the following steps:
[0046] S1. Acquire ultrasound image data from multiple external data acquisition centers and record the metadata of each acquisition center, including the ultrasound equipment brand and model, acquisition parameters, and acquisition environment information;
[0047] S2. Based on the metadata, identify the differences in equipment imaging features of different data acquisition centers. Based on the similarity of image feature distribution of each data acquisition center, select at least one reference center from multiple data acquisition centers. Based on the image feature distribution of the reference center, establish a target reference distribution. Use a feature distribution calibration algorithm to calibrate the grayscale features and texture features of the ultrasound image data of non-reference centers to the target reference distribution, eliminate the imaging differences between different equipment brands and models, and obtain the ultrasound image data after equipment calibration.
[0048] S3. Based on image clarity and integrity indicators, perform quality assessment on the ultrasound image data after the device is calibrated, identify ultrasound images with quality defects and their defect types, including image blurring, imaging artifacts and composite defects, and use corresponding image restoration strategies for different defect types to perform quality restoration processing to obtain a standardized ultrasound image dataset.
[0049] S4. Use the standardized ultrasound image dataset for AI ultrasound large model testing.
[0050] In one embodiment of the present invention, the data collection scope of step S1 covers at least 6 medical centers of different levels, including 2 tertiary hospitals, 2 secondary hospitals, and 2 primary hospitals, ensuring that the data sources cover different medical resource allocation scenarios. The equipment covers at least 4 mainstream ultrasound equipment brands, with each brand containing 2-3 different models. The data types cover 3 major categories: abdominal ultrasound, cardiovascular ultrasound, and superficial organ ultrasound, with a sample size of no less than 1200 cases for each category. Each data case includes a raw image in DICOM format, a clinical diagnostic report, and a device model parameter document. Clearly invalid data, such as images with completely black screens or no anatomical structure, are removed during collection.
[0051] The ultrasound equipment brand and model information includes the equipment brand and model number. Acquisition parameters include probe frequency, gain, dynamic range, and imaging mode. Acquisition environment information includes the operator's title, acquisition site, acquisition time, and the patient's age and gender. Probe frequency is in MHz, gain is in dB, dynamic range is in dB, and imaging modes include B-mode ultrasound, M-mode ultrasound, and Doppler ultrasound. Operator titles are categorized as attending physician or above, or resident physician. Acquisition sites are recorded as specific anatomical locations such as the right lobe of the liver and the left lobe of the thyroid gland. Acquisition time is accurate to the minute. Patient age and gender have been anonymized, retaining only age range and gender information. A database is used to store the mapping relationship between data IDs and metadata, supporting rapid matching of equipment parameters during subsequent algorithm calls.
[0052] In one embodiment of the present invention, such as Figure 2 As shown, step S2 specifically includes the following steps:
[0053] S21. Extract features from the ultrasound images of each data acquisition center to obtain the image feature distribution of each data acquisition center. The image feature distribution includes gray-level statistical distribution, texture feature distribution and frequency domain feature distribution.
[0054] Specifically, the gray-level statistical distribution is obtained by calculating the gray-level mean, standard deviation, skewness, and kurtosis of the image. Skewness reflects the asymmetry of the gray-level distribution, and kurtosis reflects the sharpness of the gray-level distribution. The texture feature distribution is obtained by extracting the contrast, correlation, energy, and homogeneity of the image through the gray-level co-occurrence matrix. The gray-level co-occurrence matrix describes the spatial relationship of gray levels in the image, contrast reflects the gray-level difference between adjacent pixels, energy reflects the uniformity of the texture, and homogeneity reflects the smoothness of the texture. The frequency domain feature distribution is obtained by transforming the image to the frequency domain space through Fourier transform and extracting spectral features. A two-dimensional Fourier transform is performed on the image at each center to obtain the frequency domain feature vector. The gray-level statistical distribution, texture feature distribution, and frequency domain feature distribution are concatenated to obtain the first... The comprehensive feature vector of each center .
[0055] S22. Calculate the image feature distribution distance between each data acquisition center, construct a distribution distance matrix between the centers, calculate the coverage index based on the distribution distance matrix, and select at least one reference center from multiple data acquisition centers according to the coverage index. The coverage index represents the average of the minimum distribution distances of all data acquisition centers to the reference center combination.
[0056] Specifically, the Wasserstein distance is used to calculate the distance between image feature distributions of each data acquisition center. This Wasserstein distance measures the transmission cost between the image feature distributions of two centers. Compared to KL divergence or Euclidean distance, the Wasserstein distance can capture the geometric structure information of the distribution and is more sensitive to local changes in the distribution. For multi-center ultrasound data, the image feature distributions of different centers may differ in multiple dimensions such as frequency domain, texture, and grayscale statistics. The Wasserstein distance can comprehensively measure these multi-dimensional differences. The calculation of the Wasserstein distance is implemented using the Sinkhorn iterative algorithm, which transforms the original optimal transmission problem into a convex optimization problem through entropy regularization.
[0057] Construct the distribution distance matrix between centers , where matrix elements For the first The center and the first The Wasserstein distance between each center is used. Based on the distance matrix, the centers are divided into several clusters using a hierarchical clustering algorithm. Centers within each cluster have similar feature distributions, and the average linkage method is used for clustering. The candidate baseline center combination with the smallest coverage index is selected as the final baseline center. The formula for calculating the coverage index is:
[0058] ;
[0059] in, The number of all data acquisition centers. For candidate benchmark center combinations, For the first Image feature distribution of each data acquisition center For the first in the reference center combination Image feature distribution at each center, For the first The center and the first Wasserstein distance between the centers Indicates the first The coverage index is the minimum Wasserstein distance from a given center to all centers in the baseline center set. The coverage index measures the average minimum distance from all centers to the baseline set; a smaller coverage value indicates that the baseline set is more representative of all centers.
[0060] This algorithm iterates through all possible benchmark combinations, including combinations with 1, 2, 3, and 4 centers, calculating the coverage index for each candidate benchmark group. The benchmark group with the smallest coverage is selected as the final benchmark group. The benchmark group size is typically between 2 and 4 centers; too few benchmark centers may result in insufficient coverage, while too many benchmark centers will increase the computational complexity of subsequent calibrations. This dynamic benchmark selection algorithm is based on quantified distribution distance and coverage indexes, and the selection process has a clear mathematical objective, eliminating subjectivity. The benchmark group size is dynamically determined according to the characteristics of the data distribution. By using the Wasserstein distance metric, it can identify center combinations that are complementary in multiple dimensions such as frequency domain, texture, and grayscale, thereby improving the coverage capability of the benchmark group.
[0061] S23. Establish a target reference distribution based on the image feature distribution of the reference center.
[0062] Specifically, the consensus benchmark feature is calculated by taking the weighted average feature of the benchmark group. The weighting formula is as follows: The weight According to the The data quality score and equipment representativeness of each benchmark center are determined. Equipment representativeness is defined as the degree of difference between the equipment brand of this center and the equipment brand of other benchmark centers. If the benchmark group includes three brands A, B, and C, then the equipment representativeness of each brand is 1. If it includes two brands A and one brand B, then the equipment representativeness of brand A is reduced in weight. The weights satisfy the normalization condition. .
[0063] S24. The ultrasound image data of the non-reference center is calibrated to the target reference distribution using a feature distribution calibration algorithm, wherein the feature distribution calibration algorithm includes grayscale distribution alignment and texture feature mapping, to obtain the ultrasound image data after calibration by the device.
[0064] The grayscale distribution alignment employs a histogram matching method. Ultrasound images outside the reference center are grouped according to equipment brand. For each equipment brand group, a grayscale histogram distribution is extracted, and the cumulative distribution function (CDF) of the grayscale histogram distribution is calculated. This CDF is then mapped to the CDF corresponding to the target reference distribution, establishing a grayscale value mapping relationship and achieving grayscale distribution alignment for all images within the same equipment brand group. Specifically, images from the external acquisition center are grouped according to equipment brand, with different models of the same brand grouped together. For each equipment group, the grayscale histogram distribution of all images within that group is extracted, and the CDF is calculated. For each grayscale level of the equipment group images... By solving Obtain the target gray level This enables grayscale value mapping.
[0065] The texture feature mapping employs a multinomial regression method to establish a regression model between device parameters and texture features. This regression model is used to eliminate residual feature differences caused by device brand and model. Probe frequency is extracted from device metadata. Gain value Dynamic range As a device parameter vector For the image after histogram matching, extract its grayscale mean. Gray standard deviation Boundary sharpness calculated using the Canny algorithm , constitute image feature vector Boundary sharpness is defined as the ratio of the number of boundary pixels to the total number of pixels in the image. A second-order polynomial regression model is established to map device parameters and image features to the consensus benchmark feature space. The mapping formula is as follows:
[0066] ;
[0067] in and This is the weight matrix. For bias vectors, This indicates element-wise multiplication, also known as element-wise squaring. Indicates will Squaring each element and then... The data is concatenated to form second-order features, expanding the dimension from 3 to 6. 10% of the data is randomly selected from each device group as the calibration set, and the remaining 90% is used as the test set. Mean squared error (MSE) is used as the loss function.
[0068] ;
[0069] in The number of samples in the calibration set is j, and j is the sample index, which ranges from 1 to N. The image feature vector after calibration for the j-th sample; This refers to the statistical characteristics part of the consensus benchmark features, namely... ,in The baseline grayscale mean, The baseline grayscale standard deviation, To establish the baseline boundary clarity, the weight matrix and bias vector are iteratively optimized using gradient descent until the loss function converges.
[0070] By using a hybrid calibration method combining histogram matching and multinomial regression, the differences in mean grayscale, standard deviation, and boundary sharpness of images from different brands and models of ultrasound equipment were significantly reduced, thus achieving effective unification of data distribution.
[0071] In one embodiment of the present invention, such as Figure 3 and 4 As shown, step S3 specifically includes the following steps:
[0072] S31. Perform a quality assessment on the calibrated ultrasound image data of the device. The quality assessment includes sharpness assessment, integrity assessment and artifact-free assessment to obtain a comprehensive quality score.
[0073] Specifically, the sharpness assessment obtains a sharpness score by calculating the average gradient magnitude of the image at multiple scales; the integrity assessment obtains an integrity score by calculating the proportion of the effective field of view area to the total area in the image; the artifact-free assessment obtains an artifact-free score by detecting the number and intensity of abnormal peaks in the frequency domain; and the overall quality score is a weighted sum of the sharpness score, integrity score, and artifact-free score.
[0074] Sharpness scoring is achieved through multi-scale gradient intensity evaluation. A standard deviation of is applied to each image. , , A Gaussian filter is applied to obtain a smoothed image. The gradient magnitudes of the image at each scale are then calculated.
[0075] ;
[0076] in Let the gradient magnitude be the value at the k-th scale. The image after applying a Gaussian filter with a standard deviation of 1. The standard deviation of the Gaussian filter. ; For the image in x Partial derivative in the direction (horizontal direction), For the image in y Partial derivative in the horizontal direction. Sharpness score is defined as the normalized average of the gradient magnitudes at three scales:
[0077] ;
[0078] in The sharpness score ranges from 0 to 1, with 1 indicating the sharpest sharpness. The average value of the gradient magnitude at the k-th scale; This represents the maximum value of the average gradient magnitude at this scale across all images.
[0079] Integrity scoring is achieved through visual field coverage assessment. The ultrasound image is divided into a 3x3 grid of 9 regions, and the grayscale entropy is calculated for each region.
[0080] ;
[0081] in grayscale The normalized histogram probability for this region. To avoid misclassifying dense artifact regions as valid regions, a grayscale range is added to assist in the judgment. If a region simultaneously satisfies the information entropy... (Threshold set to 4.0 bits) and the average grayscale value of the region If two conditions are met (typical grayscale range for effective ultrasound imaging), the region is considered to contain effective ultrasound information. The number of effective regions that simultaneously meet both conditions is counted. Integrity score is defined as The value ranges from 0 to 1, with 1 indicating complete coverage.
[0082] Artifact-free performance scoring is achieved through frequency domain anomaly peak detection. A two-dimensional discrete Fourier transform is performed on the calibrated image to obtain its frequency domain representation. Calculate the power spectrum ,in This is the frequency domain representation after the two-dimensional discrete Fourier transform. The horizontal coordinates are in the frequency domain. The vertical coordinates are in the frequency domain. Reverberation artifacts and acoustic shadowing in ultrasound images manifest as abnormal peaks at specific frequencies in the frequency domain. Gaussian smoothing of the power spectrum yields the baseline power spectrum. ,in Standard deviation Gaussian kernel, This represents the convolution operation. It calculates the power spectral deviation. Significant peak values are detected in the mid-to-high frequency region, and the peak value determination criterion is defined as follows. ,in This is the peak threshold coefficient. Indicates the standard deviation. Counts the number of significant peaks. Total frequency points ratio The artifact-free rating is defined as follows: , where 0.05 is the normalization factor.
[0083] An adaptive weighting mechanism is designed to calculate the overall quality score for different ultrasound application scenarios. For example, a balanced weighting method is used for abdominal ultrasound. For cardiac ultrasound, due to the high requirements for temporal resolution and clarity, the weight of the clarity score can be increased, such as by using a weighted average. For ultrasound of superficial organs, where the integrity of anatomical structures is crucial, the weight of the integrity score can be increased, for example, by using a weighted average. Set quality threshold ,when Data is considered acceptable at that time. The data was deemed substandard at that time.
[0084] S32. Divide the calibrated ultrasound image data of the device into regions, calculate the quality score of each region, identify ultrasound images with a comprehensive quality score lower than the quality threshold, and construct a quality gradient field based on the quality score difference between adjacent regions.
[0085] Specifically, for mildly poor-quality data with an overall quality score between 0.4 and 0.6, a partitioned adaptive repair process based on defect spatial localization is performed. The ultrasound images are divided into... Each region is a block, and the size of the block is adaptively adjusted according to the image resolution. For example, for a typical 512×512 pixel ultrasound image, the block size is set to 16×16 pixels, i.e. For high-resolution images, the block size is adjusted to 32×32 pixels to ensure that each block contains sufficient texture information for quality assessment.
[0086] Calculate the three-dimensional quality score independently for each region block. The calculation method is the same as the global quality score, but the evaluation scope is limited to this region. The quality score of the region is obtained by combining the three dimensions. (Taking abdominal ultrasound as an example.)
[0087] Construct a quality gradient field. The quality gradient field reflects the rate of change of quality scores in space and is used to identify boundary regions where quality changes drastically. The quality gradient field is defined as:
[0088] ;
[0089] in, Let f(m) be the mass gradient field of the region in the m-th row and n-th column, which is a two-dimensional vector. Let be the partial derivative of the quality score in the row direction; The partial derivative of the quality score in the column direction; This is the row index for the region block; This is the column index for the region block.
[0090] The partial derivatives are calculated using the central difference approximation:
[0091] ;
[0092] ;
[0093] in, The overall quality score for the region block in row m and column n. Assign a quality score to the adjacent next row block. The quality score is given to the adjacent block in the previous row. Assign a quality score to the adjacent next column of regions. The quality score is given to the adjacent block in the previous column.
[0094] For boundary blocks, forward or backward differencing is used instead of center differencing. The mass gradient magnitude is calculated as follows:
[0095] ;
[0096] The larger the mass gradient magnitude, the more drastic the mass change around the block.
[0097] In ultrasound images, regions with drastic changes in mass often correspond to important anatomical boundaries or areas of abrupt changes in imaging conditions. These areas require extra caution during restoration to avoid over-restoration that could damage boundary information. The mass gradient field extends the mass score from a scalar field to a vector field with spatial variation information, providing spatial guidance for subsequent adaptive adjustment of restoration intensity.
[0098] S33. Determine the defect type based on the quality score of each region block. The defect types include fuzzy defects, artifact defects, composite defects, and no defects.
[0099] Specifically, define the defect type determination rules: if and The block is determined to have a blur defect, meaning the image sharpness is insufficient but there are no obvious artifacts. If and This indicates that the block contains artifacts, meaning that artifact interference is severe but the image itself is relatively clear. If and The block was determined to have a composite defect, meaning it was both blurry and contained artifacts. If The block is determined to be an invalid region, meaning it lacks valid ultrasonic information. A spatial distribution map of the defects is generated, and the defect type of each region is labeled.
[0100] S34. Appropriate image inpainting algorithms are used to repair regions with different defect types, and the repair intensity of each region is dynamically adjusted according to the quality gradient field.
[0101] Specifically, for blurred defect areas, the BM3D algorithm is used for denoising. BM3D is a highly efficient image denoising algorithm that achieves denoising through block matching and collaborative filtering. The noise intensity parameter of BM3D is jointly adjusted based on the sharpness score and quality gradient field of the block, where the quality gradient field is used to identify boundary regions where the quality gradient magnitude is greater than a preset gradient threshold. The formula for calculating the repair intensity parameter for blurred defect areas is as follows:
[0102] ;
[0103] in, For the first Line number Repair strength parameters for column area blocks, Basic repair parameters, For clarity weighting coefficients, For the first Line number Clarity score for column area blocks. These are the gradient field weight coefficients. For the first Line number The magnitude of the quality gradient of the column region block. This is the gradient field sensitivity threshold.
[0104] A value of 15 is suitable for noise levels in general ultrasound images. This is the sharpness weighting coefficient, which is used when the sharpness score is... The lower, The larger the value, the stronger the noise reduction. The gradient field weight coefficient controls the effect of the quality gradient on the repair intensity. This is the gradient field sensitivity threshold. When the mass gradient magnitude exceeds this threshold, the repair intensity is significantly reduced. When the quality score difference between adjacent blocks exceeds 0.1, the region is considered to be on the boundary of quality change, and the repair intensity needs to be reduced.
[0105] When the mass gradient magnitude When I was very young, When the value is close to 1, the restoration strength is close to the level determined solely by the sharpness score. When the quality gradient magnitude is large, the exponential term tends to 0, the gradient field modulating term becomes ineffective, and the restoration strength is mainly determined by... This decision ensures that sufficient restoration force can still be applied to areas with severely insufficient sharpness. However, when the gradient magnitude is at a moderate level, the exponential decay term smoothly reduces the restoration intensity, avoiding over-restoration in boundary areas that could lead to blurred boundaries or distortion of tissue structure.
[0106] For areas with artifact defects, a frequency domain filtering algorithm is used for processing. The image is converted to the frequency domain using a two-dimensional Fourier transform, and then scored based on artifact-free performance. The frequency locations of anomalous peaks are determined, and band-stop filters are designed to suppress these frequency components. The image is then converted back to the spatial domain using an inverse Fourier transform. The filter's cutoff frequency and stopband width are adaptively adjusted according to the artifact type. For reverberation artifacts, low-frequency periodic components are suppressed; for acoustic shadowing artifacts, high-frequency components in specific directions are suppressed.
[0107] For regions with complex defects, frequency domain filtering and BM3D denoising are performed sequentially. First, frequency domain filtering removes artifact interference, and then the denoising algorithm improves sharpness. The intensity parameters of both steps are modulated by the mass gradient field.
[0108] For invalid regions, i.e. For regions that are not repaired, the original pixel values are preserved or marked as invalid regions to avoid introducing false textures in areas with no information.
[0109] This dynamic adjustment mechanism for restoration intensity achieves spatially adaptive restoration based on a quality gradient field. In internal regions with uniform quality, the restoration intensity is primarily determined by the sharpness score, fully restoring quality defects. In boundary regions with drastic quality changes, the restoration intensity is suppressed by a gradient field adjustment term, protecting boundary information from excessive smoothing. This mechanism improves overall image quality while preserving crucial diagnostic boundary features.
[0110] S35. Perform quality verification on each repaired region block, use a partition verification mechanism to judge the repair effect of each region block, retain the successfully repaired region blocks, restore the unrepaired region blocks to their original pixel values, and obtain the standardized ultrasound image dataset.
[0111] Specifically, the Canny operator is used to extract the edge images of the regions before and after restoration, and the number of edge pixels is calculated. The edge change rate is defined as:
[0112] ;
[0113] in, For the first Line number The number of edge pixels in the column region before patch repair. This represents the number of edge pixels after repair. Set the edge change rate threshold. This means that the number of edge pixels is allowed to vary by no more than 20%. If the repair is successful, the repaired pixel values are retained; otherwise... If the repair fails, the area will be restored to its original pixel values before the repair.
[0114] A partitioned verification mechanism is used to prevent the restoration algorithm from damaging the boundaries of anatomical structures. For areas where restoration fails, the reason for the failure and its location are recorded for subsequent manual review. The restoration success rate is statistically analyzed; when the overall image restoration success rate is below 70%, the image is deemed unsuitable for algorithmic restoration and marked as an image for manual review.
[0115] Furthermore, the sharpness score, integrity score, and artifact-free score of the restored image are recalculated to obtain the overall quality score of the restored image. Set the final quality threshold. ,filter The images formed a standardized ultrasound image dataset. Images that still fall below the threshold after restoration were marked as requiring manual review and were not included in the standardized dataset. Data pass rates were statistically analyzed for each center, equipment brand, and acquisition site to ensure a balanced distribution of the standardized dataset across different dimensions.
[0116] In one embodiment of the present invention, a standardized ultrasound image dataset is input into the AI ultrasound large model to be evaluated for testing, the diagnostic results output by the model are obtained, and the model performance index is calculated based on the clinical diagnostic report as the gold standard. The model's cross-center and cross-device generalization ability is evaluated by analyzing the coefficient of variation of the model performance index among different data acquisition centers.
[0117] Specifically, step S4 groups the standardized ultrasound image dataset according to dimensions such as data acquisition center, equipment brand, and acquisition site, forming multiple sub-test sets. Each sub-test set contains at least 100 data points to ensure statistical significance. Each sub-test set is sequentially input into the AI ultrasound model to be evaluated, and the model's output diagnostic results are recorded, including disease classification, lesion localization, and severity rating. Using the clinical diagnostic report as the gold standard, the core performance indicators of the model on each test set are calculated. These performance indicators include accuracy, sensitivity, and specificity. Accuracy is defined as the ratio of correctly diagnosed samples to the total number of samples. Sensitivity is defined as the ratio of true positive samples to the actual number of positive samples, reflecting the model's ability to detect diseases. Specificity is defined as the ratio of true negative samples to the actual number of negative samples, reflecting the model's ability to identify normal samples.
[0118] The coefficient of variation for each performance metric across different test sets is calculated using the following formula:
[0119] ;
[0120] in, Let be the standard deviation of a certain indicator across all test sets. This is the mean of the indicator. The coefficient of variation is denoted as . A smaller coefficient of variation indicates more stable model performance across data from different centers and on different devices, and stronger generalization ability. A generalization ability evaluation criterion is set: when the accuracy coefficient of variation... Sensitivity coefficient of variation Specificity coefficient of variation At that time, the decision model has good cross-center and cross-device generalization ability.
[0121] To verify the effectiveness of the standardization method of this invention, the same AI ultrasound large-scale model was tested on both unstandardized raw multicenter data and data standardized using the method of this invention. The coefficient of variation (COP) of performance indicators was calculated for both cases. In the raw data test, due to data distribution shifts caused by equipment differences and quality inconsistencies, the COP of model performance indicators was typically high; for example, the COP of accuracy reached 25%, indicating significant differences in model performance across different centers. In the standardized data test, after eliminating equipment differences and standardizing quality, the COP of accuracy dropped to below 10%, demonstrating that the standardization method of this invention effectively eliminated the interference of equipment and quality factors on the evaluation results, making the evaluation results more accurately reflect the model's cross-center and cross-equipment generalization ability.
[0122] A significant decrease in the coefficient of variation can distinguish whether the performance differences in models stem from insufficient generalization ability or from variations in the equipment and inconsistent quality of the test data. Models with a still high coefficient of variation on standardized data indicate that their generalization ability is indeed insufficient, requiring further optimization of the model architecture or expansion of the training data. Models with a significantly reduced coefficient of variation on standardized data indicate that their generalization ability is inherently good, and the performance fluctuations in the original data tests are mainly due to data factors. This provides a scientifically reliable testing basis for the clinical application of large-scale AI ultrasound models.
[0123] This invention achieves standardized processing of multi-center ultrasound image data and accurate evaluation of the generalization ability of AI ultrasound large-scale models through the above four steps. Step S1 ensures the diversity of data sources and the integrity of metadata, providing sufficient parameter basis for subsequent standardized processing. Step S2 eliminates imaging differences between different equipment brands and models based on dynamic benchmark selection and feature distribution calibration, achieving data distribution unification. Step S3 improves image quality through a quality gradient field-guided partitioned adaptive repair algorithm, repairing various quality defects while protecting critical diagnostic information. Step S4 accurately evaluates the model's cross-center and cross-device generalization ability through coefficient of variation analysis, effectively distinguishing the influence of model factors and data factors on test results, providing a scientific testing scheme for the clinical application and regulatory approval of AI ultrasound large-scale models.
[0124] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A data standardization processing method for AI ultrasound large model testing, characterized in that: Includes the following steps: S1. Acquire ultrasound image data from multiple external data acquisition centers and record the metadata of each acquisition center, including the ultrasound equipment brand and model, acquisition parameters, and acquisition environment information; S2. Based on the metadata, identify the differences in equipment imaging features of different data acquisition centers. Based on the similarity of image feature distribution of each data acquisition center, select at least one reference center from multiple data acquisition centers. Based on the image feature distribution of the reference center, establish a target reference distribution. Use a feature distribution calibration algorithm to calibrate the grayscale features and texture features of the ultrasound image data of non-reference centers to the target reference distribution, eliminate the imaging differences between different equipment brands and models, and obtain the ultrasound image data after equipment calibration. S3. Based on image clarity and integrity indicators, perform quality assessment on the ultrasound image data after the device is calibrated, identify ultrasound images with quality defects and their defect types, including image blurring, imaging artifacts and composite defects, and use corresponding image restoration strategies for different defect types to perform quality restoration processing to obtain a standardized ultrasound image dataset. S4. Use the standardized ultrasound image dataset for AI ultrasound large model testing; Step S3 specifically includes: S31. Perform a quality assessment on the calibrated ultrasound image data of the device. The quality assessment includes sharpness assessment, integrity assessment and artifact-free assessment to obtain a comprehensive quality score. S32. Divide the calibrated ultrasound image data of the device into regions, calculate the quality score of each region, identify ultrasound images with a comprehensive quality score lower than the quality threshold, and construct a quality gradient field based on the quality score difference between adjacent regions. S33. Determine the defect type based on the quality score of each region block. The defect types include fuzzy defects, artifact defects, composite defects, and no defects. S34. Appropriate image inpainting algorithms are used to repair regions with different defect types, and the repair intensity of each region is dynamically adjusted according to the mass gradient field. S35. Perform quality verification on each repaired region block, use a partition verification mechanism to judge the repair effect of each region block, retain the successfully repaired region blocks, restore the unrepaired region blocks to their original pixel values, and obtain the standardized ultrasound image dataset.
2. The data standardization processing method for AI ultrasound large model testing as described in claim 1, characterized in that: In step S1, the ultrasound equipment brand and model include the equipment brand and model number; the acquisition parameters include probe frequency, gain value, dynamic range and imaging mode; and the acquisition environment information includes the operator's professional title, acquisition site, acquisition time, and patient's age and gender.
3. The data standardization processing method for AI ultrasound large model testing as described in claim 2, characterized in that: Step S2 specifically includes: S21. Extract features from the ultrasound images of each data acquisition center to obtain the image feature distribution of each data acquisition center. The image feature distribution includes gray-level statistical distribution, texture feature distribution, and frequency domain feature distribution. S22. Calculate the image feature distribution distance between each data acquisition center, construct the distribution distance matrix between the centers, calculate the coverage index based on the distribution distance matrix, and select at least one reference center from multiple data acquisition centers according to the coverage index. The coverage index represents the average of the minimum distribution distances of all data acquisition centers to the reference center combination. S23. Establish a target reference distribution based on the image feature distribution of the reference center; S24. The ultrasound image data of the non-reference center is calibrated to the target reference distribution using a feature distribution calibration algorithm, wherein the feature distribution calibration algorithm includes grayscale distribution alignment and texture feature mapping, to obtain the ultrasound image data after calibration by the device.
4. The data standardization processing method for AI ultrasound large model testing as described in claim 3, characterized in that: In step S21, the gray-level statistical distribution is obtained by calculating the gray-level mean, gray-level standard deviation, skewness, and kurtosis of the image; the texture feature distribution is obtained by extracting the contrast, correlation, energy, and homogeneity of the image through the gray-level co-occurrence matrix; and the frequency domain feature distribution is obtained by converting the image to the frequency domain space through Fourier transform and extracting spectral features.
5. The data standardization processing method for AI ultrasound large model testing as described in claim 3, characterized in that: In step S22, the Wasserstein distance is used to calculate the distance between image feature distributions of each data acquisition center. The Wasserstein distance measures the transmission cost between the image feature distributions of two centers. The candidate reference center combination with the smallest coverage index is selected as the final reference center. The formula for calculating the coverage index is as follows: ; in, The number of all data acquisition centers. For candidate benchmark center combinations, Let F be the image feature distribution of the k-th data acquisition center. j The image feature distribution of the j-th center in the reference center combination. Let the Wasserstein distance be the distance between the k-th center and the j-th center. This represents the minimum Wasserstein distance from the k-th center to all centers in the combination of reference centers.
6. The data standardization processing method for AI ultrasound large model testing as described in claim 3, characterized in that: In step S24, the grayscale distribution alignment adopts the histogram matching method. The ultrasound images of non-reference center are grouped according to the equipment brand. The grayscale histogram distribution of each equipment brand group is extracted, the cumulative distribution function of the grayscale histogram distribution is calculated, and the cumulative distribution function is mapped to the cumulative distribution function corresponding to the target reference distribution to establish the grayscale value mapping relationship and realize the grayscale distribution alignment of all images in the same equipment brand group. The texture feature mapping adopts a multinomial regression method to establish a regression model between device parameters and texture features. The regression model is used to eliminate residual feature differences caused by device brand and model.
7. The data standardization processing method for AI ultrasound large model testing as described in claim 1, characterized in that: In step S31, the sharpness assessment obtains a sharpness score by calculating the average gradient magnitude of the image at multiple scales; the integrity assessment obtains an integrity score by calculating the proportion of the effective field of view area to the total area in the image; the artifact-free assessment obtains an artifact-free score by detecting the number and intensity of abnormal peaks in the frequency domain; and the comprehensive quality score is a weighted sum of the sharpness score, the integrity score, and the artifact-free score.
8. The data standardization processing method for AI ultrasound large model testing as described in claim 1, characterized in that: In step S34, a denoising algorithm is used to process the blurred defect region, a filtering algorithm is used to process the artifact defect region, and filtering and denoising are performed sequentially on the composite defect region. The mass gradient field is used to identify boundary regions where the mass gradient magnitude is greater than a preset gradient threshold. The repair intensity is reduced in the boundary regions, and a standard repair intensity is used in the internal regions with uniform mass. The formula for calculating the repair intensity parameter of the blurred defect region is as follows: ; in, The repair strength parameter for the region block in row m and column n is... Based on the repair parameters, For clarity weighting coefficients, Score the sharpness of the region block in row m and column n. These are the gradient field weight coefficients. The mass gradient magnitude of the region in the m-th row and n-th column. This is the gradient field sensitivity threshold.
9. The data standardization processing method for AI ultrasound large model testing as described in claim 1, characterized in that: In step S35, the partition verification mechanism includes: Calculate the edge intensity of each region before and after restoration, and obtain the edge intensity change rate. The edge intensity change rate is defined as the ratio of the difference in edge intensity before and after restoration to the edge intensity before restoration. For regions where the edge intensity change rate exceeds a preset threshold, the restoration of the region is determined to be unsuccessful, and the region is restored to its original pixel value before restoration. For regions where the edge intensity change rate does not exceed the preset threshold, the restoration of the region is determined to be successful, and the restored pixel value is retained. After all regions are verified, image fusion is performed.