Medical image artifact recognition and elimination method based on big data technology
By employing a medical image artifact recognition and removal method based on big data technology, and utilizing techniques such as grayscale normalization and semantic segmentation, the problem of artifact recognition and removal in PACS systems has been solved, improving the accuracy and consistency of medical image processing and supporting more efficient lesion detection and segmentation.
Patent Information
- Application Number
- CN202511439472.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-10
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2045-10-10
AI Technical Summary
In existing technologies, PACS systems have failed to effectively address the in-depth mining and predictive analysis of unstructured data in image annotation, dataset classification, and AI model training. Furthermore, they lack insights into the construction of structured feature libraries and the closed loop of disease prediction, and cannot effectively identify and remove artifacts in medical images.
A big data-based approach is adopted, which involves grayscale normalization, semantic segmentation, ROI localization, lesion segmentation and model validation. Combined with DICOM gateway, N4 bias field correction and adaptive window width and level adjustment, convolutional neural network is used for image preprocessing. Artifact recognition and removal are performed through ROI localization model and lesion segmentation model, and distributed deployment is achieved.
It effectively reduces the false alarm rate of artifacts, improves the accuracy of organ segmentation, reduces equipment dependence differences, enhances the accuracy of lesion detection and segmentation adaptability, assists in surgical planning, and reduces boundary positioning errors.
Smart Images

Figure CN120894459A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of medical image processing, in particular to a medical image artifact identification and elimination method based on big data technology. BACKGROUND
[0002] In image diagnostics, PACS is a computer system specially used for storing, acquiring, sending and displaying medical images. PACS technology is proposed to realize medical image non-filmization storage and exchange, improve overall medical work efficiency, realize network, modernization and remote of doctors' work, and is a general solution in the field of medical image storage and exchange.
[0003] In the prior art, the patent discloses an AI-based PACS system, which realizes image labeling, data set classification and AI model training, but does not solve the deep mining and prediction analysis of unstructured data, and lacks inspiration for structured feature library construction and disease prediction closed loop; Therefore, a new medical image artifact identification and elimination method based on big data technology is urgently needed. SUMMARY
[0004] The present application provides a medical image artifact identification and elimination method based on big data technology, which solves the problems of the prior art.
[0005] In a first aspect, the present application provides a medical image artifact identification and elimination method based on big data technology, comprising: The collected image data is preprocessed based on gray scale normalization, and then semantic segmentation and ROI positioning are performed on the image content. The segmented image content is executed, the lesion feature is quantized and fused, and the model verification is executed, and the verified lesion segmentation model is distributed.
[0006] Further, in the preprocessing and semantic segmentation steps, the steps specifically include: The field strength inhomogeneity of MRI is corrected by a semantic segmentation model based on a convolutional neural network, and at the same time, the contrast optimization suitable for all modalities is dynamically set and adjusted according to the image modalities, the structure of the image content is verified for integrity, and the low confidence area marked by the segmentation mask is removed.
[0007] Further, the step of outputting the standardized image after preprocessing specifically further includes: Step d1, access the DICOM gateway data to obtain the original image data, and correct the N4 bias field, including: modeling the image as a raw image, a true signal, a bias field, and noise; iteratively estimating the bias field and compensating by a maximum likelihood expectation algorithm; Step d2, adjusting the window width and window level, adjusting the window level by adjusting the target tissue gray center value, adjusting the window width by adjusting the gray scale display range, automatically matching the preset window value through the DICOM label, and when the DICOM label is not marked, using adaptive histogram equalization instead of fixed window value; Step d3, loading into the encoder-decoder to perform image content segmentation, and generating image content binary mask; Step d4, performing confidence detection on the binary mask, outputting image content ROI through the verified loaded ROI positioning, and marking non-artifact data if not verified; Step d5, performing rule rejection on the artifact data; Rule 1: integrity verification, calculating the artifact volume, and if the artifact volume is less than 30% of the reference corresponding organ volume, marking as incomplete scanning and rejecting; Rule 2: morphological verification, extracting the artifact contour, calculating the circularity, and if the circularity is less than the threshold, marking as non-motion artifact and rejecting.
[0008] Further, the image data collected is preprocessed based on gray scale normalization, and then semantic segmentation and ROI positioning are performed on the image content, and the steps of ROI positioning include: processing the data after semantic segmentation into a preset voxel standardized image, loading the ROI positioning model, and outputting the coordinates of the target bounding box, the confidence and the target type label.
[0009] Further, the ROI positioning model specifically includes: The loss function of the ROI positioning model: wherein λ coord is a weight coefficient for controlling the coordinate loss of sub-millimeter level positioning required in diagnosis, λ obj is a weight coefficient for controlling the target existence loss and suppressing false negatives, L obj is a focus loss function, L coord is a complete intersection ratio loss function, λ cls is a weight coefficient for controlling the classification loss, L cls is a cross-entropy loss.
[0010] Further, the loss function of the ROI positioning model specifically includes: The focus loss function formula is: wherein, p t is the confidence of the model predicting the existence of the target, α t is the class weight, γ is the focus parameter; The formula of the complete IoU loss function is: wherein, IoU is the IoU of the prediction box and the real box, ρ 2 ( b pred , b gt ) is the squared Euclidean distance of the center points of the prediction box and the real box, b pred is the prediction box, b gt is the real box, c is the diagonal length of the minimum bounding box, v is the aspect ratio consistency factor; is the dynamic weight coefficient, = v / ((1- IoU )+ v ); When IoU is small, it means that the overlap is very low, at this time more attention to the overlap problem rather than the aspect ratio, will be automatically reduced to weaken the influence of v ; v The formula for constraining the shape of the bounding box to conform to the lesion shape is: w gt , w pred , h gt , h pred respectively represent: real width, predicted width, real height, predicted height; Classification loss function L cls The formula is: wherein, M is the total number of classes in the classification task, c is the class index, w c , y c ,p c Class weight, true label and prediction probability, respectively.
[0011] Further, the positioned image content is executed lesion segmentation, lesion feature is selected and quantified, fusion is performed, and a model is verified, and the verified lesion segmentation model is distributedly deployed, and specifically includes: Based on the lesion segmentation model adaptively built by nnUNet, the lesion ROI data is loaded, 3D convolution downsampling is performed, transposed convolution upsampling is performed, and a Voxel level mask is output; Radiomics features are extracted from the PyRadiomics library, including GLCM entropy and GLSZM region size variance, deep learning features are extracted from the lesion segmentation model, the deep learning features include high-dimensional semantic features, specifically, the last but one layer global average pooling is extracted, the radiomics features and the deep learning features are executed feature fusion, and complementary verification based on the malignant and benign discrimination AUC is performed.
[0012] Further, the model verification method specifically includes: Five-fold cross-validation is performed through stratified sampling, the average Dice coefficient is calculated, and the Dice coefficient is greater than a threshold Dice, which is qualified, and the Dice coefficient is the overall segmentation coincidence degree, that is, the segmentation accuracy; Through robustness testing, motion artifacts are simulated under Gaussian noise, and the performance decline ratio is less than 5%, which is qualified; Through generalization testing, the model's malignant and benign discrimination AUC is verified, and the model is qualified when the AUC decreases by less than 0.03.
[0013] Further, it also includes performing data drift detection on the lesion segmentation model, and when the data drift value is greater than 0.1, the lesion segmentation model is retrained; When the lesion segmentation model accuracy is less than 0.85 for thirty consecutive days, the lesion segmentation model is retrained.
[0014] In a second aspect, the present application provides a medical image artifact recognition and elimination system based on big data technology, which is used to execute the method of any one of the first aspect, and includes a gateway and a server, the gateway accesses the server, the server pre-processes MRI image data through a semantic segmentation model, an ROI positioning model outputs image content positioning, a lesion segmentation model loads ROI data to output a mask, the server extracts deep learning features from the lesion segmentation model, fuses radiomics features, and performs complementary verification based on the malignant and benign discrimination AUC, and outputs voxel-level lesion segmentation results based on the fused features, including: the prediction confidence of each voxel, the lesion volume, the surface area, and the centroid coordinates.
[0015] The application provides a medical image artifact identification and elimination method based on a big data technology, solves the problem of inconsistent gray scale caused by medical imaging device differences through bias field correction+window width and window level self-adaption, reduces the gray scale standard deviation, and eliminates device dependence differences. Through image content segmentation+morphological verification, the problem of motion artifact interference and organ boundary blurring is solved, the Dice coefficient of organ segmentation is improved, and the artifact misjudgment rate is reduced. Through a lesion detection model, the problems of high micro-nodule missed detection rate and high risk of missed diagnosis of malignant lesions are solved, the false positive rate is reduced, and the malignant lesion recall rate is improved. Through a lesion segmentation model, the segmentation adaptability to different sizes of lesions is improved, clear segmentation boundaries are obtained, surgical planning is assisted, and boundary positioning errors are reduced. BRIEF DESCRIPTION OF DRAWINGS
[0016] The drawings described herein are used to provide further understanding of the embodiments of the application, constitute a part of the application, and do not constitute a limitation on the embodiments of the application. In the drawings: Figure 1 A medical image artifact identification and elimination method based on a big data technology is provided for an exemplary embodiment of the application.
[0017] Figure 2 A process diagram in which the positioning output after data preprocessing is a standardized image in a medical image artifact identification and elimination method based on a big data technology provided for an exemplary embodiment of the application. DETAILED DESCRIPTION
[0018] Here, the exemplary embodiments will be described in detail, and examples are shown in the drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the application.
[0019] First, the terms involved in the application are explained: terminology definitions PACS system Picture Archiving and Communication System The specific application scenario of the application is the field of medical image processing technology.
[0020] The medical image artifact identification and elimination method based on a big data technology provided by the application aims to solve the above technical problems of the prior art.
[0021] The technical solutions of the application and how the technical solutions of the application solve the above technical problems will be described in detail in the following specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of the application will be described with reference to the drawings.
[0022] Embodiment 1: The embodiment provides a medical image artifact identification and elimination method based on big data technology, which is realized through the following steps, including steps D1, D2, as shown in Figure 1 Step D1: Obtain the image through the DICOM gateway, use N4 bias field correction + window width adjustment to realize gray scale normalization, and then perform semantic segmentation and ROI positioning; Here, the semantic segmentation relies on the normalization result of the previous step, because the uncorrected image will cause the organ boundary to be blurred, and the image without normalization will cause the segmentation model to fail, which specifically includes: the uneven gray scale causes the organ boundary to be blurred, and the segmentation accuracy is reduced; As shown in Figure 2 Steps d1-d5 are the process of positioning the output of the standardized image after data preprocessing, in detail: Among them, the data preprocessing process before semantic segmentation includes steps d1-d2: Step d1, model the image as an original image, a true signal, a bias field and noise, estimate the bias field by maximum expectation algorithm iteration and compensation, parameter setting: iteration number = 50, convergence threshold = 0.001; Step d2, based on the cv2.normalize operation of OpenCV, adjust the window width and window level, adjust the window level by adjusting the gray scale center value of the target tissue, adjust the window width by adjusting the gray scale display range, the dynamic adjustment strategy is: automatically match the preset window value through the DICOM label, when the DICOM label is not marked, use adaptive histogram equalization instead of fixed window value; After preprocessing, semantic segmentation is performed, that is, image content segmentation, the semantic segmentation model based on convolutional neural network used in the embodiment is U-NetDice, if the image is not normalized, the segmentation model will fail, for example, when the coefficient of U-NetDice is ≤0.7, due to uneven gray scale, the organ boundary is blurred, and the U-Net segmentation accuracy is reduced, at the same time, the non-normalization also causes the artifact to remain, which affects the identification result, for example, the motion artifact is judged as a lung nodule; The technical problem solved by normalization: the normalization process solves the inconsistency of the image gray scale range output by different types of CT / MRI equipment, such as GE, Siemens, while avoiding interference from scanning parameters. Changes in scanning parameters such as kVp, slice thickness, and other parameters can cause the gray scale value of the same tissue to drift. In practice, it is found that the bone CT value fluctuates between 2000 and 3000 HU. The bias field effect specific to MRI can cause the edge gray scale of the image to be distorted, with a maximum signal deviation of 40%. The de-noising operation, which is a complete integrity check of the structure of the image content, has good results in solving noise and artifacts. The device noise influence is removed for subsequent operations, more attention is paid to biological characteristics, and the data fidelity is improved.
[0023] Therefore, the gray scale normalization and the semantic segmentation of the organ form a cascading preprocessing relationship. The organ in the image content is segmented and artifacts are removed in the embodiment, including the following steps d3-d5: d3, load into the encoder-decoder to perform image content segmentation, and generate an image content binary mask; Input specification: 128x128x64 voxels (spatial resolution 1mm3), output result: generate organ binary mask (such as liver Mask=1, background=0); d4, perform confidence detection on the binary mask, and output the image content ROI through the verified ROI positioning, and mark the non-artifact data that does not pass the verification; Step d5, perform rule removal on the artifact data; Rule 1: integrity verification, calculate the artifact volume, if the artifact volume is less than 30% of the reference organ volume, mark it as incomplete scanning and remove it; here, the organ structure integrity check, that is, the structure of the image content is checked for integrity. In this embodiment, after the normalization process, the Dice coefficient of the segmentation model is improved by 0.15; Rule 2: morphological verification, extract the artifact contour, calculate the circularity, and if the circularity is less than 0.7, mark it as a non-motion artifact and remove it. The circularity of a normal organ is greater than 0.8.
[0024] d1-d5 cascading verification data based on 1000 CT test results of the LIDC-IDRI dataset: The ROI positioning model, i.e., the lesion detection model, has the following training process and use: The training process of the lesion detection model is as follows: The loss function of the lesion detection model is constructed as: Where λ coordis a weight coefficient (default value = 5) for controlling the coordinate loss requiring sub-millimeter level positioning in diagnosis; λ obj is a weight coefficient (default value = 1) representing the control target existence loss, inhibiting false negatives; L obj is a focus loss function for dynamically adjusting the difficulty sample weight; L coord is a complete intersection-over-union loss function, introducing a center point distance and aspect ratio penalty term, representing the complete intersection-over-union loss; λ cls is a weight coefficient (default value = 1) representing the control classification loss, used to improve the lesion type identification, for example, to distinguish between benign and malignant, such as ground glass nodule vs. solid nodule; L cls is a cross-entropy loss, representing the difference between the predicted label and the true label, in this embodiment, the recorded multi-lesion type identification is added, which can simultaneously detect nodules, calcifications, masses, etc.; wherein the formula of the complete intersection-over-union loss function is: wherein, IoU is the intersection-over-union of the predicted box and the true box, used to ensure the basic positioning accuracy; ρ 2 ( b pred , b gt ) is the squared Euclidean distance of the center points of the predicted box and the true box, b pred is the predicted box, b gt is the true box, used to ensure the accuracy of the lesion center positioning; c is the length of the diagonal of the minimum bounding box, v is an aspect ratio consistency factor, used to constrain the shape of the bounding box to conform to the lesion shape, and the formula is: w gt , w pred , h gt , h pred respectively represent: true width, predicted width, true height, predicted height; αis a weight coefficient, used for weakening the influence of IoU when the loss function is low v ; wherein the focus loss function formula is: wherein, p t is the confidence of the model predicting the existence of the target; α t is the class weight, in this embodiment, the lesion class α t is 0.8, the background class α t is 0.2, and the class weight is adaptively increased when the lesion area is extremely small; γ is the focus parameter, and the default value is 2; when the sample is a medium difficulty sample, the value is adjusted to 1.5 to avoid excessive suppression of sample data; classification loss function L cls formula is: M is the total number of classes in the classification task, such as lung nodule diagnosis M is 3, indicating malignant / benign / uncertain, c is the class index, w c , y c , p c are the class weight, the real label and the predicted probability respectively, wherein the real label adopts one-hot encoding; The lesion detection model achieves high positioning accuracy and ensures accurate positioning of the lesion center.
[0025] wherein the training parameter optimizer is AdamW, and the batch size is 32 cases / GPU; The use process of the lesion detection model is as follows: In this embodiment, YOLOv5 is used to process the input data into 128x128x64 voxel standardized images, load the ROI positioning model, and output the coordinates, confidence and target type label of the target bounding box. The target bounding box here is the lesion bounding box, which solves the problem of long time and low efficiency of manual marking of single CT by traditional radiologists; the lung nodule detection rate reaches 98.7% based on the LIDC-IDRI benchmark, and the false positive rate is <0.5 cases / scan; Step D2: After lesion detection, the lesion is segmented, the lesion segmentation model is loaded, the lesion feature is selected and quantified, fused, and the model verification is performed, and the verified lesion segmentation model is distributed.
[0026] The lesion segmentation model is built based on nnUNet self-adaptation, the input is the lesion ROI region output by YOLOv5, that is, 64x64x64 voxels, and the output is a binary mask of the lesion region with Voxel level precision, which solves the problem of fuzzy segmentation of irregular lesion boundaries. The loss function of the lesion segmentation model is a hybrid loss composed of Dice loss and Hausdorff loss, and the loss function of the lesion segmentation model is L total The formula is: L total = α ⋅ L Dice + β ⋅ L Hausdorff The sum of the weight coefficients α,β is 1, usually set to α =0.6, β =0.4, and adjusted according to the training stage; L Dice The Dice loss value is the Dice loss value. The closer the value is to 0, the higher the coincidence degree of the predicted segmentation result and the true result (gold standard), and the better the segmentation effect, and the formula is: Wherein, N is the total number of three-dimensional voxels, that is, the number of voxels of a single input image, p i is the model prediction probability, that is, the possibility of the voxel being judged as a lesion, g i is the gold standard mask, which is 0 or 1, and is the true lesion area marked by a doctor; L Dice ∈[0,1] (0=perfect match, 1=complete mismatch), when the prediction and the gold standard are completely coincident L Dice is 0, when the prediction covers 50% of the gold standard, L Dice The calculation of is 1-2*0.5 / 1+0.5=0.33; L Hausdorff The Hausdorff distance-based loss value is the Hausdorff distance-based loss value. The smaller the value, the smaller the maximum distance between the two outlines, that is, the better the boundary matching, and the formula is: where, ∂ G is the surface contour point of the ground truth lesion extracted by the binary mask, ∂ P is the surface contour point of the prediction result; from ∂ G find any point as x , for the point x, find the nearest point ∂P on the predicted contour x , then traverse all the points on the real contour y , points x , and x and y are not predefined, but Hausdorff are the key point pairs dynamically determined in the distance calculation process to find the maximum mismatch between the two contours , 95perc is the 95th percentile of the distance, and ||x-y|| is the Euclidean distance; In the loss function, the traditional Dice loss may give a high score but the boundary is rough, while the Hausdorff loss is specifically optimized for sub-millimeter edge accuracy.
[0027] In feature fusion, the feature fusion matrix has a dimension of (128, 1536), and 1536 is the concatenation of 512 dimensions from radiomics and 1024 dimensions from deep learning. The last but one layer of the lesion segmentation model extracts a dimension of 1024; from the model structure, the last layer of the segmentation network is usually a 1x1 convolution, which outputs the same number of channels as the number of classes, such as 2 channels for lesion / background binary classification. The output of this layer is the classification probability of each pixel, which belongs to highly task-specific information. Here, the last but one layer is used, which retains more abundant abstract features and has not yet been constrained by the final task, so it has stronger generalization ability. From the medical demand level, lesion features need to balance generality and specificity. A 1024-dimensional embedding vector can express both the general pattern of "malignant tumor" and the unique features of specific cases (such as calcification distribution), which is the value of high-dimensional semantic embedding. Regarding medical interpretation of high-dimensional semantics, for example: in an embedding dimension of 128, the activation pattern is edge high activation, and the associated pathological feature is lesion lobulation; in an embedding dimension of 384, the activation pattern is center-edge gradient activation, and the associated pathological feature is ground glass nodule sign; in an embedding dimension of 768, the activation pattern is point cluster high activation, and the associated pathological feature is microcalcification cluster.
[0028] Regarding model verification, five-fold cross-validation is performed through stratified sampling, and an average Dice coefficient greater than 0.90 is considered qualified. The Dice coefficient is the overall segmentation overlap, i.e., segmentation accuracy. Pass the robustness test, simulate the motion artifact under Gaussian noise, and calculate the performance reduction ratio less than 5% to pass; Pass the generalization test, verify the benign and malignant discrimination AUC of the model, and the decrease is less than 0.03 to pass.
[0029] Finally, the lesion segmentation model is distributedly deployed on multiple servers to expand the data acquisition range and enhance the model effect.
[0030] In the embodiments of the present application, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiments described above are only illustrative, The modules described as separate components can or can not be physically separated, and the components shown as modules can or can not be physical modules, i.e., they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the embodiment.
[0031] In addition, each functional module in the embodiments of the present application can be integrated into one processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The integrated module can be realized in the form of hardware or in the form of hardware plus software functional module.
[0032] Those skilled in the art should understand that the embodiments of the present application can be provided as a method or a system. Therefore, the present application can be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects.
[0033] It should be further noted that the terms "comprise", "include" or any other variant thereof are intended to cover non-exclusive inclusion, so that processes, methods, articles or devices including a series of elements do not only include those elements, but also include other elements not explicitly listed, or further include elements inherent to such processes, methods, articles or devices. Without more limitations, the element defined by the statement "comprises a" does not exclude the presence of another identical element in the process, method, article or device including the element.
[0034] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application can have various modifications and changes. Any modification, equivalent replacement, improvement, etc. within the spirit and principles of the present application shall be included in the scope of claims of the present application.
[0035] Other embodiments of the application will be apparent to those skilled in the art from consideration of the specification and practice of the application disclosed herein. It is intended that the specification and examples be considered as exemplary only, with a true scope and spirit of the application being indicated by the following claims.
[0036] It is to be understood that the application is not limited to the precise construction hereinafter described and as shown in the attached drawings, and that various changes in the details and arrangements of parts can be made by those skilled in the art without departing from the scope of the application. The scope of the application is limited only by the claims appended hereto.
Claims
1. A method for medical image artifact recognition and removal based on big data technology, characterized in that, include: The acquired image data is preprocessed based on grayscale normalization, and then semantic segmentation and ROI localization are performed on the image content. The system performs lesion segmentation on the located image content, selects lesion features for quantification and fusion, performs model verification, and then deploys the verified lesion segmentation model in a distributed manner.
2. The method for medical image artifact recognition and removal based on big data technology according to claim 1, characterized in that, The preprocessing and semantic segmentation steps in the process of preprocessing the acquired image data based on grayscale normalization and then performing semantic segmentation and ROI localization on the image content specifically include: The field intensity inhomogeneity of MRI is corrected by a semantic segmentation model based on convolutional neural networks. At the same time, the contrast optimization applicable to all modalities is dynamically set and adjusted according to the image modality. The structural integrity of the image content is verified and low-confidence regions marked by segmentation masks are removed.
3. The method for medical image artifact recognition and removal based on big data technology according to claim 2, characterized in that, The steps for outputting standardized images after preprocessing also include: Step d1: Obtain raw image data from the DICOM gateway and perform N4 offset field correction, including: The image is modeled as the original image, the real signal, the bias field, and the noise. The bias field is estimated and compensated iteratively using the expectation-maximization algorithm. Step d2: Adjust the window width and window position. Adjust the window position by adjusting the gray center value of the target organization, and adjust the window width by adjusting the gray display range. Automatically match the preset window value of C through the DICOM label. When the DICOM label is not marked, use adaptive histogram equalization to replace the fixed window value. Step d3: Load the image content into the encoder-decoder to perform image content segmentation and generate a binary mask for the image content; Step d4: Perform confidence testing on the binary mask. Load the ROI that passes the test, locate the ROI of the output image content, and mark the data that does not pass the test as artifact data. Step d5: Perform rule-based removal on the artifact data; Rule 1: Integrity verification. Calculate the artifact volume. If the artifact volume is less than 30% of the corresponding reference organ volume, mark it as incomplete scan and remove it. Rule 2: Morphological verification, extract artifact contours, calculate roundness, if the roundness is less than the threshold, mark the unmoving artifacts and remove them.
4. The method for medical image artifact recognition and removal based on big data technology according to claim 2, characterized in that, The process of preprocessing the acquired image data based on grayscale normalization, followed by semantic segmentation and ROI localization of the image content, includes the following steps for ROI localization: The semantically segmented data is processed into a pre-defined voxel-normalized image, loaded into the ROI localization model, and outputs the coordinates, confidence score, and target type label of the target bounding box.
5. The method for medical image artifact recognition and removal based on big data technology according to claim 4, characterized in that, The ROI location model specifically includes: Loss function for ROI localization model: Where, λ coord λ is a weighting coefficient used to control coordinate loss in sub-millimeter positioning required for diagnostics. obj The weighting coefficients are used to control the loss of target existence and suppress false negatives. L obj To focus on the loss function, L coord This represents the coordinate localization loss, used to measure the overlap between the predicted lesion bounding box and the actual bounding box, as well as the consistency of the center point and aspect ratio. L coord Replace it with the complete intersection-union loss function, λ cls To control the weighting coefficients of the classification loss, L cls This represents the cross-entropy loss.
6. The method for medical image artifact recognition and removal based on big data technology according to claim 5, characterized in that, The loss function of the ROI localization model specifically includes: The formula for the focusing loss function is: in, p t The confidence level for the model to predict the existence of the target. α t For category weights, γ For focusing parameters; The formula for the complete intersection-union ratio loss function is: in, IoU The intersection-union ratio (IU) of the predicted bounding box and the ground truth bounding box. ρ 2 ( b pred , b gt The squared Euclidean distance between the center points of the predicted bounding box and the ground truth bounding box is denoted as . b pred For the prediction box, b gt For the true frame, c The minimum bounding box diagonal length. v The aspect ratio consistency factor; For dynamic weighting coefficients, = v / (1- IoU )+ v ); when IoU When the value is small, it indicates a low degree of overlap. In this case, the focus is more on overlap rather than aspect ratio. It will automatically decrease to weaken v Impact; v The formula used to constrain the bounding box shape to conform to the lesion shape is: w gt , w pred , h gt , h pred These represent: actual width, predicted width, actual height, and predicted height, respectively. Classification loss function L cls The formula is: in, M This represents the total number of categories in the classification task. c For category indexing, w c , y c , p c These are the class weights, true labels, and predicted probabilities, respectively.
7. The method for medical image artifact recognition and removal based on big data technology according to claim 6, characterized in that, The process of performing lesion segmentation on the located image content, selecting and quantifying lesion features, fusing them, and performing model validation, followed by distributed deployment of the validated lesion segmentation model, specifically includes: The lesion segmentation model based on nnUNet adaptive construction loads lesion ROI data, performs 3D convolution downsampling, then transposes convolution upsampling, and outputs a Voxel-level mask. Extracting radiomics features from the PyRadiomics library includes: GLCM entropy, GLSZM regional size variance. Extracting deep learning features from the lesion segmentation model, where the deep learning features include high-dimensional semantic features, specifically extracted from the global average pooling of the penultimate layer. Performing feature fusion on the radiomics features and deep learning features, and conducting complementary verification based on the AUC of benign and malignant discrimination.
8. The method for medical image artifact recognition and removal based on big data technology according to claim 7, characterized in that, Regarding the methods of model validation, specifically including: Performing five-fold cross-validation through stratified sampling. Calculating that the average Dice coefficient is greater than the threshold Dice is considered qualified. The Dice coefficient is the overall segmentation coincidence degree, that is, the segmentation accuracy. Through robustness testing, simulating motion artifacts under Gaussian noise. Calculating that the proportion of performance degradation is less than 5% is considered qualified. Through generalization testing, validating that the decrease in the AUC of benign and malignant discrimination of the model on the external dataset is less than 0.03 is considered qualified.
9. The method for medical image artifact recognition and removal based on big data technology according to claim 8, characterized in that, It also includes performing data drift detection on the lesion segmentation model. When the data drift value is greater than 0.1, triggering retraining of the lesion segmentation model. When the accuracy of the lesion segmentation model is less than 0.85 for 30 consecutive days, triggering retraining of the lesion segmentation model.
Citation Information
Patent Citations
Artificial Intelligence-Based PACS System and Its Design Methodology
CN113053494B
AI-based medical image automatic segmentation and labeling system
CN119418061A
MRI medical image correction method and system based on convolutional neural network, and computer readable storage medium
CN119963681A
Method and system for processing intracranial large vessel image, electronic device, and medium
WO2025180094A1
Cited By
Medical staff-oriented equipment parameter configuration auxiliary method and system
CN121687434A
A medical staff-oriented device parameter configuration assisting method and system
CN121687434B