Image artifact detection method and system based on deep learning technology

Through the image artifact detection method of deep learning technology, the feature extraction, position coding and dynamic anchor box scaling mechanisms are used to solve the problem of insufficient accuracy and generalization ability of artifact detection in medical images, and the automated detection and quantitative evaluation of artifacts are realized, which improves the accuracy and efficiency of detection.

CN120298864APending Publication Date: 2025-07-11SHAN DONG MSUN HEALTH TECH GRP CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510385833.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-29
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The existing artifact detection methods in medical images have insufficient positioning accuracy under small-scale artifact recognition and complex backgrounds, high detection missed detection rate, large position deviation, and weak generalization ability, which affects the quality of artifact detection in clinical images.

Method used

The image artifact detection method based on deep learning technology is adopted, and the model space perception ability and anchor frame matching accuracy are improved through feature extraction, position encoding, encoder-decoder structure and dynamic anchor frame scaling mechanism, and feature modeling and context understanding capabilities are enhanced.

Benefits of technology

The accuracy and efficiency of artifact detection in medical images are improved, especially under small scales and complex texture artifact types, the problem of insufficient accuracy and generalization capabilities of existing methods is overcome, and the automatic detection and quantitative evaluation of artifacts are realized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298864A_ABST
    Figure CN120298864A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image processing, and provides an image artifact detection method and system based on a deep learning technology, and the method comprises the following steps: carrying out the preprocessing of an obtained to-be-processed image, and carrying out the feature extraction of the preprocessed image; position coding is carried out on the extracted feature map A, and the extracted feature map A is fused with the feature map A; performing multi-layer encoding and decoding on the fused feature map through an encoder and a decoder to obtain a feature map B containing an artifact target area; and generating an anchor frame of an artifact region for a target region in the feature map B, and dynamically adjusting the size and position of the anchor frame according to local information of the down-sampling rate, texture complexity and edge strength of the feature map B to obtain an artifact region detection result after the anchor frame is labeled. Automatic artifact detection and quantitative evaluation of the influence degree can be achieved, the accuracy and efficiency of medical image interpretation are improved, and more reliable data support is provided for clinical decision making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of image processing, and more specifically, to an image artifact detection method and system based on deep learning technology. Background Art

[0002] The statements in this section merely provide background technical information related to the present disclosure and do not necessarily constitute prior art.

[0003] With the rapid development of Magnetic Resonance Imaging (MRI) and Computed Tomography (CT) technologies, medical imaging systems have been continuously optimized in terms of hardware configuration and image reconstruction algorithms, significantly improving the resolution and quality of images. In the update and iteration of imaging systems, by adopting advanced detectors, radiofrequency coils, parallel acquisition technologies, and image reconstruction algorithms such as iterative reconstruction and deep learning reconstruction, the occurrence frequency of artifacts has been significantly reduced compared to the past.

[0004] However, even with these advancements, artifacts are still an inevitable part of MR and CT imaging, and the artifact problem still commonly exists and is inevitable in clinical images. Artifacts refer to the unrealistic features that appear in CT images, such as lines, shadows, light spots, etc., which will interfere with the correct interpretation of the original image. Artifacts refer to the non-anatomical structures that appear in images, such as stripes, bright spots, dark bands, blurred edges, etc. The formation reasons of artifacts are diverse, including both equipment-related problems such as partial volume effect, beam hardening, photon deficiency, undersampling, etc., and factors caused by patient behavior such as respiratory movement, metal implants, cardiac pulsation, etc. These artifact features often interfere with doctors' accurate judgment of lesions or normal tissue structures, thus affecting the accuracy and reliability of diagnosis.

[0005] The inventors found in their research that in medical images (such as CT images or MR images), due to the variety of artifact types, large scale differences, and strong texture interference, existing artifact detection methods have obvious deficiencies in small-scale artifact recognition, positioning accuracy in complex backgrounds, and adaptability of anchor box sizes, resulting in high detection omission rates, large position deviations, and weak generalization abilities, affecting the quality of clinical image artifact detection. Summary of the Invention

[0006] To solve the above problems, the present disclosure proposes an image artifact detection method and system based on deep learning technology, which can achieve automatic artifact detection and quantitative evaluation of the influence degree, improve the accuracy and efficiency of medical image interpretation, and provide more reliable data support for clinical decision-making.

[0007] To achieve the above object, the present disclosure adopts the following technical solutions: One or more embodiments provide an image artifact detection method based on deep learning technology, including the following steps: Preprocess the acquired image to be processed, and extract features from the preprocessed image; Perform position encoding on the extracted feature map A, and fuse the position-encoded position features with feature map A; Pass the fused feature map through an encoder and a decoder for multi-layer encoding and decoding to obtain a feature map B containing the artifact target area; Generate anchor boxes for the artifact areas in the target area of feature map B, and dynamically adjust the size and position of the anchor boxes according to the local information of the downsampling rate, texture complexity, and edge intensity of feature map B to obtain the detection result of the artifact area with labeled anchor boxes.

[0008] One or more embodiments provide an image artifact detection system based on deep learning technology, including: A feature extraction module configured to preprocess the acquired image to be processed and extract features from the preprocessed image; A position feature encoding module configured to perform position encoding on the extracted feature map A and fuse the position-encoded position features with feature map A; A feature analysis module configured to pass the fused feature map through an encoder and a decoder for multi-layer encoding and decoding to obtain a feature map B containing the artifact target area; An anchor box optimization module configured to generate anchor boxes for the artifact areas in the target area of feature map B, and dynamically adjust the size and position of the anchor boxes according to the local information of the downsampling rate, texture complexity, and edge intensity of feature map B to obtain the detection result of the artifact area with labeled anchor boxes.

[0009] An electronic device includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps in the above image artifact detection method based on deep learning technology are completed.

[0010] A computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the steps in the above image artifact detection method based on deep learning technology are completed.

[0011] Compared with the prior art, the beneficial effects of the present disclosure are: The present disclosure performs position encoding on feature maps, embeds spatial position information into the features, and enhances the model's spatial perception ability; through a dynamic anchor box scaling mechanism, it automatically adjusts the box position and size in combination with texture complexity and edge intensity, enhancing the accuracy of anchor box matching; it adopts an encoder-decoder structure combined with local and global attention to enhance feature modeling and context understanding ability, and solves the problem that the model's perception range is limited in the case of complex artifacts background, overlapping occlusion. It solves the problems of low accuracy of artifact detection, inappropriate anchor boxes, and insufficient spatial positioning ability in medical images. Especially when facing small-scale, complex texture or low-frequency artifact types, it overcomes the technical defects that existing methods are difficult to balance accuracy and generalization ability.

[0012] The advantages of the present disclosure and the advantages of additional aspects will be described in detail in the following specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The accompanying drawings forming a part of this disclosure are used to provide a further understanding of the present disclosure. The schematic embodiments and descriptions thereof of the present disclosure are used to explain the present disclosure and do not constitute a limitation to the present disclosure.

[0014] Figure 1 is a flowchart of the image artifact detection method in Embodiment 1 of the present disclosure; Figure 2 is a schematic structural diagram of the image artifact detection model in Embodiment 1 of the present disclosure; Figure 3 is the first artifact detection effect diagram in the simulation experiment of Embodiment 1 of the present disclosure; Figure 4 is the second artifact detection effect diagram in the simulation experiment of Embodiment 1 of the present disclosure; Figure 5 is the third artifact detection effect diagram in the simulation experiment of Embodiment 1 of the present disclosure; Figure 6 is the system interface diagram of image artifact detection in the simulation experiment of Embodiment 1 of the present disclosure; DETAILED DESCRIPTION OF THE EMBODIMENTS The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments.

[0015] It should be noted that the following detailed descriptions are all exemplary and are intended to provide further explanations of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure belongs.

[0016] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should also be understood that when the terms "comprise" and / or "include" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof. It should be noted that, without conflict, the various embodiments and features in the present disclosure may be combined with each other. The embodiments will be described in detail below with reference to the drawings.

[0017] Embodiment 1 In the technical solutions disclosed in one or more embodiments, as Figures 1 to 6 shown, an image artifact detection method based on deep learning technology includes the following steps: Step 1, preprocess the acquired image to be processed, and extract features from the preprocessed image; Step 2, perform position encoding on the extracted feature map A, and fuse the position-encoded position features with the feature map A; Step 3, pass the fused feature map through an encoder and a decoder for multi-layer encoding and decoding to obtain a feature map B containing the artifact target region; Step 4, generate anchor boxes for the target region in the feature map B, and dynamically adjust the size and position of the anchor boxes according to the distribution characteristics, texture complexity, and local information of the edge intensity of the artifacts in the image to obtain the detection result of the artifact region after annotating the anchor boxes; Further, Step 5: Calculate the ratio of the area of the anchor box to the area of the original image to be processed based on the dynamically adjusted anchor box size, and use it as the quantified detection result. The classification of artifact inspection can be determined by the size of the area ratio, such as severe or not severe. A ratio threshold can be set, and when it exceeds the set ratio, it is considered a severe artifact.

[0018] In this embodiment, position encoding is performed on the feature map to embed spatial position information into the features, improving the model's spatial perception ability; through the dynamic anchor box scaling mechanism, the position and size of the box are automatically adjusted in combination with texture complexity and edge intensity, enhancing the accuracy of anchor box matching; the encoder-decoder structure is combined with local and global attention to enhance the feature modeling and context understanding ability, solving the problem that the model's perception range is limited in the case of complex artifact backgrounds, overlapping occlusions. It solves the problems of low accuracy in artifact detection, inappropriate anchor boxes, and insufficient spatial positioning ability in medical images, especially when facing small-scale, texture-complex, or low-frequency artifact types, overcoming the technical defects that existing methods are difficult to balance accuracy and generalization ability.

[0019] In step 1, the image to be processed is a medical image, which can be a magnetic resonance imaging (MRI) image (abbreviated as MR image for short) and a CT image. Preprocessing the acquired image can include denoising processing and data augmentation. Specifically: Step 11: Perform image denoising to remove images that do not meet the quality requirements, such as abnormal exposure, blurred images, non-compliant sizes, missing annotations, and damaged formats; Step 12: Perform various random transformations, which can include random cropping, random rotation, random scaling, color jitter, etc., to generate diverse perspective pictures; Step 13: Convert the data format to make the dataset more suitable for the object detection task; In step 1, the preprocessed image is transmitted to the improved ResNet50 network for feature extraction. The improved ResNet50 network sets a feature retention layer at the output end of the third stage (Stage3) of the network. A multi-scale feature pyramid is connected after the feature retention layer. The multi-scale feature pyramid includes multiple sub-pyramid layers with unequal strides connected in cascade. The feature retention layer retains the feature map before the output of Stage3 and adjusts the resolution; the sub-pyramid layers perform layer-by-layer feature extraction on the feature map of the feature retention layer and then output feature map A; The ResNet50 network is a convolutional neural network structure composed of five stages: Conv1 and Stage1 to Stage5. Each stage contains multiple residual modules. In step 2, the improved ResNet50 network is as follows Figure 2 As shown, the third stage (Stage3) of the ResNet50 network is sequentially connected to the feature retention layer and multiple sub-pyramid layers with unequal strides; The feature retention layer is configured to extract the feature map output by the third stage (Stage3) of the ResNet50 network and perform upsampling, specifically adjusting it to a feature map with a resolution of 1 / 8 (i.e., 64×64); The sub-pyramid layer is configured to detect different-scale artifact regions for the feature map and implement feature extraction of artifact regions of different scales; In this embodiment, after the traditional Stage3 output, an independent sub-pyramid structure is constructed instead of continuing to enter Stage4 and Stage5 of ResNet; each sub-pyramid layer performs downsampling and transformation on the features output by Stage3 and no longer depends on the outputs of Stage4 and Stage5.

[0020] In this embodiment, the ResNet50 structure is innovatively improved by breaking its downsampling logic, and high-resolution features are extracted and enhanced at the Stage3 stage. To avoid the loss of small artifact information caused by downsampling, the feature map before the output of Stage3 is retained as the basis for unified high-resolution features. Then, sub-pyramid layers are derived on the unified high-resolution layer to enhance the detection ability for multi-scale artifacts. Without relying on cross-stage fusion, problems such as feature offset and alignment difficulties are avoided, and the overall structural stability and detection accuracy are improved.

[0021] Artifacts in medical images have obvious size diversity: from small metal buttons to large physiological motion artifacts. Each layer of the pyramid is responsible for detecting artifacts within a specific size range to ensure that small, medium, and large artifacts are not missed. In this embodiment, four sub-pyramid layers are set up, specifically as follows: The first sub-pyramid layer: at 1 / 8 resolution, detecting small artifacts of 5 - 15px; The second sub-pyramid layer: at 1 / 16 resolution, detecting medium-small artifacts of 15 - 25px; The third sub-pyramid layer: at 1 / 24 resolution, detecting medium artifacts of 25 - 40px; The fourth sub-pyramid layer: at 1 / 32 resolution, detecting large artifacts larger than 40px; The setting of the sub-pyramid layers in this embodiment can achieve multi-scale artifact detection and improve the recognition ability and localization accuracy for artifacts of different sizes.

[0022] Since the convolutional neural network itself does not have explicit absolute position perception ability and pays more attention to relative information such as local texture and edges. Even if an artifact appears in the corner or edge of the image, the CNN is not sensitive to its spatial semantics; for the same artifact shape in the feature map, if it appears in different positions, the network may make the same response, affecting the localization accuracy and false detection control.

[0023] This embodiment uses positional encoding precisely to solve this "spatial unconsciousness" problem, introducing absolute position semantics and enhancing the discriminability of the artifact area. A learnable two-dimensional positional encoding mechanism (2D Positional Encoding) is introduced into the artifact detection network to explicitly model the spatial position information of each pixel or region in the image, thereby improving the model's understanding ability of spatial relationships in complex structures.

[0024] In step 2, the obtained feature map A is subjected to positional encoding, including the following steps: Step 21, coordinate grid initialization: Construct two-dimensional coordinate matrices Grid_X and Grid_Y, representing the horizontal position and vertical position of each pixel point in the image respectively; Specifically, first construct a 64×64 basic feature grid. On the 64×64 basic feature grid, 9 anchor points in total of 3×3 are set in each cell, and the distribution of the anchor points is improved from dot-like to surface-like dense coverage, so as to enhance the detection ability for small artifacts and edge targets.

[0025] In this embodiment, through high-density anchor point coverage, the recognition rate of small target artifacts is improved, and the problem that the existing artifact detection methods have inaccurate recognition of small-scale and rare types of artifacts is solved.

[0026] Step 22: Linear fusion encoding: Based on the constructed coordinate matrix, extract the coordinates of each pixel point; input the two-dimensional coordinate vector [x, y] corresponding to each pixel point into the linear transformation network, and transform it into the channel space with the same dimension as the input feature map A to obtain a learnable spatial position vector field. In step 22, to extract the coordinate values, specifically: Take the coordinate values Grid_X and Grid_Y as two channels, and splice them to form a two-dimensional position grid Grid_Pos∈R (H×W×2) ; where H represents the height of the feature map, W represents the width of the feature map, 2 represents that each position contains two dimensions of coordinates: the horizontal coordinate (X) and the vertical coordinate (Y), corresponding to each pixel position on the feature map. All coordinate values are normalized and mapped to the range [0, 1] to generate a two-dimensional coordinate vector [x, y], representing its normalized position in the grid. Organize these [x, y] to form a shape of Grid_Pos∈R (H×W×2) A two-dimensional position grid, and the horizontal and vertical coordinates of each pixel point are placed in its channel; Step 23: Sinusoidal Encoding: Apply the Sinusoidal function to the position vectors in the learnable spatial position vector field to form a spatial encoding with periodicity and hierarchical expression ability, that is, obtain the position encoding of each pixel point in the feature map; enable the model to obtain non-linear position perception ability and enhance the recognition ability for information such as the boundary and center offset of the artifact area.

[0027] Among them, the Sinusoidal function, for example: a combination of the sine function and the cosine function; In step 2, fuse the position features after position encoding with the feature map A, specifically as follows: After performing a full connection operation and splicing the position encoding and the feature map A in the channel dimension, then perform dimensionality reduction through a 1×1 convolutional layer, so that the fused feature retains both texture and spatial information.

[0028] In step 3, an encoder-decoder structure is adopted for encoding and decoding; Step 31: Layer-by-layer encode the fused feature map, perform multi-scale feature extraction, and obtain multi-layer encoded features; The encoder in this step adopts an encoder stack structure: including multiple encoder layers cascaded in sequence, and each encoder layer includes a self-attention mechanism and a feed-forward neural network in sequence; Process the features through multiple layers of encoders. The shallow encoder uses a 32-pixel local window attention to capture small artifact details; the deep encoder is extended to a 64-pixel window to establish large-scale context associations. Each layer contains a self-attention mechanism and a feed-forward neural network. The self-attention mechanism enables the model to capture the long-term dependencies between image features globally, further enhancing the model's global perception ability of the artifact region.

[0029] In this embodiment, to improve the global modeling ability of the artifact region and the multi-scale feature integration effect, a multi-layer encoder stack (Encoder Stack) structure is introduced after the backbone feature extraction, combining the local window attention and the global self-attention mechanism to achieve layer-by-layer modeling from local details to large-scale dependencies.

[0030] Step 32: Decoder decoding: Layer-by-layer decode and gradually refine the encoded features to generate an artifact prediction map consistent with the original image space, that is, feature map B; In this embodiment, an encoder-decoder structure is used to process the feature map, which can ensure that the prediction result is highly consistent with the boundary of the real artifact region. By fitting complex structures through a multi-scale decoding process, the model can adapt to irregular artifact shapes.

[0031] Step 4 is the Dynamic Anchor Scaling mechanism, which adaptively adjusts the anchor box size in combination with the local features of the image and is the key mechanism to improve the detection sensitivity and positioning accuracy in artifact detection; In Step 4, according to the local information of the downsampling rate, texture complexity, and edge intensity of feature map B, dynamically adjust the size and position of the anchor box, including the following steps: Step 41: Initial anchor box mapping: For feature map B, initialize the anchor boxes of the reference scale and map them to the position encoding grid of the Stage3 feature map to obtain the position coordinates of each initial anchor box; Specifically, based on the basic grid constructed by the 64×64 feature map output by Stage3, initialize three anchor boxes of the reference scale: 16×16, 32×32, 64×64, corresponding to small artifacts, medium artifacts, and large artifacts respectively; map the reference anchor boxes (16×16, 32×32, 64×64) to the position encoding grid of the Stage3 feature map, and each anchor point corresponds to an 8×8 pixel area of the original image to be processed; Step 42: Perform a linear transformation on the coordinates of the initial anchor boxes: Convert the anchor boxes from the feature map space to the original image space to be processed, and perform the following linear mapping: , ; , ; where, and represent the coordinates of the anchor points in the feature map, and represent the positions of the centers of the anchor boxes in the original image; Step 43: Based on the transformed coordinates, dynamically calculate the sizes and positions of the anchor boxes according to the parameters of the extracted feature map, local texture complexity, and edge features, and update the initialized anchor boxes. Specifically, the calculation formula is as follows:

[0032] where, represents the displacement coefficient in the feature space; represents the scale factor of the feature map, which is a parameter of the feature map; represents the perturbation amplitude coefficient; : Gaussian parameter, generating a random number obeying the N(0, 0.25) distribution; represents the predicted value of the base width; represents the texture sensitivity coefficient; represents the local texture complexity; represents the predicted value of the base height; represents the edge feature intensity; : regional average feature; Furthermore, the above process can be implemented by constructing an image artifact detection model, including an improved ResNet50 network, a multi-scale feature pyramid, a position encoding and feature analysis module, and an anchor box optimization module; among them, the position encoding and feature analysis module includes: A position feature encoding module, configured to perform position encoding on the extracted feature map A and fuse the position-encoded position features with the feature map A; A feature analysis module, including an encoder and a decoder, performing multi-layer encoding and decoding to obtain a feature map B containing the artifact target region; The anchor box optimization module is configured to generate anchor boxes for the artifact regions for the target regions in Feature Map B, and dynamically adjust the size and position of the anchor boxes according to the local information of the downsampling rate, texture complexity, and edge intensity of Feature Map B in the image, so as to obtain the detection result of the artifact region after annotating the anchor boxes; Furthermore, it also includes the process of training the above-mentioned image artifact detection model. For the image artifact detection model with CT and MR as the dual-modal input, through a phased and multi-mechanism training process, the training of a high-quality artifact detection model is realized, including two stages: Step 1, Pre-training stage: Based on data of different modalities, the image artifact detection model is separately trained in stages, and the detection branches after training with data of different modalities are obtained; The modality refers to the data modality, such as CT modality data and MR modality data; First, the image artifact detection model is independently trained on CT and MR single-modal data respectively to obtain two detection branches, namely the CT branch and the MR branch; The CT branch can learn the high-contrast features of metal artifacts, and the MR branch can identify the fuzzy boundary characteristics of motion artifacts. The first 3 stages of the ResNet50 network are shared by the two branches to maintain the consistency of low-level feature extraction.

[0033] Step S2, Joint fine-tuning stage: Data of different modalities are mixed according to a set ratio, and the pre-trained image artifact detection model is jointly trained; Specifically, a dynamic batch combination strategy is adopted, and each batch contains 50% CT samples and 50% MR samples. When calculating the loss, the learning weights are automatically adjusted according to the prediction errors of each modality in the current batch: when the average loss of CT samples is higher than the average loss of MR samples, the gradient backpropagation intensity of CT samples is increased, and vice versa, the learning of MR samples is strengthened.

[0034] Furthermore, in the joint fine-tuning stage, there may be conflicts or imbalances in the loss gradients of the CT modality and the MR modality during the training process, which affects the learning effect. Aiming at the problem that the optimization objectives of the CT branch and the MR branch may conflict during the multi-modal joint training process, by introducing a gradient conflict detection and loss adaptive adjustment mechanism, the gradient directions and loss weights of each modality are dynamically coordinated and trained to obtain a multi-modal fusion detection model with consistent optimization paths and self-balanced learning weights.

[0035] Specifically, the gradient conflict detection method includes: calculating the included angle between the gradient directions of the two detection branches in real time. If the included angle is greater than 90°, it is regarded as a conflict, and the gradient projection mechanism is started to remove the reverse component of the gradient direction, and only the common direction component is retained for parameter update; Specifically, the loss adaptive adjustment mechanism includes: Establish a loss ratio dynamic adjustment factor λ, and its calculation formula is as follows:

[0036] Wherein, is the predicted loss of CT modality data, is the predicted loss of MR modality data, is a set parameter; When the predicted loss of CT modality data is significantly greater than the MR loss λ approaches 1, and the training focuses on CT modality optimization; conversely, when the MR modality loss is higher, λ approaches 0, focusing on the MR modality. This mechanism realizes the autonomous coordination between modalities through a sliding balance point.

[0037] Furthermore, the loss function constructed in the joint fine-tuning stage is as follows:

[0038] Furthermore, a cross-modal feature interaction mechanism is adopted: for the semantic difference and artifact feature complementarity problems of CT and MR modality images, a unified detection model with the ability of inter-modal alignment and collaborative enhancement of artifact recognition is obtained by constructing a contrast learning and knowledge distillation mechanism to jointly train the multi-modal image artifact detection model.

[0039] Specifically, feature space alignment, that is, contrast learning: construct a CT-MR paired dataset, including positive samples of the same part and negative sample pairs of different parts, and conduct joint training; Furthermore, a bidirectional knowledge distillation path can also be designed: after the artifact heat map output by the CT branch is processed by temperature softening, it is used as the auxiliary supervision signal of the MR branch; after the artifact heat map output by the MR branch is processed by temperature softening, it is used as the auxiliary supervision signal of the CT branch. This process enables the metal artifact discrimination knowledge learned by the CT branch to be transferred to the MR branch, and at the same time, the motion artifact detection ability of the MR branch feeds back to the CT branch. The KL divergence is used to measure the difference in the predicted distribution, and the distillation weight is set to 0.3 to avoid overfitting; Furthermore, in the training method of joint training, an adversarial training method can be adopted. The discriminator takes the features predicted by the image artifact detection model as input. The discriminator adopts a three-layer fully connected network, and the parameters of the discriminator are updated twice per round and then the parameters of the model are updated once.

[0040] In this embodiment, by constructing a phased and multi-mechanism multi-modal training method, the overall performance and practicality of the image artifact detection model are significantly improved. By pre-training on CT and MR modality data respectively, the model fully masters the modality-specific features of metal artifacts and motion artifacts, improving the accuracy and robustness of artifact detection. The joint fine-tuning stage introduces contrastive learning and bidirectional knowledge distillation mechanisms to achieve feature alignment and knowledge mutual guidance between the CT and MR branches, enhancing the cross-modal recognition ability of the model under mixed-modal inputs. Aiming at the optimization conflict problem in the multi-modal training process, a dynamic gradient coordination mechanism and a loss weight adaptive adjustment strategy are proposed, significantly improving the training stability and convergence efficiency. At the same time, through the joint optimization of modal adversarial training and structural reconstruction, the model is guided to extract modality-independent artifact essential features, enhancing the integrity and interpretability of feature expression. The finally trained model can be widely applied to CT, MR or mixed-modal image inputs to achieve unified and efficient artifact detection, meeting the actual clinical needs in multi-source data scenarios.

[0041] To verify the effectiveness of the proposed deep learning-based multi-modal image artifact detection method in this embodiment, a simulation experimental environment is constructed, and artifact detection tests are carried out on the actually collected CT and MR medical images, and the detection results are visually displayed and quantitatively analyzed.

[0042] In the experiment, a medical image dataset containing various artifact types is selected, covering typical scenarios such as metal artifacts (such as necklaces, internal fixation brackets), motion artifacts, and non-structural imaging artifacts. The original images are input into the artifact detection model constructed in the present invention, and after being processed by modules such as feature extraction, position encoding, encoding-decoding, and dynamic anchor box scaling, the prediction results of the artifact regions are output.

[0043] As Figures 4 to 6 shown, the model can accurately identify the artifact regions in different modality images, and the detection boxes precisely cover the artifact boundaries, with good positioning accuracy and boundary awareness ability. For the metal artifacts in CT images, this model can accurately capture the high-contrast edges; for the blurred motion artifacts in MR images, it can also stably identify their spatial distribution positions. The artifact regions detected by the model are marked with red boxes, intuitively showing the model's adaptability to complex artifact features.

[0044] As Figure 6 shown, for the interface diagram of the system, the system can also automatically count key parameters such as the number of artifacts, distribution layers, area range, and proportion, and will Figure 6The various detection indexes shown on the right side of the interface diagram: For example, a certain CT sequence image has a total of 125 layers, and 28 layers are detected to have artifacts, accounting for 22.4%; the maximum area of the artifacts is 657.56 mm², the total area of the artifacts is 12569.32 mm², and the proportion of the artifact area is 2.5%. The above results show that the image artifact detection model of this embodiment not only has the ability to detect artifacts at the pixel level, but also supports regional quantification and risk assessment, providing an effective reference for subsequent image quality control and diagnostic tips.

[0045] In summary, the method of this embodiment can accurately identify and quantify the artifact area in various modal medical images, and shows good generalization ability and detection stability under different artifact types, which can effectively improve the interpretation efficiency and diagnostic accuracy of medical images and has important clinical application value.

[0046] Embodiment 2 Based on Embodiment 1, an image artifact detection system based on deep learning technology is provided in this embodiment, including: A feature extraction module, configured to perform preprocessing on the acquired image to be processed and extract features from the preprocessed image; A position feature encoding module, configured to perform position encoding on the extracted feature map A and fuse the position-encoded position feature with the feature map A; A feature analysis module, configured to perform multi-layer encoding and decoding on the fused feature map through an encoder and a decoder to obtain a feature map B containing the artifact target area; An anchor box optimization module, configured to generate an anchor box for the artifact area for the target area in the feature map B, and dynamically adjust the size and position of the anchor box according to the local information of the downsampling rate, texture complexity and edge strength of the feature map B to obtain the detection result of the artifact area with the anchor box marked.

[0047] It should be noted here that each module in this embodiment corresponds to each step in Embodiment 1, and the specific implementation process is the same, so it will not be repeated here.

[0048] Embodiment 3 This embodiment provides an electronic device, including a memory and a processor, as well as computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps in the image artifact detection method based on deep learning technology described in Embodiment 1 are completed.

[0049] Embodiment 4 This embodiment provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by the processor, the steps in the image artifact detection method based on deep learning technology described in Embodiment 1 are completed.

[0050] The above are only the preferred embodiments of the present disclosure, and are not intended to limit the present disclosure. For those skilled in the art, the present disclosure may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.

[0051] Although the specific implementation manners of the present disclosure are described above in conjunction with the accompanying drawings, it is not a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that based on the technical solutions of the present disclosure, various modifications or deformations that can be made by those skilled in the art without creative efforts are still within the protection scope of the present disclosure.

Claims

1. An image artifact detection method based on deep learning technology, characterized in that, It includes the following steps: Preprocess the obtained image to be processed, and extract features from the preprocessed image; Perform position encoding on the extracted feature map A, and fuse the position features after position encoding with feature map A; Pass the fused feature map through an encoder and a decoder for multi-layer encoding and decoding to obtain a feature map B containing the artifact target area; Generate anchor boxes for the target area in feature map B, and dynamically adjust the size and position of the anchor boxes according to the local information of the downsampling rate, texture complexity, and edge intensity of feature map B to obtain the detection result of the artifact area with labeled anchor boxes.

2. The image artifact detection method based on deep learning technology according to claim 1, wherein: Calculate the ratio of the area of the anchor box to the area of the original image to be processed based on the dynamically adjusted anchor box size as the quantified detection result.

3. The image artifact detection method based on deep learning technology according to claim 1, wherein: Transmit the preprocessed image to the improved ResNet50 network for feature extraction. The improved ResNet50 network is provided with a feature retention layer at the output end of the third stage Stage3 of the network. The feature retention layer is connected to a multi-scale feature pyramid, and the multi-scale feature pyramid includes multiple sub-pyramid layers with unequal strides connected in cascade; The feature retention layer is used to retain the feature map before the output of Stage3 and adjust the resolution; The sub-pyramid layer is used to perform layer-by-layer feature extraction on the output feature map of the feature retention layer and then output feature map A.

4. The image artifact detection method based on deep learning technology according to claim 1, wherein: Performing position encoding on the extracted feature map A includes the following steps: Construct two-dimensional coordinate matrices Grid_X and Grid_Y, which respectively represent the horizontal position and vertical position of each pixel point in the image; Based on the constructed coordinate matrices, extract the coordinates of each pixel point; transform the two-dimensional coordinate vector corresponding to each pixel point into a channel space with the same dimension as the input feature map A to obtain a learnable spatial position vector field; Apply the Sinusoidal function to the position vectors in the learnable spatial position vector field to obtain the position encoding of each pixel point in the feature map.

5. The image artifact detection method based on deep learning technology according to claim 1, wherein: Fusing the position features after position encoding with feature map A is specifically as follows: After performing a full connection operation and splicing the position encoding and feature map A in the channel dimension, perform a dimensionality reduction operation through a convolutional layer.

6. The image artifact detection method based on deep learning technology according to claim 1, wherein: Dynamically adjusting the size and position of the anchor box includes the following steps: For feature map B, initialize the anchor box with the benchmark scale and map it to the position encoding grid of the Stage3 feature map to obtain the position coordinates of each initial anchor box; Convert the anchor box from the feature map space to the original image space to be processed; Based on the transformed coordinates, dynamically calculate the size and position of the anchor boxes according to the parameters of the extracted feature map, the local texture complexity, and the edge features, and update the initialized anchor boxes.

7. The method for detecting image artifacts based on deep learning technology according to claim 1, wherein: It further includes constructing an image artifact detection model, and constructing the image artifact detection model includes an improved ResNet50 network, a multi-scale feature pyramid, a position encoding and feature analysis module, and an anchor box optimization module; among them, the position encoding and feature analysis module includes: A position feature encoding module, configured to perform position encoding on the extracted feature map A, and fuse the position-encoded position features with the feature map A; A feature analysis module, including an encoder and a decoder, performing multi-layer encoding and decoding to obtain a feature map B containing the artifact target area; An anchor box optimization module, configured to generate anchor boxes for the artifact area for the target area in the feature map B, and dynamically adjust the size and position of the anchor boxes according to the local information of the downsampling rate, texture complexity, and edge intensity of the feature map B in the image, to obtain the detection result of the artifact area after labeling the anchor boxes.

8. An image artifact detection system based on deep learning technology, characterized in that, It includes: A feature extraction module, configured to preprocess the acquired image to be processed and extract features from the preprocessed image; A position feature encoding module, configured to perform position encoding on the extracted feature map A, and fuse the position-encoded position features with the feature map A; A feature analysis module, configured to pass the fused feature map through an encoder and a decoder, perform multi-layer encoding and decoding, to obtain a feature map B containing the artifact target area; An anchor box optimization module, configured to generate anchor boxes for the artifact area for the target area in the feature map B, and dynamically adjust the size and position of the anchor boxes according to the local information of the downsampling rate, texture complexity, and edge intensity of the feature map B, to obtain the detection result of the artifact area after labeling the anchor boxes.

9. An electronic device, characterized in that, It includes a memory, a processor, and computer instructions stored on the memory and running on the processor. When the computer instructions are run by the processor, the steps in the method for detecting image artifacts based on deep learning technology according to any one of claims 1-7 are completed.

10. A computer-readable storage medium, characterized in that, For storing computer instructions, when the computer instructions are executed by the processor, the steps in the method for detecting image artifacts based on deep learning technology according to any one of claims 1-7 are completed.

Citation Information

Cited By

  • Real-time monitoring method for copper-plated steel strip electroplating production line

    CN120894364A

  • A real-time monitoring method for a copper-plated steel strip electroplating production line

    CN120894364B

  • Medical image quality enhancement and artifact correction method and system based on artificial intelligence

    CN121685770A

  • Artificial intelligence-based medical image quality enhancement and artifact correction method and system

    CN121685770B

  • Multi-modal medical image artifact intelligent detection method and system based on deep learning

    CN121767355A