Quality control method and system for intra-frame weighted medical image with similar specific attention

By employing an intra-frame weighted medical image quality control method, the problem of weak discrimination capability in 3D medical image processing is solved, the sensitivity to subtle artifacts and positional deviations is improved, multi-dimensional and precise image quality control is achieved, and the quality control efficiency and equipment adaptability are enhanced.

CN120953289AActive Publication Date: 2025-11-14安徽影联云享医疗科技有限公司

Patent Information

Application Number
CN202511484979.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2025-11-14
Estimated Expiration
2045-10-17

AI Technical Summary

Technical Problem

Existing medical image processing methods have weak discrimination capabilities when dealing with 3D medical images, and are less sensitive to subtle artifacts and positional deviations, resulting in weak recognition capabilities.

Method used

A neural network-based intra-frame weighted medical image quality control method is adopted. Through data preprocessing, spatial feature extraction, temporal enhancement feature extraction, category-guided intra-frame weighted score calculation and model optimization, class activation maps are generated to focus on key anatomical regions and improve the model's sensitivity to subtle artifacts and positional deviations.

Benefits of technology

It enables multi-dimensional, precise, and real-time quality assessment of 3D medical images, improves quality control efficiency, reduces duplicate scanning rates, and is adaptable to various image types and devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120953289A_ABST
    Figure CN120953289A_ABST
Patent Text Reader

Abstract

The invention discloses a quality control method and system for an intra-frame weighted medical image with similar specific attention, relates to the technical field of medical images, and solves the technical problems that when an existing medical image processing method is used for processing a 3D medical image, the discrimination capability is weak, and the sensitivity to subtle artifacts or body position deviation is low, so that the recognition capability is weak. Comprising the steps that medical image data are acquired, and the medical image data comprise a 3D medical image and an evaluation system corresponding to the 3D medical image; performing data labeling on the 3D medical image based on an evaluation system to obtain a plurality of sequence data; dividing the plurality of sequence data into a training set, a verification set and a test set; training the medical image quality control model based on the training set, the verification set and the test set; the training process comprises data preprocessing, spatial feature extraction, time sequence enhancement feature extraction, class-guided intra-frame weighted score calculation, model optimization and class activation graph generation; and the model identification capability is enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of medical imaging technology, specifically a method and system for intra-frame weighted medical image quality control based on specific attention. Background Technology

[0002] Medical image quality control is a crucial component of modern clinical medical imaging work. Through the automated analysis of images such as CT, MRI, and X-rays, it provides key quality control data for medical images, laying the foundation for high-quality subsequent diagnoses. With the development of artificial intelligence technology, medical image quality control methods have evolved from traditional manual feature extraction to end-to-end deep learning. Early methods heavily relied on manually designed feature descriptors (such as HOG and SIFT) and machine learning classifiers (such as SVM), resulting in limited feature representation capabilities and weak generalization performance.

[0003] In recent years, the introduction of deep learning technologies, especially convolutional neural networks (CNNs) and the Transformer architecture, has significantly improved the accuracy and robustness of medical image analysis. CNNs can automatically learn the local texture and shape features of images through local receptive fields and hierarchical feature extraction mechanisms; while Transformers overcome the limitation of CNNs' limited receptive fields by achieving global context modeling through self-attention.

[0004] However, current technologies still face significant challenges when dealing with three-dimensional medical imaging data. Medical images are essentially continuous three-dimensional spatial data (such as axial, sagittal, and coronal slice sequences from CT scans) and often include a temporal dimension (such as four-dimensional data from dynamic contrast-enhanced MRI). Most existing methods decompose 3D data into 2D slices for processing or directly apply computationally intensive 3D deep learning models. However, 2D image classification methods have inherent limitations in 3D medical image analysis. 2D convolutional neural networks (such as ResNet and VGG) perform well in natural image classification, but they suffer from spatial information collapse, temporal dependency loss, annotation ambiguity, and inefficiency when processing 3D medical images. While SWIN Transformer and its 3D variant (Video SWIN Transformer) have been introduced to address long-range dependency issues, they face several problems in medical image applications, such as computational complexity explosion, local overfitting, and data starvation. In particular, the global attention mechanism of Transformer over-smooths local features, making it difficult to capture subtle artifacts in medical images. Especially in artifact recognition tasks common in medical images (such as metal artifacts, motion artifacts, and radio frequency artifacts), existing methods may lead to loss of spatial context information, low computational efficiency, and missed detection of rare pathological features, thus failing to effectively identify them. Therefore, a frame-specific attention-based intra-weighted medical image quality control method and system is needed. Summary of the Invention

[0005] This application provides a type of attention-specific intra-frame weighted medical image quality control method and system, which solves the technical problem that existing medical image processing methods have weak discrimination ability when dealing with 3D medical image processing, and low sensitivity to subtle artifacts or positional deviations, resulting in weak recognition ability.

[0006] To achieve the above objectives, this application adopts the following technical solution: Firstly, a neural network-based intra-frame weighted medical image quality control method is provided, comprising: Acquire medical imaging data, including 3D medical images and their corresponding evaluation system; the 3D medical images are CT, MRI, or X-ray scans, such as chest CT images; the evaluation system includes quality control standards, category labels, and annotation labels corresponding to the 3D medical images. Based on the evaluation system, several sequence data are obtained by data annotation of 3D medical images; the several sequence data are divided into training set, validation set and test set; The medical image quality control model is trained based on the training set, validation set, and test set. The training process includes data preprocessing, spatial feature extraction, temporal enhancement feature extraction, category-guided intra-frame weighted score calculation, model optimization, and class activation map generation.

[0007] Based on the above technical solutions, this application provides a method and system for intra-frame weighted medical image quality control based on specific attention. Addressing the lack of spatial attention in the quality control process of CT and MR data in 3D image algorithms, this application acquires medical image data, including 3D medical images and their corresponding evaluation systems; annotates the 3D medical images according to the evaluation systems to obtain several sequence data; divides the sequence data into training, validation, and test sets; trains the medical image quality control model based on the training, validation, and test sets; the training process includes data preprocessing, spatial feature extraction, temporal enhancement feature extraction, category-guided intra-frame weighted score calculation, model optimization, and class activation map generation; the category-guided intra-frame weighted score calculation includes calculating the intra-frame weighted score for each spatial location; by calculating the response intensity of each spatial location to a specific quality control category, class-specific weighted features are generated, thereby highlighting the most discriminative region for quality control judgment in each frame, and visualizing 3D medical image quality control through frame perception capabilities. This effectively focuses on key anatomical regions and improves the model's sensitivity to subtle artifacts or positional deviations. Meanwhile, by integrating the spatial feature extraction of CNN, the temporal modeling of RNN, and the region focusing capabilities of the category-guided intra-frame weighting mechanism, it achieves multi-dimensional, precise, and real-time quality assessment of medical images. Furthermore, through frame-level evidence extraction, it provides interpretable frame numbers and heatmaps while giving quality control labels. This not only significantly improves quality control efficiency and reduces the rate of repeated scans, but also adapts to various image types and devices.

[0008] In conjunction with the first aspect above, in one possible implementation, the data annotation of 3D medical images based on the evaluation system to obtain several sequence data includes: The 3D medical image is frame-by-frame extracted to obtain several frames. The frames are then integrated according to their corresponding frame order to generate sequence data. The sequence data is labeled to obtain labeled sequence data. The data labeling includes category labeling and annotation labeling. The labeled sequence data includes several frames and the category of each spatial location. The frame order is the chronological order of the frames in the 3D medical image.

[0009] In conjunction with the first aspect above, in one possible implementation, the data preprocessing includes: Extract each frame image from the sequence data, convert each frame image to RGB mode, and sort them according to the original frame order to obtain serialized data.

[0010] In conjunction with the first aspect above, in one possible implementation, the spatial feature extraction includes: Frame images are extracted from the serialized data, and spatial features are extracted from the frame images using a convolutional neural network to obtain the spatial feature map of the frame images; the spatial feature map is flattened according to spatial location to obtain the spatial features corresponding to each spatial location.

[0011] In conjunction with the first aspect above, in one possible implementation, the temporal enhancement feature extraction includes: The process involves obtaining spatial features corresponding to each spatial location in each frame image, and using a recurrent neural network (RNN) to perform temporal modeling on the spatial features of the same spatial location in different frame images to obtain the temporal enhancement features of the corresponding spatial location. This process is repeated to obtain the temporal enhancement features of each spatial location sequentially. Specifically, to capture the temporal dependencies between video frames, a recurrent neural network (RNN), such as LSTM, GRU, or Mamba, can be introduced to perform temporal modeling on the spatial feature sequences of all frames, outputting the temporal enhancement features of each frame. One formula for calculating the temporal enhancement features is as follows:

[0012] in, The temporal enhancement feature for the j-th spatial location in the t-th frame image; The temporal enhancement feature is the spatial location of the j-th frame in the (t-1)-th frame image; The spatial features of the j-th spatial location in the t-th frame image; In conjunction with the first aspect above, in one possible implementation, the category-guided intra-frame weighted score calculation includes the following steps: A target category is set, and temporal enhancement features at each spatial location of the frame image are obtained. Based on each temporal enhancement feature, the intra-frame weighted score of the target category at each spatial location in the frame image is calculated. ; in, Let be the intra-frame weighted score of target category i at the j-th spatial position in the t-th frame image, which is the probability of target category i appearing at the j-th spatial position in the t-th frame image, and ∈[0,1], used to measure the importance of different spatial locations for category recognition; Let be the classifier vector corresponding to target category i, and ∈ B is the set temperature parameter used to control the degree of sharpening and smoothing; it can be fixed during training. The temporal enhancement feature corresponding to the j-th spatial location of the t-th frame image; Let be the temporal enhancement feature corresponding to the k-th spatial location of the t-th frame image, where k∈[1, hw] and hw is the total number of spatial locations in the feature map.

[0013] In conjunction with the first aspect above, in one possible implementation, the model optimization includes: Obtain the intra-frame weighted score of the target category at all spatial locations in each frame image, as well as the temporal enhancement features at all spatial locations in each frame image; generate the classification score of the target category based on the intra-frame weighted score and temporal enhancement features at all spatial locations in each frame image. The classification result is determined based on the classification score, the set label set is obtained, and the loss value is calculated based on the classification result and the label set. If the loss value is less than a set threshold, the model training is completed; otherwise, the relevant parameters are adjusted and the training is repeated.

[0014] In conjunction with the first aspect above, in one possible implementation, the classification score of the target category is generated based on the intra-frame weighted scores of all spatial locations in each frame image and the temporal enhancement features, including: Obtain the intra-frame weighted score corresponding to all spatial locations in the frame image, and perform weighted summation on the temporal enhancement features corresponding to each spatial location based on the intra-frame weighted score corresponding to each spatial location to obtain the frame-level class-specific features of the corresponding target category; The global class-specific features of the target type are obtained by temporally fusing the frame-level class-specific features corresponding to all frame images of the target type. Global average pooling is performed on the temporal enhancement features of the spatial locations in each frame image to obtain the background global features; The global class-specific features and background global features are fused through residual training to obtain the final features of the target class; The final feature is multiplied by the classification vector of the target category to obtain the category score corresponding to the target category.

[0015] In conjunction with the first aspect above, in one possible implementation, the class activation graph generation includes: Extract frame-level class-specific features corresponding to each frame image; using the formula: ; The response intensity of the target category i corresponding to the t-th frame image is calculated. The higher the response intensity value, the more likely the frame contains discriminative information most relevant to target category i; where, For frame-level class-specific features of target category i in the t-th frame image The frame-level class-specific feature sequence; ; The frame image with the largest response value is selected as the keyframe image corresponding to the target category; specifically, this is done using the formula: ; Select keyframe images, where t The number corresponding to the keyframe image; Obtain the spatial feature map corresponding to the keyframe image; recalculate the spatial weight map using the class vector; bilinearly interpolate to the original resolution, assign pseudo-color, and superimpose it on the original image to obtain the class activation map.

[0016] Secondly, this application provides a class-specific attention intra-weighted medical image quality control device, comprising: a processor and a storage medium; the storage medium includes instructions, and the processor is configured to execute the instructions to implement the method described in the first aspect and any possible implementation thereof. This class-specific attention intra-weighted medical image quality control device may be an electronic device or a chip within an electronic device.

[0017] Thirdly, this application provides a type of attention-specific intra-frame weighted medical image quality control system, including: a data acquisition module, a data partitioning module, and a model training module; The data acquisition module is used to acquire several medical image data. The data partitioning module is used to annotate 3D medical images based on an evaluation system to obtain several sequence data; and to partition the several sequence data into a training set, a validation set, and a test set. The model training module trains the medical image quality control model based on the training set, validation set, and test set. The training process includes data preprocessing, spatial feature extraction, temporal enhancement feature extraction, category-guided intra-frame weighted score calculation, model optimization, and class activation map generation.

[0018] Fourthly, this application provides a computer-readable storage medium storing instructions that, when executed on a class-specific attention intra-weighted medical image quality control device, cause the class-specific attention intra-weighted medical image quality control device to perform the methods described in the first aspect and any possible implementation thereof.

[0019] Fifthly, this application provides a computer program product containing instructions that, when run on a class-specific attention intra-weighted medical image quality control device, cause the class-specific attention intra-weighted medical image quality control device to perform the methods described in the first aspect and any possible implementation thereof.

[0020] This application provides a class-specific attention-based intra-frame weighted medical image quality control method and system. It acquires medical image data, including 3D medical images and their corresponding evaluation systems; annotates the 3D medical images according to the evaluation systems to obtain several sequence data; divides the sequence data into training, validation, and test sets; and trains a medical image quality control model based on the training, validation, and test sets. The training process includes data preprocessing, spatial feature extraction, temporal enhancement feature extraction, category-guided intra-frame weighted score calculation, model optimization, and class activation map generation. The category-guided intra-frame weighted score calculation includes calculating the intra-frame weighted score for each spatial location; by calculating the response intensity of each spatial location to a specific quality control category, class-specific weighted features are generated, thereby highlighting the most discriminative region for quality control judgment in each frame, and visualizing 3D medical image quality control through frame perception capabilities. This effectively focuses on key anatomical regions and improves the model's sensitivity to subtle artifacts or positional deviations. Meanwhile, by integrating the spatial feature extraction of CNN, the temporal modeling of RNN, and the region focusing capabilities of the category-guided intra-frame weighting mechanism, it achieves multi-dimensional, precise, and real-time quality assessment of medical images. Furthermore, through frame-level evidence extraction, it provides interpretable frame numbers and heatmaps while giving quality control labels. This not only significantly improves quality control efficiency and reduces the rate of repeated scans, but also adapts to various image types and devices.

[0021] It should be understood that the descriptions of technical features, technical solutions, beneficial effects, or similar language in this application do not imply that all features and advantages can be achieved in any single embodiment. Rather, it is understood that the description of a feature or beneficial effect means that a specific technical feature, technical solution, or beneficial effect is included in at least one embodiment. Therefore, the descriptions of technical features, technical solutions, or beneficial effects in this specification do not necessarily refer to the same embodiment. Furthermore, the technical features, technical solutions, and beneficial effects described in this embodiment can be combined in any suitable manner. Those skilled in the art will understand that embodiments can be implemented without one or more specific technical features, technical solutions, or beneficial effects of a particular embodiment. In other embodiments, additional technical features and beneficial effects may be identified in specific embodiments that do not embody all embodiments. Attached Figure Description

[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0023] Figure 1 This is a schematic diagram illustrating the steps of the medical image quality control method in this application; Figure 2 This is a flowchart illustrating the medical image quality control method in this application; Figure 3 This is a schematic diagram of the module connections of the medical image quality control system in this application. Detailed Implementation

[0024] The technical solutions of this application will be clearly and completely described below with reference to the embodiments. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0025] Please see Figures 1-2 The first aspect of this application provides a neural network-based intra-frame weighted medical image quality control method, comprising: Acquire medical imaging data, including 3D medical images and their corresponding evaluation systems; 3D medical images refer to CT, MRI, or X-ray scans, such as chest CT images and X-ray images; the evaluation system includes the quality control standards, category labels, and annotation labels corresponding to the 3D medical images used. Based on the evaluation system, several sequence data are obtained by data annotation of 3D medical images; the several sequence data are divided into training set, validation set and test set; The medical image quality control model is trained based on training, validation, and test sets. The training process includes data preprocessing, spatial feature extraction, temporal enhancement feature extraction, category-guided intra-frame weighted score calculation, model optimization, and class activation map generation. The medical image quality control model is used to perform quality control on 3D medical image data. The medical image quality control model is a model for image quality control composed of convolutional neural network models and other models in a set order. The recurrent neural network model can be replaced by network models with the same function, such as LSTM, GRU, and Mamba.

[0026] Understandably, quality control is a process based on the image itself, classifying images according to imaging expertise and relevant quality control standards to determine potential quality issues. Medical image quality control models, on the other hand, are based on imaging experts summarizing relevant quality control standards and experience into expressible knowledge, labeling images according to this knowledge, and then performing image classification tasks based on these labels. This method takes as input sequential images observable by the human eye, which is closer to the actual focus of experts; it is neither an independent image nor unprocessed volumetric data, thus differing from traditional 2D image input and 3D voxel data.

[0027] Based on the above technical solution, in the intra-frame weighted medical image quality control method and system for class-specific attention provided in this application, medical image data is acquired, including 3D medical images and their corresponding evaluation system; the 3D medical images are annotated according to the evaluation system to obtain several sequence data; the several sequence data are divided into training set, validation set and test set; the medical image quality control model is trained according to the training set, validation set and test set; the training process includes data preprocessing, spatial feature extraction, temporal enhancement feature extraction, category-guided intra-frame weighted score calculation, model optimization and class activation map generation; the category-guided intra-frame weighted score calculation includes calculating the intra-frame weighted score of each spatial location; by calculating the response intensity of each spatial location to a specific quality control category, class-specific weighted features are generated, thereby highlighting the region in each frame that has the strongest discriminative power for quality control judgment, and realizing the visualization of 3D medical image quality control through frame perception capability; effectively focusing on key anatomical regions and improving the model's sensitivity to subtle artifacts or positional deviations.

[0028] Meanwhile, by integrating the spatial feature extraction of CNN, the temporal modeling of RNN, and the region focusing capabilities of the category-guided intra-frame weighting mechanism, it achieves multi-dimensional, precise, and real-time quality assessment of medical images. Furthermore, through frame-level evidence extraction, it provides interpretable frame numbers and heatmaps while giving quality control labels. This not only significantly improves quality control efficiency and reduces the rate of repeated scans, but also makes it well applicable to various image types and devices.

[0029] In one possible implementation, the step of annotating 3D medical images based on an evaluation system to obtain several sequence data includes: extracting frames from the 3D medical images to obtain several frame images; integrating the frame images according to their corresponding frame order to generate sequence data; annotating the sequence data to obtain annotated sequence data; the data annotation includes the annotation of category labels and the annotation of annotation tags; the annotated sequence data includes several frame images and the categories of each spatial location, etc.; the frame order refers to the order in which the slice images of the 3D medical image data are sliced ​​along a specific spatial direction, such as axial, sagittal, or coronal, and are followed during the original scanning acquisition. Specifically, this embodiment refers to the Chest CT Image Quality Control Guidelines and has developed an evaluation system that includes four major categories of quality control standards. The corresponding category labels include the evaluation system for scanning range, image off-center, scanning position, and image artifacts. Each standard has a specific evaluation sub-category, i.e., a label. This embodiment sets up a total of 15 labels; Table 1 is the corresponding quality control standards, category labels and label correspondence table in the evaluation system of this embodiment, which lists the quality control standards and their corresponding labels in detail.

[0030] In this embodiment, the data annotation is performed by an annotation team of three or more radiology experts who independently annotate each tomographic sequence of each image. The annotation results are determined through multiple rounds of consensus. After the annotation is completed, the dataset can be divided in a custom way as needed. For example, the dataset can be randomly divided into a 70% training set, a 10% validation set, and a 20% test set.

[0031] Table 1: Correspondence between Quality Control Standards, Category Labels, and Marking Labels

[0032] In one possible implementation, the data preprocessing includes: extracting each frame image from the sequence data, converting each frame image to RGB mode, and sorting them according to the original frame order to obtain serialized data. Specifically, the input 3D sequence data is processed by reading DICOM data, i.e., medical digital imaging and communication data. In this embodiment, the DICOM data is CT image data, simulating the visual effect of a doctor directly viewing it on a PACS (picture archiving and communication system) workstation. The DICOM data is converted into a visualized RGB three-channel image according to the window width and window level values ​​recorded in the Tag values ​​in the DICOM data, and the sequence data is sorted according to the Instance Number and combined into directional serialized data.

[0033] ; Where T is the number of frames; H and W are the spatial resolutions; V is the extracted visual feature; R is the number of real number spaces to which the feature belongs; and C is the number of channels, such as three channels in RGB.

[0034] In one possible implementation, the spatial feature extraction includes: extracting frame images from the serialized data; using a convolutional neural network to extract spatial features from the frame images to obtain a spatial feature map of the frame images; flattening the spatial feature map according to spatial location to obtain spatial features corresponding to each spatial location; specifically, using a CNN to extract image features, for the features corresponding to each frame image in the 3D image data, a convolutional neural network, such as a classic model like ResNet, is used to extract the spatial features of that frame to obtain a feature map. ;in, d is the feature map corresponding to the t-th frame image; d is the feature dimension, such as 2048 in ResNet50; h and w are the spatial dimensions of the feature map, such as 7x7 in ResNet50; Flattening the above 3D feature map according to spatial location yields the feature set for all spatial locations in this frame; flattened features: ; in, Let be the spatial feature corresponding to the first spatial location in the t-th frame image; hw is the total number of spatial locations in the feature map; Let d be the feature set under feature dimension d.

[0035] In one possible implementation, temporal enhancement feature extraction includes: acquiring spatial features corresponding to each spatial location in each frame image; using a recurrent neural network (RNN) to perform temporal modeling on the spatial features of the same spatial location in different frame images to obtain the temporal enhancement features of the corresponding spatial location; and sequentially acquiring the temporal enhancement features of each spatial location. Specifically, to capture the temporal dependencies between video frames, a recurrent neural network (RNN), such as LSTM, GRU, and Mamba, can be introduced to perform temporal modeling on the spatial feature sequences of all frames, outputting the temporal enhancement features of each frame. One formula for calculating the temporal enhancement features is as follows:

[0036] in, The temporal enhancement feature for the j-th spatial location in the t-th frame image; The temporal enhancement feature is the spatial location of the j-th frame in the (t-1)-th frame image; Let be the spatial features of the j-th spatial location in the t-th frame image; it is understandable that the recurrent neural network used to extract temporal enhancement features can be a basic RNN, or a specific variant of models such as LSTM and GRU.

[0037] In this embodiment, the selective state-space model Mamba is selected as the time series modeler. It can maintain the linear complexity constant and has better parameter and computational cost than other RNN modules with the same hidden dimension.

[0038] In one possible implementation, the category-guided intra-frame weighted score calculation includes the following steps: setting a target category, obtaining temporal enhancement features at each spatial location of the frame image, and calculating the intra-frame weighted score of the target category at each spatial location of the frame image based on each temporal enhancement feature. ; in, Let be the intra-frame weighted score of target category i at the j-th spatial position in the t-th frame image, which is the probability of target category i appearing at the j-th spatial position in the t-th frame image, and ∈[0,1], used to measure the importance of different spatial locations for category recognition; Let be the classifier vector corresponding to target category i, and ∈ B is the set temperature parameter used to control the degree of sharpening and smoothing; it can be fixed during training. The temporal enhancement feature corresponding to the j-th spatial location of the t-th frame image; Let be the temporal enhancement feature corresponding to the k-th spatial location of the t-th frame image, where k∈[1, hw] and hw is the total number of spatial locations in the feature map.

[0039] In one possible implementation, model optimization includes: obtaining the intra-frame weighted scores of the target category at all spatial locations in each frame image, and the temporal enhancement features at all spatial locations in each frame image; generating a classification score for the target category based on the intra-frame weighted scores and temporal enhancement features at all spatial locations in each frame image. The classification result is determined based on the classification score, a set of labels is obtained, and the loss value is calculated based on the classification result and the label set. Specifically, the formula for calculating the loss value is as follows: ; Where L is the loss value; N is the total number of samples in the quality control image set of the training set. For the true label of the i-th sample, The model predicts the probability that the i-th sample is a positive sample.

[0040] If the loss value is less than a set threshold, the model training is completed; otherwise, the relevant parameters are adjusted and the training is repeated.

[0041] In one possible implementation, a classification score for the target category is generated based on the intra-frame weighted scores and temporal enhancement features at all spatial locations in each frame image. This includes: obtaining the intra-frame weighted scores corresponding to all spatial locations in the frame image, and performing a weighted summation of the temporal enhancement features corresponding to each spatial location based on the intra-frame weighted scores to obtain the frame-level class-specific features for the corresponding target category; specifically, combining the attention score, the temporal enhancement features at all spatial locations in each frame image are weighted and summed to obtain the frame-level class-specific features for category i of the frame image; the formula for obtaining the frame-level class-specific features is: ; in, For target feature i, it is a frame-specific feature in the t-th frame image; The intra-frame weighted score of target category i at the j-th spatial location in the t-th frame image; Let be the temporal enhancement feature corresponding to the j-th spatial location of the t-th frame image; and j∈[1, hw]; The global class-specific features of the target type are obtained by temporally fusing the frame-level class-specific features corresponding to all frames of the target image type. Specifically, the global class-specific features of the entire image sequence for class i are obtained by temporally fusing the frame-level class-specific features of all frames of the 3D image data (i.e., the image sequence). The temporal fusion formula is as follows: ; in, For target type i, the global class-specific feature in the image sequence; Global average pooling is performed on the temporal enhancement features at the spatial locations in each frame to obtain the background global features; specifically, to preserve the overall global information of the video, global average pooling is performed on the temporal enhancement features at all frames and all spatial locations to obtain background-like global features: the formula for calculating the background-like global features is as follows: ; Where g represents the global feature of the background class; Global class-specific features and background global features are fused through residual fusion to obtain the final features of the target class. Specifically, class-specific features and class-background global features are fused through residual connections, which preserves global information while highlighting key class-related information. The residual fusion formula is as follows: ; in, The final feature after fusion for target category i; These are weight parameters used to adjust the contribution of class-specific features in the fusion result; they can be learned during training or set manually. The final feature is multiplied by the classification vector of the target category to obtain the category score corresponding to the target category; specifically, the fused final feature of target category i is multiplied by the classifier vector of target category i to obtain the classification score of the image sequence belonging to category i. ; in, The classification score is the score of the image sequence belonging to target category i. The higher the score, the greater the probability that the sequence data belongs to that category. Let be the classifier vector for target category i at temperature B.

[0042] In one possible implementation, class activation map generation includes: extracting frame-level class-specific features corresponding to each frame image; and using the formula: ; The response intensity of the target category i corresponding to the t-th frame image is calculated. The larger the response intensity value, the more relevant the frame image is to the target category i; where, For frame-level class-specific features of target category i in the t-th frame image The frame-level class-specific feature sequence; and ; The frame image with the largest response value is selected as the keyframe image corresponding to the target category; specifically, this is done using the formula: ; Select keyframe images, where t The number corresponding to the keyframe image; Obtain the spatial feature map corresponding to the keyframe image; recalculate the spatial weight map using the class vector; perform bilinear interpolation to the original resolution, assign pseudo-color, and overlay it onto the original image to obtain the class activation map; specifically, the following steps are included: Step 1: Recalculate the spatial weight map using the category vector; the spatial feature map of keyframe image A is as follows. Where h×w is the feature map size, c is the number of channels, and the class vector is... Let i be the target category. The formula for calculating the spatial weight map is: ,in It is the k-th channel at position activation value, It is the weight of category i for the k-th channel.

[0043] Step 2: Bilinear interpolation to the original resolution; using upsampling, convert the spatial weight map M to the original image size to obtain a high-resolution heatmap. ; It is a single-channel grayscale image, and the value range is usually [0, 1]. Step 3: Apply pseudo-color; using a predefined color lookup table, convert the heatmap, i.e., the grayscale image, into a pseudo-color image. For example, red.

[0044] Step 4: Superimpose the pseudo-color image C onto the original image to obtain the class activation map; superimpose and fuse the pseudo-color image C and the original keyframe image to obtain the class activation map.

[0045] The final output includes the numerical values ​​corresponding to the keyframe images and the Base64 encoding of the CAM images, which can be displayed by the front-end browser without plugins.

[0046] Secondly, this application provides a class-specific attention intra-weighted medical image quality control device, comprising: a processor and a storage medium; the storage medium includes instructions, and the processor is configured to execute the instructions to implement the method described in the first aspect and any possible implementation thereof. This class-specific attention intra-weighted medical image quality control device may be an electronic device or a chip within an electronic device.

[0047] Please see Figure 3 Thirdly, this application provides a type of attention-specific intra-frame weighted medical image quality control system, including: a data acquisition module, a data partitioning module, and a model training module; Data acquisition module: used to acquire a number of medical image data; Data partitioning module: used to annotate 3D medical images based on the evaluation system to obtain several sequence data; and to partition the several sequence data into training set, validation set and test set; Model training module: The medical image quality control model is trained based on the training set, validation set and test set; the training process includes data preprocessing, spatial feature extraction, temporal enhancement feature extraction, category-guided intra-frame weighted score calculation, model optimization and class activation map generation.

[0048] Fourthly, this application provides a computer-readable storage medium storing instructions that, when executed on a class-specific attention intra-weighted medical image quality control device, cause the class-specific attention intra-weighted medical image quality control device to perform the methods described in the first aspect and any possible implementation thereof.

[0049] Fifthly, this application provides a computer program product containing instructions that, when run on a class-specific attention intra-weighted medical image quality control device, cause the class-specific attention intra-weighted medical image quality control device to perform the methods described in the first aspect and any possible implementation thereof.

[0050] Some of the data in the above formula are calculated by removing dimensions and taking their numerical values. The formula is the closest to the real situation obtained by software simulation of a large amount of collected data. The preset parameters and preset thresholds in the formula are set by those skilled in the art according to the actual situation or obtained through simulation of a large amount of data.

[0051] How this application works: By acquiring medical image data, including 3D medical images and their corresponding evaluation system, and annotating the 3D medical images according to the evaluation system to obtain several sequence data, the sequence data is divided into training set, validation set, and test set. The medical image quality control model is trained based on the training set, validation set, and test set. The training process includes data preprocessing, spatial feature extraction, temporal enhancement feature extraction, category-guided intra-frame weighted score calculation, model optimization, and class activation map generation. The category-guided intra-frame weighted score calculation includes calculating the intra-frame weighted score for each spatial location. By calculating the response intensity of each spatial location to a specific quality control category, class-specific weighted features are generated, thereby highlighting the region in each frame that has the strongest discriminative power for quality control judgment, and realizing the visualization of 3D medical image quality control through frame perception capability. It effectively focuses on key anatomical regions, enhancing the model's sensitivity to subtle artifacts or positional deviations. Simultaneously, by integrating CNN spatial feature extraction, RNN temporal modeling, and category-guided intra-frame weighting mechanisms for region focusing, it achieves multi-dimensional, precise, and real-time quality assessment of medical images. Furthermore, through frame-level evidence extraction, it provides interpretable frame numbers and heatmaps along with quality control labels. This not only significantly improves quality control efficiency and reduces duplicate scan rates but also adapts to various image types and devices.

[0052] The above embodiments are only used to illustrate the technical methods of this application and are not intended to limit it. Although this application has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical methods of this application without departing from the spirit and scope of the technical methods of this application.

Claims

1. A method for intra-frame weighted medical image quality control based on attention-specific features, characterized in that, include: Acquire medical imaging data, including 3D medical images and their corresponding evaluation system; the evaluation system includes the quality control standards, category labels, and annotation labels corresponding to the 3D medical images used. Based on the evaluation system, several sequence data are obtained by data annotation of 3D medical images; the several sequence data are divided into training set, validation set and test set; A medical image quality control model is trained based on training, validation, and test sets. The training process includes data preprocessing, spatial feature extraction, temporal enhancement feature extraction, category-guided intra-frame weighted score calculation, model optimization, and class activation map generation. The medical image quality control model is used for quality control of 3D medical images. The category-guided intra-weighted score calculation includes calculating the intra-weighted score for each spatial location.

2. The intra-frame weighted medical image quality control method based on specific attention as described in claim 1, characterized in that, The data annotation of 3D medical images based on the evaluation system yields several sequence data, including: The 3D medical image is frame-by-frame extracted to obtain several frame images. The frame images are then integrated according to their corresponding frame order to generate sequence data. The sequence data is labeled to obtain labeled sequence data. The data labeling includes the labeling of category labels and the labeling of annotation labels. The frame order is the order in which the slice images are obtained during the original scanning after the 3D medical image data is sliced ​​along a specific spatial direction.

3. The intra-frame weighted medical image quality control method based on specific attention as described in claim 1, characterized in that, The data preprocessing includes: Extract each frame image from the sequence data, convert each frame image to RGB mode, and sort them according to the original frame order to obtain serialized data.

4. The intra-frame weighted medical image quality control method based on specific attention as described in claim 3, characterized in that, The spatial feature extraction includes: Frame images are extracted from the serialized data, and spatial features are extracted from the frame images using a convolutional neural network to obtain the spatial feature map of the frame images; the spatial feature map is flattened according to spatial location to obtain the spatial features corresponding to each spatial location.

5. The intra-frame weighted medical image quality control method for specific attention as described in claim 4, characterized in that, The temporal enhancement feature extraction includes: Spatial features corresponding to each spatial location in each frame image are obtained. A recurrent neural network is used to perform temporal modeling of the spatial features of the same spatial location in different frame images, resulting in temporal enhancement features for the corresponding spatial locations. The temporal enhancement features for each spatial location are then obtained sequentially. One formula for calculating the temporal enhancement features is as follows: ; in, The temporal enhancement feature for the j-th spatial location in the t-th frame image; The temporal enhancement feature is the spatial location of the j-th frame in the (t-1)-th frame image; Let be the spatial feature of the j-th spatial location in the t-th frame image.

6. The intra-frame weighted medical image quality control method based on specific attention as described in claim 5, characterized in that, The category-guided intra-frame weighted score calculation includes the following steps: A target category is set, and temporal enhancement features at each spatial location of the frame image are obtained. Based on each temporal enhancement feature, the intra-frame weighted score of the target category at each spatial location in the frame image is calculated. ; in, The intra-frame weighted score of target category i at the j-th spatial location in the t-th frame image; B is the classifier vector corresponding to target category i; B is the set temperature parameter. The temporal enhancement feature corresponding to the j-th spatial location of the t-th frame image; Let be the temporal enhancement feature corresponding to the k-th spatial location of the t-th frame image, where k∈[1, hw] and hw is the total number of spatial locations in the feature map.

7. The intra-frame weighted medical image quality control method based on specific attention as described in claim 1, characterized in that, The model optimization includes: Obtain the intra-frame weighted score of the target category at all spatial locations in each frame image, as well as the temporal enhancement features at all spatial locations in each frame image; generate the classification score of the target category based on the intra-frame weighted score and temporal enhancement features at all spatial locations in each frame image. The classification result is determined based on the classification score, the set label set is obtained, and the loss value is calculated based on the classification result and the label set. If the loss value is less than a set threshold, the model training is completed; otherwise, the relevant parameters are adjusted and the training is repeated.

8. The intra-frame weighted medical image quality control method based on specific attention as described in claim 7, characterized in that, A classification score for the target category is generated based on the intra-frame weighted scores of all spatial locations in each frame image and temporal enhancement features, including: Obtain the intra-frame weighted score corresponding to all spatial locations in the frame image, and perform weighted summation on the temporal enhancement features corresponding to each spatial location based on the intra-frame weighted score corresponding to each spatial location to obtain the frame-level class-specific features of the corresponding target category; The global class-specific features of the target type are obtained by temporally fusing the frame-level class-specific features corresponding to all frame images of the target type. Global average pooling is performed on the temporal enhancement features of the spatial locations in each frame image to obtain the background global features; The global class-specific features and background global features are fused using residual magnitude to obtain the final features of the target class; The final feature is multiplied by the classification vector of the target category to obtain the category score corresponding to the target category.

9. The intra-frame weighted medical image quality control method based on specific attention as described in claim 1, characterized in that, The class activation graph generation includes: Extract frame-level class-specific features corresponding to each frame image; using the formula: ; The response intensity of the target category i corresponding to the t-th frame image is calculated. ;in, The frame-level class-specific feature sequence is composed of frame-level class-specific features of target category i in the t-th frame image; Select the frame image with the largest response intensity value as the keyframe image corresponding to the target category; Obtain the spatial feature map corresponding to the keyframe image; recalculate the spatial weight map using the class vector; perform bilinear interpolation to the original resolution, assign pseudo-color, and overlay it onto the original image to obtain the class activation map; including: Step 1: Recalculate the spatial weight map using the category vector; the spatial feature map of keyframe image A is as follows. Where h×w is the feature map size, c is the number of channels, and the class vector is... Let i be the target category. The formula for calculating the spatial weight map is: ,in It is the k-th channel at position activation value, It is the weight of category i with respect to the k-th channel; Step 2: Bilinear interpolation to the original resolution; using upsampling, convert the spatial weight map M to the original image size to obtain a high-resolution heatmap. H and W represent spatial resolution. Step 3: Apply pseudo-color; use a predefined color lookup table to convert the heatmap into a pseudo-color image. ; Step 4: Superimpose the pseudo-color image C onto the original image to obtain the class activation map; superimpose and fuse the pseudo-color image C and the original keyframe image to obtain the class activation map.

10. A type of attention-based intra-frame weighted medical image quality control system, applied to the type of attention-based intra-frame weighted medical image quality control method according to any one of claims 1 to 9, characterized in that, include: Data acquisition module, data partitioning module, and model training module; The data acquisition module is used to acquire several medical image data. The data partitioning module is used to annotate 3D medical images based on an evaluation system to obtain several sequence data; and to partition the several sequence data into a training set, a validation set, and a test set. The model training module trains the medical image quality control model based on the training set, validation set, and test set. The training process includes data preprocessing, spatial feature extraction, temporal enhancement feature extraction, category-guided intra-frame weighted score calculation, model optimization, and class activation map generation.

Citation Information

Patent Citations

  • CT image geometric artifact evaluation method based on residual network

    CN113870375A

  • Medical image classification method based on multi-modal feature fusion

    CN119339128A

  • Medical image definition quality control evaluation method

    CN119359629A

  • Medical image classification processing system and method based on artificial intelligence

    WO2019052063A1

Cited By

  • Medical image sequence positioning method and system

    CN121544625A