Heart magnetic resonance image detection method based on convolutional neural network
Through the cardiac magnetic resonance image detection method based on convolutional neural network, combined with wavelet transform denoising, feature map fusion and graph theory method optimization, the problem of low detection accuracy of image key points in the prior art is solved, and higher detection accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510205775.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In the prior art, the detection of image key points is problematic, especially in cardiac magnetic resonance images, noise, left ventricle shape differences and high similarity between the ventricle and peripheral tissues make it difficult to adapt to the detection algorithm, increasing the risk of false detection.
The cardiac magnetic resonance image detection method based on convolutional neural network is adopted to realize key point detection through wavelet transform denoising, feature map fusion, regression network bounding box prediction, segmentation network segmentation, graph theory method optimization and local maximum search.
It significantly reduces noise interference in MR images, improves the accuracy of key point detection, enhances the adaptability to the differences in left ventricular shapes of different patients, reduces the risk of false detection, and improves the reliability of key points.
Smart Images

Figure CN120047429A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image detection, and particularly to a method for detecting cardiac magnetic resonance images based on a convolutional neural network. Background Art
[0002] With the wide application of computed tomography (CT) and magnetic resonance imaging (MR) in disease diagnosis, treatment planning, and clinical research, computer-aided diagnosis (CAD) of medical images has become an important step in the daily work of doctors such as clinical diagnosis and determining treatment plans. And the detection of anatomical structure key points is an important research hotspot of this technology. The key point detection technology can help doctors quickly locate the positions of lesions, organs and other interested targets, improving the diagnosis efficiency. Although the key point detection algorithm has achieved good detection results in medical image processing, it still has certain limitations. In the prior art, the invention patent with the publication number CN111144486 B discloses a method for detecting key points of cardiac magnetic resonance images based on a convolutional neural network, which uses a deep learning model to process cardiac MR images to achieve the detection and segmentation of key points and improves the accuracy of left ventricular volume estimation by fusing information from different perspectives. However, there may be noise in MR images, and there are differences in the shape of the left ventricle: the differences in the shape of the left ventricle among different patients may make it difficult for the detection algorithm to adapt; the high similarity between the ventricle and surrounding tissues: the ventricle and surrounding tissues may be difficult to distinguish in the image, increasing the risk of false detection. In summary, there is a problem of low accuracy in the detection of image key points in the prior art. Summary of the Invention
[0003] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide a method for detecting cardiac magnetic resonance images based on a convolutional neural network, and the present invention solves the problem of low accuracy in the detection of image key points in the prior art.
[0004] To achieve the above object, the present invention provides the following solutions:
[0005] A method for detecting cardiac magnetic resonance images based on a convolutional neural network, comprising:
[0006] Collecting the cardiac MR image to be detected;
[0007] Denosing the MR image by using a wavelet transform algorithm and a wavelet inverse transform algorithm to obtain a denoised MR image;
[0008] Inputting the denoised MR image into a trained convolutional neural network to obtain a feature map of the intermediate layer;
[0009] Predicting the position of the bounding box of the key point for the denoised MR image by using a regression network algorithm to obtain the position of the bounding box;
[0010] Segment the MR image using a segmentation network to obtain a segmentation mask of the left ventricle;
[0011] Fuse the segmentation mask of the left ventricle with the feature map of the intermediate layer to obtain a comprehensive feature map;
[0012] Use graph theory methods and the position of the bounding box to optimize the key points of the comprehensive feature map to obtain an optimized image;
[0013] Perform a local maximum search on the optimized image to obtain the key point detection result.
[0014] Preferably, the using of the wavelet transform algorithm and the inverse wavelet transform algorithm to denoise the MR image to obtain a denoised MR image includes:
[0015] Perform a wavelet transform on the MR image to decompose the MR image into a high-frequency image and a low-frequency image;
[0016] Process the low-frequency image with a non-local means denoising algorithm to obtain a low-frequency denoised image;
[0017] Perform local pixel grouping on the low-frequency denoised image to obtain a sample set of similar local pixel blocks;
[0018] Perform a fast Fourier transform on the sample set to obtain a reconstructed low-frequency image;
[0019] Perform block decomposition on the high-frequency image to obtain a high-frequency block image;
[0020] Input the high-frequency block image into a denoising autoencoder to obtain a reconstructed high-frequency block image and aggregate it to obtain a reconstructed high-frequency image;
[0021] Use the inverse wavelet transform algorithm to combine the reconstructed high-frequency image and the reconstructed low-frequency image to obtain a denoised MR image.
[0022] Preferably, the performing of the fast Fourier transform on the sample set to obtain a reconstructed low-frequency image includes:
[0023] Perform a fast Fourier transform on each pixel block in the sample set to obtain a corresponding frequency domain set;
[0024] Perform spectral data analysis on each frequency domain in the frequency domain set to obtain a frequency domain feature set;
[0025] According to a preset threshold, screen the frequency domain feature set to obtain a target frequency domain set;
[0026] Perform an inverse fast Fourier transform on the target frequency domain set to obtain a reconstructed low-frequency block image;
[0027] Aggregate the reconstructed low-frequency fast images to obtain the reconstructed low-frequency images.
[0028] Preferably, the expression of the sample set is:
[0029] S = {s i | i = 1, 2, …, N};
[0030] where S is the sample set, s i is the pixel block, i is the label of the pixel block, and N is a natural number;
[0031] The feature of the pixel block is represented as:
[0032] f i = CNN(s i )
[0033] where f i is the feature representation of the pixel block, and the corresponding feature extraction time of the pixel block is;
[0034] The calculation expression in the frequency domain is:
[0035]
[0036] where M is the pixel block size, k is the frequency index, i.e., the frequency distribution, and n is the sample index;
[0037] The calculation expression of the target frequency domain is:
[0038]
[0039] where is the target frequency domain, w f (k) is the adaptive weight, σ is the standard deviation of the noise, T(k; σ) is the dynamic threshold function, is the indicator function;
[0040] The calculation expression of the reconstructed low-frequency block image is:
[0041]
[0042] where w i is the weight of the correlation of the reconstructed block;
[0043] The calculation expression of the reconstructed low-frequency image is:
[0044]
[0045] where w a (i) is the aggregation weight, and φ(f i ) is the feature weighting function.
[0046] Preferably, inputting the high-frequency block image into the denoising autoencoder, obtaining the reconstructed high-frequency block image and aggregating it to obtain the reconstructed high-frequency image includes:
[0047] Extracting the high-frequency features of the high-frequency block image by using the denoising autoencoder;
[0048] Obtaining the reconstructed high-frequency block image according to the high-frequency features of the high-frequency block image;
[0049] Performing weighting and aggregation on the reconstructed high-frequency block image to obtain the initial reconstructed high-frequency image;
[0050] Performing global optimization on the initial reconstructed high-frequency image to obtain the final reconstructed high-frequency image.
[0051] Preferably, the calculation expression of the final reconstructed high-frequency image is:
[0052]
[0053] where arg min H is to minimize the objective function, H is the matrix of the high-frequency image, H(j) is the pixel value of the j-th pixel in the reconstructed high-frequency image, X is the total number of pixels, E is the feature extraction function, D is the denoising autoencoder function, H i is the input high-frequency block image, is the reconstructed high-frequency block image, W is the weight calculation function, and λ is the regularization hyperparameter.
[0054] Preferably, using the regression network algorithm to predict the bounding box position of the key points for the denoised MR image to obtain the bounding box position includes:
[0055] Collecting and annotating the denoised MR image to obtain the annotated image training set;
[0056] Setting the loss function, combining the feature maps of the intermediate layers, and training the regression network according to the annotated image training set to obtain the trained regression network;
[0057] Using the trained regression network to predict the bounding box position of the key points for the denoised MR image to obtain the initial prediction result;
[0058] Performing non-maximum suppression algorithm calculation on the initial prediction result to obtain the bounding box position.
[0059] Preferably, using the graph theory method and the bounding box position to optimize the key points of the comprehensive feature map to obtain the optimized image includes:
[0060] Determining the candidate key point positions according to the bounding box position;
[0061] Construct a graph theory model, where the nodes of the graph theory model are candidate key points;
[0062] Use the shortest path algorithm to select the optimal key point path and obtain the optimization result;
[0063] Based on the optimization result, determine the optimized key points based on the candidate key points;
[0064] Determine the optimized image according to the optimized key points.
[0065] The present invention discloses the following technical effects:
[0066] The present invention provides a method for detecting cardiac magnetic resonance images based on a convolutional neural network, including: collecting a cardiac MR image to be detected; denoising the MR image using a wavelet transform algorithm and an inverse wavelet transform algorithm to obtain a denoised MR image; inputting the denoised MR image into a trained convolutional neural network to obtain a feature map of an intermediate layer; using a regression network algorithm to predict the bounding box position of key points for the denoised MR image to obtain the bounding box position; using a segmentation network to segment the MR image to obtain a segmentation mask of the left ventricle; fusing the segmentation mask of the left ventricle with the feature map of the intermediate layer to obtain a comprehensive feature map; using a graph theory method and the bounding box position to optimize the key points of the comprehensive feature map to obtain an optimized image; performing a local maximum search on the optimized image to obtain the key point detection result. The denoising process of the present invention significantly reduces the noise interference in the MR image, enhances the usability of the image, and thus improves the accuracy of key point detection. Through feature map fusion and multi-level feature extraction, the adaptability of the algorithm to the shape differences of the left ventricles of different patients is enhanced, and the risk of false detection is reduced. Using a graph theory method to optimize key point localization effectively reduces the confusion caused by the similarity between the ventricle and surrounding tissues and improves the reliability of the key points. Description of the Drawings
[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0068] Figure 1 It is a flowchart of a method for detecting cardiac magnetic resonance images based on a convolutional neural network provided by an embodiment of the present invention;
[0069] Figure 2 It is a flowchart of denoising based on a convolutional neural network provided by an embodiment of the present invention;
[0070] Figure 3 A flowchart for calculating the position of a bounding box based on a convolutional neural network provided by an embodiment of the present invention. Detailed implementation manners
[0071] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0072] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.
[0073] As Figure 1 shown, the present invention provides a method for detecting cardiac magnetic resonance images based on a convolutional neural network, including:
[0074] Step 100: Collect the cardiac MR image to be detected;
[0075] Step 200: Denoise the MR image using the wavelet transform algorithm and the inverse wavelet transform algorithm to obtain the denoised MR image;
[0076] Step 300: Input the denoised MR image into the trained convolutional neural network to obtain the feature map of the intermediate layer;
[0077] Step 400: Use the regression network algorithm to predict the position of the bounding box of the key points for the denoised MR image to obtain the position of the bounding box;
[0078] Step 500: Segment the MR image using the segmentation network to obtain the segmentation mask of the left ventricle;
[0079] Specifically, input the MR image to be segmented into the trained segmentation network for inference.
[0080] Generate a segmentation mask: The network outputs a segmentation mask of the same size as the input image, where the value of each pixel represents the probability that the position belongs to the left ventricle. Usually, a threshold (such as 0.5) can be selected to convert the probability map into a binary image.
[0081] Perform morphological operations (such as dilation and erosion) on the generated binary segmentation mask to remove noise and small isolated regions, thereby optimizing the final segmentation result.
[0082] Step 600: Fuse the segmentation mask of the left ventricle with the feature map of the intermediate layer to obtain a comprehensive feature map;
[0083] Step 700: Optimize the key points of the comprehensive feature map by using graph theory method and the position of the bounding box to obtain an optimized image;
[0084] Step 800: Perform a local maximum search on the optimized image to obtain the key point detection result.
[0085] Further, as Figure 2 shown, denoising the MR image by using the wavelet transform algorithm and the inverse wavelet transform algorithm to obtain a denoised MR image, including:
[0086] Step 201: Perform wavelet transform on the MR image to decompose the MR image into a high-frequency image and a low-frequency image;
[0087] Step 202: Process the low-frequency image with a non-local means denoising algorithm to obtain a low-frequency denoised image;
[0088] Step 203: Perform local pixel grouping on the low-frequency denoised image to obtain a sample set of similar local pixel blocks;
[0089] Step 204: Perform fast Fourier transform on the sample set to obtain a reconstructed low-frequency image;
[0090] Step 205: Decompose the high-frequency image into high-frequency block images;
[0091] Step 206: Input the high-frequency block image into a denoising autoencoder to obtain a reconstructed high-frequency block image and aggregate it to obtain a reconstructed high-frequency image;
[0092] Step 207: Use the inverse wavelet transform algorithm to combine the reconstructed high-frequency image and the reconstructed low-frequency image to obtain a denoised MR image.
[0093] Specifically, wavelet transform decomposition: Apply wavelet transform to the cardiac MR image to be processed, and decompose it into a high-frequency image (capturing details and edges) and a low-frequency image (containing main structural information and background). This decomposition helps to more effectively identify and remove image noise.
[0094] Low-frequency image processing: Apply a non-local means denoising algorithm to the low-frequency image. Considering the similarity of pixels in the image, find the pixels similar to each pixel in its neighborhood, and use these similar pixels for weighted averaging to achieve a better denoising effect.
[0095] Local pixel grouping: Perform local pixel grouping on the processed low-frequency denoised image, and collect similar local pixel blocks into a sample set for subsequent processing. This process can be achieved by calculating the similarity metric (such as mean square error) between local blocks to ensure effective aggregation of similar pixel blocks.
[0096] Fast Fourier Transform Application: For each pixel block in the sample set, perform the Fast Fourier Transform (FFT) to obtain its frequency domain representation. This process can capture the frequency characteristics of the image and provide necessary information for denoising and reconstruction.
[0097] High-frequency Image Block Decomposition: Perform block decomposition on the high-frequency image to divide it into multiple smaller blocks for subsequent processing. This operation helps analyze detailed features and lays a foundation for subsequent high-frequency information reconstruction.
[0098] Input to the High-frequency Block Denoising Autoencoder: Input the high-frequency block image into the denoising autoencoder, which is specially trained to learn and reconstruct useful high-frequency features in the input block. Through this step, noise in the high-frequency block can be effectively removed without losing important edge information.
[0099] Aggregation of High-frequency Images: Reconstructed high-frequency blocks are output from the denoising autoencoder, and these blocks are aggregated to form the final reconstructed high-frequency image. This step ensures that the features of each block can be globally and consistently combined to improve the image quality.
[0100] Inverse Wavelet Transform: Combine the reconstructed high-frequency image and the low-frequency image denoised by non-local means, and use the inverse wavelet transform algorithm to combine them into the final denoised MR image. This step can restore the denoised structural information and the retained detailed information to the spatial domain through the inverse transform in the wavelet domain, generating a clear and detailed MR image.
[0101] Furthermore, performing the Fast Fourier Transform on the sample set to obtain the reconstructed low-frequency image includes:
[0102] Perform the Fast Fourier Transform on each pixel block in the sample set to obtain the corresponding frequency domain set;
[0103] Perform spectral data analysis on each frequency domain in the frequency domain set to obtain the frequency domain feature set;
[0104] According to a preset threshold, screen the frequency domain feature set to obtain the target frequency domain set;
[0105] Perform the inverse Fast Fourier Transform on the target frequency domain set to obtain the reconstructed low-frequency block image;
[0106] Aggregate the reconstructed low-frequency block images to obtain the reconstructed low-frequency image.
[0107] Specifically, apply the Fast Fourier Transform (FFT) to each pixel block in the sample set to transform it from the spatial domain to the frequency domain. This process uses the FFT algorithm to efficiently process the pixel values of each block and obtain the corresponding frequency-domain representation. The FFT result of each block contains amplitude and phase information, usually stored in complex form; perform spectral data analysis on each frequency domain in the obtained frequency-domain set to calculate the frequency-domain feature set. Common analysis methods include calculating the amplitude spectrum and phase spectrum of each frequency-domain component; screen the frequency-domain feature set according to a preset threshold (which can be a certain proportion or a specific value of the amplitude), and extract the frequency components above the threshold to form the target frequency-domain set. This can be achieved through conditional judgment. For example, verify whether the frequency-domain value is greater than the set threshold and save the frequency components that meet the conditions; apply the Inverse Fast Fourier Transform (IFFT) to the target frequency-domain set to convert the frequency-domain information back to the spatial domain to obtain the reconstructed low-frequency block image; aggregate all the reconstructed low-frequency block images and merge them into a complete reconstructed low-frequency image. This can be achieved by splicing, averaging, or placing the blocks in the appropriate positions according to the spatial position information. This step ensures the consistency of all reconstructed blocks in size and position to form a continuous and seamless low-frequency image.
[0108] Specifically, the expression of the sample set is:
[0109] S = {s i | i = 1, 2, …, N};
[0110] where S is the sample set, s i is the pixel block, i is the label of the pixel block, and N is a natural number;
[0111] The feature representation of the pixel block is:
[0112] f i = CNN(s i )
[0113] where f i is the feature representation of the pixel block, and the corresponding feature extraction time of the pixel block is;
[0114] The calculation expression of the frequency domain is:
[0115]
[0116] where M is the pixel block size, k is the frequency index, that is, the frequency distribution, which refers to the different frequency components in the signal when performing the Fourier transform. Its value range is usually from 0 to M - 1, used to specify the frequency distribution in the frequency domain. n is the sample index, usually used to represent a certain specific point in the time series or signal. In the inverse Fourier transform, it is used to represent the spatial domain coordinates for restoring the image
[0117] The calculation expression of the target frequency domain is as follows:
[0118]
[0119] Among them, is the target frequency domain, w f (k) is the adaptive weight, σ is the standard deviation of the noise, and T(k;σ) is the dynamic threshold function. is the indicator function;
[0120] The calculation expression of the reconstructed low-frequency block image is as follows:
[0121]
[0122] Among them, w i is the weight of the correlation of the reconstructed block;
[0123] The calculation expression of the reconstructed low-frequency image is as follows:
[0124]
[0125] Among them, w a (i) is the aggregation weight, and φ(f i ) is the feature weighting function.
[0126] Furthermore, inputting the high-frequency block image into the denoising autoencoder to obtain the reconstructed high-frequency block image and performing aggregation to obtain the reconstructed high-frequency image includes:
[0127] Extracting the high-frequency features of the high-frequency block image by using the denoising autoencoder;
[0128] Obtaining the reconstructed high-frequency block image according to the high-frequency features of the high-frequency block image;
[0129] Performing weighting and aggregation on the reconstructed high-frequency block image to obtain the initial reconstructed high-frequency image;
[0130] Performing global optimization on the initial reconstructed high-frequency image to obtain the final reconstructed high-frequency image.
[0131] Specifically, the high-frequency feature extraction process is introduced as follows:
[0132] (1) Input the high-frequency block image: Input each high-frequency block image into the pre-trained denoising autoencoder (DAE). The autoencoder consists of two main parts: an encoder and a decoder.
[0133] (2) Encoder processing: In the denoising autoencoder, the encoder part gradually extracts the features of the high-frequency block image through multiple convolutional layers. This process includes downsampling the image to reduce the dimension. During this process, the model attempts to retain the key information in the image while suppressing noise.
[0134] (3) Application of activation function: Use appropriate activation functions (such as ReLU or LeakyReLU) in each layer of the encoder to increase the non-linearity of the network and help the model learn complex features.
[0135] The reconstruction of the high-frequency block image is introduced as follows:
[0136] (1) Decoder part: The high-frequency features extracted by the encoder are processed by the decoder and gradually reconstructed back into a reconstructed high-frequency block image of the same size as the original high-frequency block image. The decoder usually uses upsampling (such as transposed convolution or interpolation) to convert the low-dimensional feature map back to the high-dimensional space of the original image.
[0137] (2) Output correction: In the last layer of the decoder, activation functions such as Sigmoid or Tanh can be used to limit the output value within a reasonable range (such as 0 - 1 or -1 to 1) to ensure the validity and usability of the image data.
[0138] The weighting and aggregation process is as follows:
[0139] (1) Reconstruction block weighting: After obtaining multiple reconstructed high-frequency block images, weights can be assigned to each reconstructed block according to the importance of these blocks in the original high-frequency image. For example, weights can be set according to the brightness, texture information of the block or its correlation with other blocks.
[0140] (2) Aggregation operation: The reconstructed high-frequency block images are weighted and averaged according to their weights and merged into an initial reconstructed high-frequency image. This process can effectively integrate the characteristic information of each high-frequency block and improve the overall image quality.
[0141] Global optimization is introduced as follows:
[0142] (1) Application of global optimization algorithm: Apply a global optimization algorithm (such as Stochastic Gradient Descent (SGD) or Adam optimization) to the initial reconstructed high-frequency image to adjust the pixel values to make the image more consistent. This step aims to eliminate possible inflection points and incoherences and comprehensively improve the visual effect of the image.
[0143] (2) Loss function setting: Set an appropriate loss function (such as mean squared error or Structural Similarity Index (SSIM)) to evaluate the difference between the optimized image and the target image to guide the optimization process.
[0144] (3) Optimization iteration: Update the image through multiple iterations to ensure that the output reconstructed high-frequency image reaches the optimal state, enhancing details and features.
[0145] Furthermore, the calculation expression of the final reconstructed high-frequency image is:
[0146]
[0147] where arg min H is to minimize the objective function, H is the matrix of the high-frequency image, H(j) is the pixel value of the j-th pixel in the reconstructed high-frequency image, X is the total number of pixels, E is the feature extraction function, D is the denoising autoencoder function, H i is the input high-frequency block image, is the reconstructed high-frequency block image, W is the weight calculation function, and λ is the regularization hyperparameter.
[0148] Furthermore, as Figure 3 shown, using the regression network algorithm to predict the bounding box position of the key points for the denoised MR image to obtain the bounding box position includes:
[0149] Step 401: Collect and annotate the denoised MR image to obtain an annotated image training set;
[0150] Step 402: Set the loss function, combine the feature maps of the intermediate layers, and train the regression network according to the annotated image training set to obtain a trained regression network;
[0151] Step 403: Use the trained regression network to predict the bounding box position of the key points for the denoised MR image to obtain an initial prediction result;
[0152] Step 404: Perform non-maximum suppression algorithm calculation on the initial prediction result to obtain the bounding box position.
[0153] Specifically, the construction process of the annotated image training set is as follows:
[0154] (1) MR image collection: Collect a group of denoised MR images to ensure sample diversity, including different patients, different angles, and image qualities.
[0155] (2) Image annotation: Invite medical professionals to accurately annotate the key points in the MR image, and define the corresponding bounding box (Bounding Box) for each key point. Each bounding box is represented by the coordinates of the upper left corner and the lower right corner. The annotation results form an annotated image training set, which is used as the training data for the regression network.
[0156] The process of setting the loss function is as follows:
[0157] (1) Loss function design: Select a loss function suitable for the regression task, such as mean squared error (MSE) or SmoothL1 loss. This loss function is used to measure the error between the predicted bounding box and the actual annotated bounding box.
[0158] (2) Combine feature maps: When designing the network, combine the intermediate layer feature maps of the denoised MR image with the input image to enhance the feature representation. This step can use feature fusion (such as concatenation or weighted average) to provide more context information.
[0159] The training process of the regression network is as follows:
[0160] (1) Network design and construction: Construct a deep regression network (for example, using a convolutional neural network as the basis), which can learn the context and spatial information of key point detection from the feature maps.
[0161] (2) Training process: Train the network with the annotated image training set, optimize the loss function, and make the model better fit the data. In the training process, use batch gradient descent (such as Adam or SGD optimizer) and learning rate adjustment strategy to improve the convergence speed.
[0162] The prediction process of the bounding box position is as follows:
[0163] (1) Input the denoised MR image: Use the trained regression network to predict the bounding box position of the denoised MR image to obtain the initial prediction result. The model will output the bounding box coordinates of each key point.
[0164] (2) Representation of the prediction result: The prediction result is usually presented in the form of the center coordinates, width, and height of each bounding box, indicating the predicted key point positions.
[0165] Non-maximum suppression (NMS):
[0166] Application of the NMS algorithm: Apply the non-maximum suppression algorithm to the initial prediction result to avoid repeated detection of key points at similar positions. NMS deletes the bounding boxes with too high overlap by setting a threshold (such as the IOU threshold), and only retains the bounding box with the highest confidence.
[0167] Bounding box optimization: After NMS processing, the finally output bounding box position is the optimized key point bounding box, realizing the accurate positioning of key points.
[0168] Furthermore, the use of the graph theory method and the bounding box position to optimize the key points of the comprehensive feature map to obtain an optimized image includes:
[0169] Determine the candidate key point positions according to the bounding box position;
[0170] Construct a graph theory model, where the nodes of the graph theory model are candidate key points;
[0171] Use the shortest path algorithm to select the best key point path and obtain the optimization result;
[0172] Based on the optimization result, determine the optimized key points based on the candidate key points;
[0173] Determine the optimized image according to the optimized key points.
[0174] Specifically, the process of determining the candidate key point positions is as follows:
[0175] (1) Bounding box information parsing: Use the bounding box positions obtained from the regression network to extract the preliminary positions of each key point in the bounding box. Each bounding box contains one or more candidate key point positions (usually the center position of the border or predefined relative positions).
[0176] (2) Generation of candidate point sets: According to the bounding box positions, generate multiple candidate key points through regular grid generation or neighboring pixel methods to form a candidate key point set. These candidate points should cover the possible key point positions within the bounding box.
[0177] The process of constructing the graph theory model is as follows:
[0178] (1) Node definition: Regard the candidate key points as the nodes of the graph and create a graph theory model. Each candidate key point corresponds to a node, and a unique identifier is assigned to each node.
[0179] (2) Edge construction: Define the edges between the nodes. The weights of the edges are usually related to the spatial distance, feature similarity, or image features (such as gradient information) between the candidate key points. The smaller the weight, the more similar the candidate points are in features or the closer the distance, forming the connectivity of the graph.
[0180] The selection of the shortest path algorithm is as follows:
[0181] (1) Application of the shortest path algorithm: Use the shortest path algorithm (such as Dijkstra's algorithm or A* search algorithm) to find the best path among the candidate key points. The goal is to find the path that can best represent the relationship and features of the key points, while considering the cost (i.e., weight) of the path.
[0182] (2) Path optimization result: The path returned by the shortest path algorithm gives the optimization result, identifying the best set of key points, which reflects the optimal connection relationship between the nodes.
[0183] The process of determining the optimized key points is as follows:
[0184] (1) Extract optimization key points: Based on the optimization results obtained from the shortest path algorithm, corresponding optimization key points are selected from the candidate key point set. These points are the key points for the final optimization and can better reflect the image features and their interrelationships.
[0185] (2) Weight assignment: Importance scores can be assigned to the optimization key points according to the weights of the edges to further assist subsequent processing.
[0186] The process of optimizing image generation is as follows:
[0187] (1) Image reconstruction: Use the optimization key points to generate an optimized image. This step can refill the pixels in the image between the optimization key points through interpolation methods (such as bilinear interpolation or cubic interpolation) to make the image structure more coherent and complete.
[0188] (2) Post - processing techniques: Adjust the image details according to the positions of the optimization key points, which may include operations such as edge smoothing and noise removal to further improve the image quality and obtain the final optimized image.
[0189] The various embodiments in this specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other.
[0190] Specific examples are used in this article to elaborate on the principles and implementation methods of the present invention. The descriptions of the above embodiments are only used to help understand the method of the present invention and its core idea; at the same time, for those of ordinary skill in the art, based on the idea of the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A cardiac magnetic resonance image detection method based on convolutional neural network, characterized in that: include: Acquiring a cardiac MR image to be detected; De-noising the MR image using a wavelet transform algorithm and an inverse wavelet transform algorithm to obtain a de-noised MR image; Inputting the denoised MR image into a trained convolutional neural network to obtain a feature map of an intermediate layer; Using a regression network algorithm to predict the bounding box position of key points on the denoised MR image to obtain the bounding box position; Segmenting the MR image using a segmentation network to obtain a segmentation mask of the left ventricle; fusing the segmentation mask of the left ventricle with the feature map of the intermediate layer to obtain a comprehensive feature map; Optimizing key points of the comprehensive feature map using a graph theory method and the position of the bounding box to obtain an optimized image; A local maximum search is performed on the optimized image to obtain a key point detection result.
2. A cardiac magnetic resonance image detection method based on convolutional neural network according to claim 1, characterized in that: The denoising of the MR image by using a wavelet transform algorithm and an inverse wavelet transform algorithm to obtain a denoised MR image includes: Performing wavelet transform on the MR image to decompose the MR image into a high-frequency image and a low-frequency image; Processing the low-frequency image using a non-local mean denoising algorithm to obtain a low-frequency denoised image; Performing local pixel grouping processing on the low-frequency denoised image to obtain a sample set of similar local pixel blocks; Performing fast Fourier transform on the sample set to obtain a reconstructed low-frequency image; Decomposing the high-frequency image into blocks to obtain a high-frequency block image; Inputting the high-frequency block image into a denoising autoencoder to obtain a reconstructed high-frequency block image and aggregating the image to obtain a reconstructed high-frequency image; The reconstructed high-frequency image and the reconstructed low-frequency image are combined by using an inverse wavelet transform algorithm to obtain a denoised MR image.
3. The cardiac magnetic resonance image detection method based on convolutional neural network according to claim 1, characterized in that: The step of performing a fast Fourier transform on the sample set to obtain a reconstructed low-frequency image includes: Performing a fast Fourier transform on each pixel block in the sample set to obtain a corresponding frequency domain set; Performing spectrum data analysis on each frequency domain in the frequency domain set to obtain a frequency domain feature set; According to a preset threshold, the frequency domain feature set is screened to obtain a target frequency domain set; Performing an inverse fast Fourier transform on the target frequency domain set to obtain a reconstructed low-frequency block image; The reconstructed low-frequency fast image is aggregated to obtain a reconstructed low-frequency image.
4. The cardiac magnetic resonance image detection method based on convolutional neural network according to claim 3, characterized in that: The expression of the sample set is: S={s i ∣i=1,2,…,N}; Among them, S is the sample set, s i is a pixel block, i is the number of the pixel block, and N is a natural number; The characteristic representation of the pixel block is: f i =CNN(s i ) Among them, f i is the feature representation of the pixel block, and the feature extraction time of the pixel block corresponds to; The calculation expression in the frequency domain is: Among them, M is the pixel block size, k is the frequency index, that is, the frequency distribution, and n is the sample index. The calculation expression of the target frequency domain is: in, is the target frequency domain, w f (k) is the adaptive weight, σ is the standard deviation of the noise, T(k;σ) is the dynamic threshold function, is the indicator function; The calculation expression for reconstructing the low-frequency block image is: Among them, w i is the weight of the relevance of the reconstructed block; The calculation expression for reconstructing the low-frequency image is: Among them, w a (i) is the aggregation weight, φ(f i ) is the feature weighting function.
5. The cardiac magnetic resonance image detection method based on convolutional neural network according to claim 3, characterized in that: The step of inputting the high-frequency block image into a denoising autoencoder to obtain a reconstructed high-frequency block image and aggregating the image to obtain a reconstructed high-frequency image includes: Extracting high-frequency features of the high-frequency block image using a denoising autoencoder; Obtaining a reconstructed high-frequency block image according to the high-frequency features of the high-frequency block image; Weighting and aggregating the reconstructed high-frequency block image to obtain an initial reconstructed high-frequency image; The initial reconstructed high-frequency image is globally optimized to obtain a final reconstructed high-frequency image.
6. A cardiac magnetic resonance image detection method based on convolutional neural network according to claim 5, characterized in that: The calculation expression of the final reconstructed high-frequency image is: Among them, arg min H To minimize the objective function, H is the matrix of the high-frequency image, H(j) is the pixel value of the jth pixel in the reconstructed high-frequency image, X is the total number of pixels, E is the feature extraction function, D is the denoising autoencoder function, and H i is the input high-frequency block image, To reconstruct the high-frequency block image, W is the weight calculation function and λ is the regularization hyperparameter.
7. The cardiac magnetic resonance image detection method based on convolutional neural network according to claim 1, characterized in that: The method of using a regression network algorithm to predict the bounding box position of key points on the denoised MR image to obtain the bounding box position includes: Acquiring and annotating the denoised MR images to obtain an annotated image training set; Setting a loss function, combining the feature graph of the intermediate layer, and training the regression network according to the labeled image training set to obtain a trained regression network; Using the trained regression network to perform bounding box position prediction on the denoised MR image, to obtain an initial prediction result; The initial prediction result is calculated using a non-maximum suppression algorithm to obtain a bounding box position.
8. The cardiac magnetic resonance image detection method based on convolutional neural network according to claim 1, characterized in that: The step of optimizing the key points of the comprehensive feature map by using the graph theory method and the bounding box position to obtain an optimized image includes: Determine the candidate key point position according to the bounding box position; Constructing a graph theory model, wherein the graph theory model nodes are candidate key points; The shortest path algorithm is used to select the best key point path and obtain the optimization result; According to the optimization result, determining the optimization key points based on the candidate key points; An optimized image is determined according to the optimization key points.
Citation Information
Patent Citations
Keypoint Detection Method for Cardiac MRI Images Based on Convolutional Neural Networks
CN111144486B