Medical image automatic diagnosis method and system based on deep learning
Through a two-channel convolutional neural network based on deep learning and a knowledge graph-driven diagnostic inference engine, the problem of difficult fusion of multimodal medical image data is solved, and efficient and accurate lesion diagnosis and diagnostic suggestions are achieved, improving diagnostic efficiency and safety.
Patent Information
- Application Number
- CN202510445044.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing medical imaging diagnosis technology faces the problems of difficulty in direct fusion of multimodal imaging data and incomplete feature extraction, which leads to high diagnostic complexity, low efficiency, poor accuracy and consistency, and lack of effective diagnostic reasoning mechanisms.
Using a deep learning-based method, spatial and dynamic sequence features are extracted through a dual-channel convolutional neural network, combined with cross-modal feature fusion, adaptive attention weight allocation and cascading long and short-term memory network, we model the lesion evolution law, and use a probability graph model and a knowledge graph-driven diagnostic inference engine to generate diagnostic suggestions.
It improves the accuracy and efficiency of medical imaging diagnosis, provides comprehensive and structured diagnostic references, reduces the probability of misdiagnosis and missed diagnosis, and ensures the security and privacy of data through multimodal imaging acquisition, distributed computing and security audits.
Smart Images

Figure CN120340784A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of medical image diagnosis, and particularly to an automatic medical image diagnosis method and system based on deep learning. Background Art
[0002] Medical image diagnosis is a crucial link in modern medicine. Doctors rely on medical images to detect diseases, judge the development of the condition, and formulate treatment plans. However, traditional medical image diagnosis faces many challenges.
[0003] In terms of data acquisition, various modalities of medical images, such as CT images, MRI images, and ultrasound images, have extremely different data formats due to different imaging principles. For CT images obtained by different devices, there are differences in Hounsfield units, and the problem of spatial distortion is also relatively prominent; the multi-sequence intensities of MRI images lack a unified standard, and motion artifacts often interfere with diagnosis; ultrasound images have problems such as a large dynamic range, a lot of speckle noise, and insufficient tissue boundary clarity. These differences make it difficult to directly analyze and fuse image data effectively, increasing the complexity of diagnosis.
[0004] From the perspective of the diagnosis process, manual diagnosis relies on doctors' professional experience and knowledge reserves. The amount of medical image data is huge and complex. Doctors need to spend a lot of time and energy carefully observing and analyzing image details, which is not only inefficient but also easily leads to doctor fatigue, thereby affecting the accuracy of diagnosis. Moreover, the diagnostic levels of different doctors vary, and for some difficult cases, the diagnostic results may vary greatly, making it difficult to ensure the consistency and reliability of diagnosis.
[0005] With the development of deep learning technology, although some studies have tried to apply it to medical image diagnosis, there are still many deficiencies. Existing deep learning models are not comprehensive enough in feature extraction and are difficult to simultaneously consider the spatial features and dynamic sequence features of medical images. In cross-modal feature fusion, the advantages of different modality images cannot be fully integrated, resulting in information loss. And the ability to capture the evolution law of lesions is limited and cannot well reflect the development and changes of diseases in the time dimension. In terms of abnormal detection and segmentation, the segmentation accuracy and efficiency need to be improved, and the boundaries of abnormal regions cannot be accurately marked and precise lesion segmentation masks cannot be generated. In addition, there is a lack of an effective diagnostic reasoning mechanism and comprehensive and structured diagnostic suggestions cannot be provided for doctors. Summary of the Invention
[0006] The purpose of the present invention is to provide an automatic medical image diagnosis method and system based on deep learning to solve the problems raised in the above background art.
[0007] To achieve the above object, the present invention provides the following technical solutions: A method for automatic medical image diagnosis based on deep learning, the method comprising:
[0008] Obtain multi-modal medical image data of a target object, including CT images, MRI images and ultrasound images, and perform format standardization processing on the medical image data;
[0009] Construct an image feature extraction module based on a dual-channel convolutional neural network, where the first channel uses a three-dimensional convolutional kernel to extract spatial features, and the second channel uses a temporal convolutional kernel to extract dynamic sequence features, and perform parallel feature extraction on the standardized medical image data;
[0010] Establish a cross-modal feature fusion layer, perform cross-dimensional splicing on the spatial features and dynamic sequence features extracted through the dual channels to generate a fused feature map;
[0011] Through an adaptive attention weight allocation mechanism, dynamically adjust the feature weights of different regions in the fused feature map to highlight lesion-related features;
[0012] Use a cascaded bidirectional long short-term memory network to perform temporal dependence modeling on the adjusted feature sequence to capture the evolution law of lesions;
[0013] Design an anomaly detection module based on a probabilistic graphical model, generate a potential lesion probability distribution map according to the temporal dependence modeling result, and mark the boundary of the abnormal region;
[0014] Integrate a multi-scale context information enhancement module to perform local refinement segmentation on the boundary of the abnormal region, and generate a final lesion segmentation mask in combination with global semantic constraints;
[0015] Construct a knowledge graph-driven diagnostic inference engine, associate and map the lesion segmentation result with a preset medical entity relationship graph to generate a structured diagnostic suggestion.
[0016] Preferably, the format standardization processing specifically includes:
[0017] Perform Hounsfield unit standardization on the obtained CT images, and use a non-rigid registration algorithm to eliminate spatial distortion between different scanning devices;
[0018] Perform multi-sequence intensity normalization processing on MRI images, and eliminate motion artifacts through an artifact suppression model based on a generative adversarial network;
[0019] Perform dynamic range compression and speckle noise suppression processing on ultrasound images, and use a phase consistency edge enhancement algorithm to improve the clarity of tissue boundaries.
[0020] Preferably, the first channel of the dual-channel convolutional neural network adopts the following structure:
[0021] Embed a deformable convolutional layer in the three-dimensional convolutional kernel to adjust the receptive field shape of the convolutional kernel through adaptive deformation parameters;
[0022] Connect a channel recalibration module after each layer of convolution to dynamically adjust the activation weights of each feature channel;
[0023] Adopt a pyramid pooling structure to extract multi-resolution spatial features at multiple scales;
[0024] The calculation formula for the activation weight of the channel recalibration module is:
[0025] s c = σ(W·F c + b)
[0026] In the formula, s c represents the activation weight scalar value of the c-th feature channel, σ is the sigmoid activation function, is the learnable weight matrix, is the bias term, is the input feature map, and C, H, W, D represent the number of channels, height, width, and depth respectively.
[0027] Preferably, the implementation method of the cross-modal feature fusion layer includes:
[0028] Perform a tensor expansion operation on the spatial feature matrix and the dynamic sequence feature matrix, and perform interleaved splicing along the channel dimension;
[0029] Insert a learnable position encoding vector into the spliced fusion matrix to retain the original spatial position information of different modal features;
[0030] Adopt a sparse constraint regularization method to reduce the dimension of the fusion matrix and eliminate redundant feature correlations.
[0031] Preferably, the adaptive attention weight allocation mechanism is specifically:
[0032] Construct a spatial-channel dual attention subnetwork to calculate the significance score of each spatial position and the importance coefficient of each channel in the feature map respectively;
[0033] Perform a Hadamard product operation on the spatial significance score and the channel importance coefficient to generate a composite attention weight matrix;
[0034] Dynamically update the weight matrix through a gated recurrent unit, and adjust the current attention distribution according to the historical feature state;
[0035] The expression of the Hadamard product operation is:
[0036] A i,j,k = S i,j ⊙ C k
[0037] In the formula, A i,j,k represents the composite attention weight of the fused feature map at the spatial position (i, j) and channel k, is the spatial saliency score matrix, is the channel importance coefficient vector, ⊙ represents the element-wise multiplication operation, and i, j, k respectively represent the height, width, and channel dimension indices.
[0038] Preferably, the cascaded bidirectional long short-term memory network includes:
[0039] The first-level network extracts local temporal features in a sliding window manner and generates multi-scale temporal feature segments;
[0040] The second-level network performs cross-window correlation analysis on the temporal feature segments to capture long-range dependencies;
[0041] A residual skip connection is set at the output end of each level of the network to perform weighted fusion of the original input features and the network learning features.
[0042] Preferably, the construction method of the probability graph model includes:
[0043] Using a conditional random field to model the spatial constraint relationship between adjacent pixels and defining the continuity prior probability of the lesion area;
[0044] Using a variational autoencoder to generate a latent feature distribution and aligning the latent space with the real data distribution through the KL divergence constraint;
[0045] Jointly optimizing the prior probability and the latent distribution to generate a probability heat map with spatial consistency.
[0046] Preferably, the execution steps of the multi-scale context information enhancement module include:
[0047] Extracting multi-level regions of interest at the boundary of the abnormal area and using dilated convolutions with different dilation rates to extract local detail features;
[0048] Constructing a graph convolutional network to model the global semantic relationship and performing information transfer between the local detail features and the global semantic nodes;
[0049] Filtering cross-scale features through a feature distillation mechanism and retaining the key features highly relevant to the lesion anatomical structure.
[0050] Preferably, the workflow of the knowledge graph-driven diagnostic inference engine includes:
[0051] Match the anatomical location and morphological parameters in the lesion segmentation result with the disease-anatomy mapping rules in the medical entity atlas;
[0052] Derive the potential complications and risks of secondary lesions based on the pathological causal chain in the atlas;
[0053] Use a fuzzy logic inference engine to handle uncertain diagnosis situations and generate multiple candidate diagnosis schemes with confidence scores.
[0054] Preferably, the present invention further includes a system for implementing the above-mentioned automatic medical image diagnosis method based on deep learning, and the system includes:
[0055] Data preprocessing module: used to obtain multi-modal medical image data of the target object, including CT images, MRI images and ultrasound images, and perform format standardization processing on the medical image data;
[0056] Image feature extraction module: construct a dual-channel convolutional neural network, where the first channel uses a three-dimensional convolutional kernel to extract spatial features, and the second channel uses a temporal convolutional kernel to extract dynamic sequence features, and perform parallel feature extraction on the standardized medical image data;
[0057] Cross-modal feature fusion layer: perform cross-dimensional splicing on the spatial features and dynamic sequence features extracted through the dual channels to generate a fused feature map;
[0058] Adaptive attention weight allocation module: dynamically adjust the feature weights of different regions in the fused feature map through an adaptive attention weight allocation mechanism to highlight the lesion-related features;
[0059] Temporal dependence modeling module: use a cascaded bidirectional long short-term memory network to perform temporal dependence modeling on the adjusted feature sequence to capture the evolution law of the lesion;
[0060] Anomaly detection module: design an anomaly detection module based on a probabilistic graph model, generate a potential lesion probability distribution map according to the temporal dependence modeling result, and mark the boundary of the abnormal area;
[0061] Multi-scale context information enhancement module: integrate a multi-scale context information enhancement module, perform local refinement segmentation on the boundary of the abnormal area, and generate a final lesion segmentation mask in combination with global semantic constraints;
[0062] Diagnostic inference engine: construct a knowledge graph-driven diagnostic inference engine, associate and map the lesion segmentation result with a preset medical entity relationship graph, and generate a structured diagnostic suggestion.
[0063] Compared with the prior art, the beneficial effects of the present invention are:
[0064] In the data processing stage, specialized format standardization methods are adopted according to the characteristics of different modality medical images. For CT images, Hounsfield unit standardization and non-rigid registration algorithms are used to unify density representation and eliminate spatial distortion, making subsequent analysis more accurate. For MRI images, multi-sequence intensity normalization and artifact suppression based on generative adversarial networks improve image quality and avoid interference from motion artifacts in diagnosis. For ultrasound images, dynamic range compression, speckle noise suppression, and phase consistency edge enhancement algorithms enhance image details and tissue boundary clarity, providing a reliable data basis for accurate diagnosis.
[0065] In terms of feature extraction and fusion, an image feature extraction module based on a dual-channel convolutional neural network extracts spatial and dynamic sequence features in parallel through 3D convolutional kernels and temporal convolutional kernels, capable of comprehensively capturing key information in medical images. The cross-modal feature fusion layer splices different modality features across dimensions, inserts position encoding vectors, and uses sparse constraint regularization for dimensionality reduction, which not only preserves the original spatial position information but also eliminates redundant features, making the fused feature map more valuable.
[0066] The adaptive attention weight allocation mechanism dynamically adjusts the feature weights of different regions in the fused feature map, highlighting lesion-related features and effectively avoiding interference from other irrelevant information, greatly improving the accuracy and pertinence of diagnosis. The cascaded bidirectional long short-term memory network models the temporal dependence of the feature sequence, captures the evolution law of lesions, provides doctors with dynamic information on disease development, helps to understand the condition more comprehensively, and formulate a more scientific treatment plan.
[0067] The anomaly detection module based on the probabilistic graph model generates a potential lesion probability distribution map and marks the boundaries of abnormal regions, providing a clear target for subsequent segmentation. The multi-scale context information enhancement module combines local refined segmentation and global semantic constraints to generate the final lesion segmentation mask, with high segmentation accuracy, capable of accurately outlining the lesion range and assisting doctors to more intuitively observe the lesion morphology.
[0068] The knowledge graph-driven diagnostic reasoning engine correlates and maps the lesion segmentation results with the medical entity relationship graph to generate structured diagnostic suggestions. By matching anatomical positions and morphological parameters, inferring potential complications and the risk of secondary lesions, and using a fuzzy logic inference engine to handle uncertain diagnostic situations, it provides doctors with comprehensive and scientific diagnostic references, reduces the probability of misdiagnosis and missed diagnosis, and improves the diagnostic efficiency and quality.
[0069] In addition, the medical image automatic diagnosis system of the present invention is equipped with a multi-modal image acquisition interface module, a distributed feature calculation cluster, a visualization interaction terminal, and a security audit module. The multi-modal image acquisition interface module facilitates the access of DICOM protocol data streams from different medical imaging devices, enabling rapid data acquisition; the distributed feature calculation cluster is configured with parallel computing nodes accelerated by GPUs, greatly improving the inference speed of deep learning models; the visualization interaction terminal provides a three-dimensional lesion reconstruction view and an editable interface for diagnostic reports, facilitating intuitive observation and operation by doctors; the security audit module records the operation logs of all data processing processes and implements a differential privacy protection mechanism, ensuring the security and privacy of data. BRIEF DESCRIPTION OF THE DRAWINGS
[0070] Figure 1 is the working principle diagram of the medical image automatic diagnosis method of the present invention;
[0071] Figure 2 is the flowchart of the implementation of the cross-modal feature fusion layer;
[0072] Figure 3 is the working flowchart of the adaptive attention weight allocation mechanism;
[0073] Figure 4 is the working flowchart of the knowledge graph-driven diagnostic inference engine. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0074] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0075] Please refer to Figures 1-4 , the present invention provides a technical solution: a medical image automatic diagnosis method based on deep learning, specifically including the following steps:
[0076] Medical image data acquisition and format standardization processing: Obtain multi-modal medical image data of the target object, including CT images, MRI images, and ultrasound images. To ensure the accuracy and consistency of subsequent processing, perform format standardization processing on these medical image data.
[0077] Construct an image feature extraction module: Construct an image feature extraction module based on a dual-channel convolutional neural network. The first channel of this module uses three-dimensional convolutional kernels to extract spatial features; the second channel uses temporal convolutional kernels to extract dynamic sequence features. Through these two channels, parallel feature extraction is performed on the standardized medical image data, enabling the full exploration of key information in the images from different dimensions.
[0078] Establish a cross-modal feature fusion layer: Establish a cross-modal feature fusion layer to perform cross-dimensional splicing on the spatial features and dynamic sequence features extracted through the dual channels. This splicing method can integrate the advantages of different modal features, generate a fused feature map, and provide a more comprehensive information basis for subsequent analysis.
[0079] Adaptive attention weight assignment: Through an adaptive attention weight assignment mechanism, dynamically adjust the feature weights of different regions in the fused feature map. This mechanism can highlight lesion-related features, making subsequent analysis more focused on areas that may have lesions and improving the accuracy of diagnosis.
[0080] Temporal dependence modeling: Use a cascaded bidirectional long short-term memory network to perform temporal dependence modeling on the adjusted feature sequence. This network can capture the evolution law of lesions, taking into account the development and changes of diseases in the time dimension, and providing more dynamic and forward-looking information for diagnosis.
[0081] Anomaly detection and marking: Design an anomaly detection module based on a probabilistic graphical model to generate a probability distribution map of potential lesions according to the results of temporal dependence modeling. This distribution map intuitively shows the probability distribution of potential lesions and marks the boundaries of abnormal regions, providing a clear target area for subsequent segmentation and diagnosis.
[0082] Lesion segmentation and mask generation: Integrate a multi-scale context information enhancement module to perform local refinement segmentation on the boundaries of abnormal regions. At the same time, combine global semantic constraints to generate the final lesion segmentation mask, accurately outlining the scope of the lesions.
[0083] Diagnostic reasoning and recommendation generation: Construct a knowledge graph-driven diagnostic reasoning engine to associate and map the lesion segmentation results with a pre-set medical entity relationship graph. Through this mapping, structured diagnostic recommendations can be generated based on medical knowledge and logic, providing valuable diagnostic references for doctors.
[0084] The present invention will be further described below in conjunction with Embodiments 1 to 6:
[0085] Embodiment 1:
[0086] In the format standardization process of medical image data, specific processing methods were adopted for medical images of different modalities.
[0087] For CT images, Hounsfield unit normalization is performed first. Since different CT scanning devices may have differences in measuring the Hounsfield values of the same substance during imaging, through normalization processing, the Hounsfield units of all CT images are unified to a standard scale, making the density representations of CT images obtained by different devices consistent. After that, a non-rigid registration algorithm is used to eliminate spatial distortions between different scanning devices. The non-rigid registration algorithm can flexibly deform and adjust the images. By establishing the corresponding relationships between images, it corrects the spatial position deviations and shape deformations caused by factors such as different scanning devices and scanning angle differences, ensuring the accurate spatial positions of organs, tissues and other structures in CT images and laying a foundation for subsequent precise feature extraction and diagnostic analysis.
[0088] For MRI images, multi-sequence intensity normalization processing is implemented. MRI imaging includes multiple sequences, such as T1-weighted images, T2-weighted images, etc. The intensity distributions of images in different sequences are different. Performing multi-sequence intensity normalization enables the intensities of images in different sequences to be compared and analyzed on the same scale, enhancing the consistency of image features. At the same time, an artifact suppression model based on a generative adversarial network is used to eliminate motion artifacts. Motion artifacts are relatively common in MRI imaging, mainly caused by involuntary movements of patients during scanning. The generative adversarial network consists of a generator and a discriminator. The generator is responsible for generating images after removing artifacts, and the discriminator judges whether the generated images are real. Through the adversarial training of the two, it can effectively identify and remove motion artifacts in MRI images and improve the image quality.
[0089] For ultrasound images, dynamic range compression and speckle noise suppression processing are performed. In ultrasound images, the dynamic range of the signal is large, which makes it difficult to observe image details. Through dynamic range compression, the range of gray values of the image is adjusted to highlight the details of the region of interest. Speckle noise is a unique noise in ultrasound imaging, which reduces the clarity and contrast of the image. Appropriate filtering algorithms are used to suppress speckle noise and improve the image quality. In addition, the phase consistency edge enhancement algorithm is used to enhance the clarity of tissue boundaries. The phase consistency algorithm can detect the phase changes of different frequency components in the image, highlighting the true edge information in the image and making the boundaries of tissues in ultrasound images clearer, facilitating subsequent feature extraction and diagnostic analysis.
[0090] Example 2:
[0091] The first channel of the dual-channel convolutional neural network adopts a series of optimized structures to improve the feature extraction ability. A deformable convolutional layer is embedded in the three-dimensional convolutional kernel. The key of the deformable convolutional layer lies in its adaptive deformation parameters. These parameters can dynamically adjust the receptive field shape of the convolutional kernel according to the content of the input image. For example, when facing irregularly shaped lesions, the deformable convolutional layer can make the receptive field of the convolutional kernel better fit the shape of the lesion, so as to more accurately extract the spatial features of the lesion. Compared with the traditional convolutional kernel with a fixed shape, it greatly improves the ability to capture features of targets with complex shapes.
[0092] After each layer of convolution, a channel recalibration module is connected. This module highlights important features and suppresses redundant features by dynamically adjusting the activation weights of each feature channel. The calculation formula of the activation weight is: s c = σ(W·F c + b). Among them, s c represents the activation weight scalar value of the c-th feature channel, which determines the importance of this channel in subsequent processing. σ is the sigmoid activation function, which maps the input value to the range between 0 and 1, playing the role of normalization and non-linear transformation, making the weight value interpretable and adjustable. is a learnable weight matrix, where C represents the number of channels. This matrix is continuously learned and adjusted through network training to adapt to different image features and achieve weighted fusion of features of each channel. is the bias term, which adds a learnable constant to the calculation of each channel, helping the model better fit the data. is the input feature map, where H, W, and D respectively represent the height, width, and depth of the feature map, representing the spatial dimension information of the input data, which is the object processed by the channel recalibration module.
[0093] In addition, a pyramid pooling structure is adopted to extract multi-resolution spatial features at multiple scales. The pyramid pooling structure performs pooling operations on the input feature map through pooling kernels of different sizes to obtain feature representations at different scales. Small pooling kernels can retain the detailed information of the image, while large pooling kernels can obtain the global information of the image. By fusing these features at different scales, the network can simultaneously take into account the local details and global features of the image, more comprehensively extract the spatial features in medical images, and provide rich information for subsequent diagnostic analysis.
[0094] Example 3:
[0095] The cross-modal feature fusion layer adopts a variety of technical means in implementation. First, a tensor unfolding operation is performed on the spatial feature matrix and the dynamic sequence feature matrix, and they are interleaved and spliced along the channel dimension. The spatial feature matrix contains information such as the spatial position and shape of objects in medical images, while the dynamic sequence feature matrix reflects the changes in the images in the time dimension. Through tensor unfolding, these two feature matrices are transformed into a form suitable for splicing, and then interleaved and spliced along the channel dimension, enabling the features of different modalities to be fused with each other in the channel dimension, forming a new fusion matrix. This interleaved splicing method can fully preserve the original information of the two-modal features and make them complement each other during the fusion process.
[0096] Learnable position encoding vectors are inserted into the spliced fusion matrix. The purpose of this step is to retain the original spatial position information of different-modal features. The position encoding vectors assign unique encodings to each position. After inserting these vectors into the fusion matrix, the network can identify and utilize the spatial position relationships of the features based on this encoding information during subsequent processing, avoiding the loss of important spatial position information during the feature fusion process and ensuring the integrity and accuracy of the features.
[0097] A sparse constraint regularization method is used to reduce the dimension of the fusion matrix and eliminate the correlation of redundant features. Medical image data usually contains a large number of features, and some of these features may be redundant or highly correlated, which not only increases the computational load but may also affect the performance of the model. The sparse constraint regularization method constrains the fusion matrix, making the weights of some unimportant features approach 0, thereby achieving the purpose of dimensionality reduction. At the same time, this method can effectively eliminate the correlation between redundant features, improve the quality and independence of the features, and make the subsequent diagnostic analysis more efficient and accurate.
[0098] Example 4:
[0099] As a key link in improving the diagnostic accuracy of the present invention, the adaptive attention weight allocation mechanism realizes the dynamic adjustment of the feature weights in different regions of the fused feature map by constructing a spatial-channel dual attention sub-network, and precisely highlights the features related to the lesion.
[0100] When constructing the spatial-channel dual attention sub-network, its main task is to calculate the saliency scores of each spatial position and the importance coefficients of each channel in the feature map respectively. For the calculation of spatial saliency scores, the network analyzes each spatial position in the feature map. In medical images, lesion regions often have unique features such as texture and shape, which make the spatial positions where the lesions are located more prominent in the whole map. The network assigns higher spatial saliency scores to the lesion regions by learning these features. For example, in a lung CT image, if there is a tumor lesion, the texture of the local region where the tumor is located is significantly different from that of the surrounding normal lung tissue, and the network will recognize this difference and assign a higher spatial saliency score to this region.
[0101] The calculation of channel importance coefficients is based on the importance of the features represented by different channels for diagnosis. Different channels may represent different information in medical images. For example, some channels focus on reflecting the density information of tissues, while others are more sensitive to the edge information of tissues. In actual diagnosis, for different diseases, certain specific channel information is more crucial. Taking the diagnosis of liver diseases as an example, the channels that reflect the density changes of liver tissues have higher importance for detecting diseases such as liver cysts and liver tumors, and the network will accordingly assign higher importance coefficients to these channels.
[0102] After calculating the spatial saliency scores and channel importance coefficients, the Hadamard product operation is used to generate a composite attention weight matrix, and its operation expression is:
[0103] A i,j,k =S i,j ⊙C k
[0104] In this formula, A i,j,k represents the composite attention weight at the spatial position (i, j) and channel k of the fused feature map. It combines the importance information of both the spatial position and the channel dimensions and is the key parameter finally used to adjust the feature weights. Among them, is the spatial saliency score matrix. H and W represent the height and width of the space respectively. Each element in the matrix represents the saliency score of the corresponding spatial position, reflecting the importance of this position in the whole space. For example, in a medical image feature map with a size of H×W, S i,j represents the spatial saliency score at the position with coordinates (i, j). Let \(C\) be the channel importance coefficient vector, where \(K\) represents the number of channels. Each element in the vector represents the importance coefficient of the corresponding channel, reflecting the importance degree of this channel among all channels. \(\odot\) represents the element-wise multiplication operation. Through this operation mode, the spatial saliency score matrix \(S\) and the channel importance coefficient vector \(C\) are multiplied element by element, combining the importance of spatial positions and channels, thereby generating a composite attention weight matrix that can more accurately reflect the comprehensive importance of different regions and channels.
[0105] To further optimize the attention weights, the present invention uses a gated recurrent unit to dynamically update the weight matrix. The gated recurrent unit has a memory function and can remember the previous feature state information. During the medical image diagnosis process, as time goes by or different image sequences are analyzed, the characteristics of the lesion may change. The gated recurrent unit will dynamically adjust the weight matrix according to the current input features and the previously remembered historical feature states. For example, when analyzing the dynamic MRI images of the heart, the shape and function of the heart change in different cardiac cycles. The gated recurrent unit can adjust the attention distribution at the current moment according to the feature states of the previous cardiac cycles, enabling the model to pay more attention to the key feature changes of the heart in the current state, thereby more accurately highlighting the lesion features related to the current condition and greatly improving the accuracy and timeliness of diagnosis.
[0106] Example 5:
[0107] The first-level network of the cascaded bidirectional long short-term memory network extracts local temporal features in a sliding window manner and generates multi-scale temporal feature segments. The sliding window slides on the time series, extracting features within a certain time range each time. By adjusting the window size and step length, local temporal features with different time resolutions can be obtained. Generating multi-scale temporal feature segments can observe the changes of the lesion from different time scales, such as rapid changes on a short time scale and slow evolution on a long time scale, providing more comprehensive temporal information for subsequent analysis.
[0108] The second-level network performs cross-window correlation analysis on the temporal feature segments generated by the first-level network to capture long-range dependencies. Through this cross-window analysis, the network can discover the correlations between different time windows and capture the evolution trends of the lesion over a long time span, such as the development process of the disease and the changes in treatment effects.
[0109] At the output end of each network level, a residual skip connection is set up to perform weighted fusion of the original input features and the network-learned features. The residual skip connection can avoid the problem of gradient vanishing or gradient explosion when the network depth increases. At the same time, it directly transmits the original input features to subsequent layers, so that the network can learn new features without losing the original important information. Through weighted fusion, the relative importance of the original features and the learned features can be adjusted according to the requirements of different tasks, improving the performance of the network.
[0110] The construction method of the probabilistic graph model includes multiple key steps. A conditional random field is used to model the spatial constraint relationship between adjacent pixels, and the prior probability of the continuity of the lesion area is defined. The conditional random field takes into account the adjacent relationship between pixels. By constructing an energy function, the probability that adjacent pixels have similar labels (i.e., belonging to the same lesion area or non-lesion area) is higher, thus ensuring the spatial continuity of the lesion area and avoiding unreasonable segmentation results.
[0111] A variational autoencoder is used to generate the latent feature distribution, and the KL divergence is used to constrain the alignment of the latent space and the real data distribution. The variational autoencoder consists of an encoder and a decoder. The encoder maps the input data to the latent space, and the decoder reconstructs the data from the latent space. The KL divergence is used to measure the difference between the latent space distribution and the real data distribution. By minimizing the KL divergence, the latent space can better reflect the feature distribution of the real data, improving the generation ability and generalization performance of the model.
[0112] The prior probability and the latent distribution are jointly optimized to generate a probability heat map with spatial consistency. Through joint optimization, the prior probability defined by the conditional random field and the latent feature distribution generated by the variational autoencoder are combined, so that the generated probability heat map can not only reflect the spatial continuity of the lesion area but also embody the latent features of the data, thus more accurately showing the probability distribution of potential lesions.
[0113] Example 6:
[0114] When the multi-scale context information enhancement module performs local refinement segmentation on the boundary of the abnormal area, the following steps are executed. Multiple levels of regions of interest are extracted at the boundary of the abnormal area, and dilated convolutions with different dilation rates are used to extract local detail features respectively. Dilated convolution can expand the receptive field of the convolution kernel without increasing the number of parameters and computational complexity by introducing holes in the convolution kernel. Dilated convolutions with different dilation rates can obtain local detail information at different scales. Dilated convolutions with smaller dilation rates can capture finer edge details, while dilated convolutions with larger dilation rates can obtain a larger range of context information. By fusing these local detail features at different scales, the boundary of the abnormal area can be described more comprehensively.
[0115] Construct a graph convolutional network to model the global semantic relationships and perform information transfer between local detailed features and global semantic nodes. The graph convolutional network represents the semantic relationships in medical images in the form of a graph, where nodes represent different regions or features, and edges represent the relationships between them. Through graph convolutional operations, information can be propagated globally, integrating local detailed features with global semantic information, so that the segmentation results not only consider the local image features but also incorporate the overall semantic background, improving the accuracy and rationality of segmentation.
[0116] Screen cross-scale features through a feature distillation mechanism and retain the key features highly relevant to the anatomical structure of the lesion. The feature distillation mechanism extracts important features from complex feature sets by comparing features at different scales, removing redundant and unimportant features. This can reduce the dimensionality of features, improve computational efficiency, and at the same time retain the key features closely related to the anatomical structure of the lesion, making the final segmentation results more in line with medical anatomical principles and providing strong support for accurate diagnosis.
[0117] The workflow of the knowledge graph-driven diagnostic inference engine includes multiple key steps. Match the anatomical locations and morphological parameters in the lesion segmentation results with the disease-anatomy mapping rules in the medical entity graph. A large amount of medical knowledge, including the corresponding relationships between various diseases and anatomical structures, is pre-stored in the medical entity graph. By matching the anatomical locations and morphological parameters in the segmentation results with the rules in the graph, the possible disease types can be initially judged.
[0118] Deduce the potential complications and risks of secondary lesions based on the pathological causal relationship chain in the graph. The pathological causal relationship chain describes the development and evolution relationships between diseases, such as other complications or secondary lesions that a certain disease may cause. Using these relationships, on the basis of the initial diagnosis, further analyze the potential health risks that the patient may face, providing a reference for formulating a comprehensive treatment plan.
[0119] Adopt a fuzzy logic inference engine to handle uncertain diagnostic situations and generate multiple candidate diagnostic solutions with confidence scores. In medical diagnosis, due to the complexity and uncertainty of imaging data, there may be multiple diagnostic possibilities. The fuzzy logic inference engine can handle this uncertainty, reason and evaluate various diagnostic possibilities according to different evidences and rules, generate multiple candidate diagnostic solutions, and assign confidence scores to each solution, helping doctors understand the diagnostic situation more comprehensively and make more accurate decisions.
[0120] It should be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device.
[0121] Although the embodiments of the present invention have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention, and the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An automatic medical image diagnosis method based on deep learning, characterized in that, It includes the following steps: Obtain multi-modal medical image data of the target object, including CT images, MRI images, and ultrasound images, and perform format standardization processing on the medical image data; Construct an image feature extraction module based on a dual-channel convolutional neural network, where the first channel uses a three-dimensional convolutional kernel to extract spatial features, and the second channel uses a temporal convolutional kernel to extract dynamic sequence features, and perform parallel feature extraction on the standardized medical image data; Establish a cross-modal feature fusion layer, perform cross-dimensional splicing on the spatial features and dynamic sequence features extracted through the dual channels to generate a fused feature map; Through an adaptive attention weight allocation mechanism, dynamically adjust the feature weights of different regions in the fused feature map to highlight lesion-related features; Use a cascaded bidirectional long short-term memory network to perform temporal dependence modeling on the adjusted feature sequence to capture the evolution law of lesions; Design an anomaly detection module based on a probabilistic graph model, generate a potential lesion probability distribution map according to the temporal dependence modeling result, and mark the boundary of the abnormal region; Integrate a multi-scale context information enhancement module to perform local refinement segmentation on the boundary of the abnormal region, and generate a final lesion segmentation mask in combination with global semantic constraints; Construct a knowledge graph-driven diagnostic inference engine, associate and map the lesion segmentation result with a preset medical entity relationship graph to generate a structured diagnostic recommendation.
2. The medical image automatic diagnosis method according to claim 1, characterized in that The format standardization processing specifically includes: Perform Hounsfield unit standardization on the obtained CT images, and use a non-rigid registration algorithm to eliminate spatial distortion between different scanning devices; Perform multi-sequence intensity normalization processing on MRI images, and eliminate motion artifacts through an artifact suppression model based on a generative adversarial network; Perform dynamic range compression and speckle noise suppression processing on ultrasound images, and use a phase consistency edge enhancement algorithm to improve the clarity of tissue boundaries.
3. The medical image automatic diagnosis method according to claim 2, wherein The first channel of the dual-channel convolutional neural network adopts the following structure: Embed a deformable convolutional layer in the three-dimensional convolutional kernel, and adjust the receptive field shape of the convolutional kernel through adaptive deformation parameters; Connect a channel recalibration module after each layer of convolution to dynamically adjust the activation weights of each feature channel; Adopt a pyramid pooling structure to extract multi-resolution spatial features at multiple scales; The calculation formula for the activation weight of the channel recalibration module is: s c = σ(W·F c + b) where s c represents the activation weight scalar value of the c-th feature channel, σ is the sigmoid activation function, is the learnable weight matrix, is the bias term, is the input feature map, and C, H, W, D represent the number of channels, height, width, and depth respectively.
4. The medical image automatic diagnosis method according to claim 1, characterized in that The implementation method of the cross-modal feature fusion layer includes: Perform tensor expansion operation on the spatial feature matrix and the dynamic sequence feature matrix, and perform interleaved splicing along the channel dimension; Insert a learnable position encoding vector into the spliced fusion matrix to retain the original spatial position information of different modality features; Use a sparse constraint regularization method to reduce the dimension of the fusion matrix and eliminate redundant feature correlations.
5. The medical image automatic diagnosis method according to claim 4, characterized in that The adaptive attention weight allocation mechanism is specifically: Construct a spatial-channel dual attention sub-network to calculate the saliency score of each spatial position in the feature map and the importance coefficient of each channel respectively; Perform Hadamard product operation on the spatial saliency score and the channel importance coefficient to generate a composite attention weight matrix; Dynamically update the weight matrix through a gated recurrent unit, and adjust the current attention distribution according to the historical feature state; The Hadamard product operation expression is as follows: A i,j,k = S i,j ⊙C k Where, A i,j,k represents the composite attention weight of the fused feature map at the spatial position (i, j) and channel k, is the spatial saliency score matrix, is the channel importance coefficient vector, ⊙ represents the element-wise multiplication operation, and i, j, k represent the height, width, and channel dimension indices respectively.
6. The medical image automatic diagnosis method according to claim 1, characterized in that The cascaded bidirectional long short-term memory network includes: The first-level network extracts local temporal features in a sliding window manner and generates multi-scale temporal feature segments; The second-level network performs cross-window correlation analysis on the temporal feature segments to capture long-range dependencies; At the output end of each level of the network, a residual skip connection is set to perform weighted fusion of the original input features and the network learning features.
7. The medical image automatic diagnosis method according to claim 1, characterized in that The construction method of the probabilistic graphical model includes: Using a conditional random field to model the spatial constraint relationship between adjacent pixels and defining the continuity prior probability of the lesion area; Using a variational autoencoder to generate a latent feature distribution, and aligning the latent space with the real data distribution through KL divergence constraint; Jointly optimizing the prior probability and the latent distribution to generate a probability heat map with spatial consistency.
8. The medical image automatic diagnosis method according to claim 5, characterized in that The execution steps of the multi-scale context information enhancement module include: Extracting multi-level regions of interest at the boundary of the abnormal area, and respectively using dilated convolutions with different dilation rates to extract local detail features; Constructing a graph convolutional network to model the global semantic relationship, and performing information transfer between the local detail features and the global semantic nodes; Screening cross-scale features through a feature distillation mechanism, and retaining key features highly relevant to the anatomical structure of the lesion.
9. The medical image automatic diagnosis method according to claim 7, wherein The working process of the knowledge graph-driven diagnostic inference engine includes: Matching the anatomical location and morphological parameters in the lesion segmentation result with the disease-anatomy mapping rules in the medical entity graph; Deriving the potential complication and secondary lesion risks based on the pathological causal relationship chain in the graph; Using a fuzzy logic inference engine to process uncertain diagnosis situations and generating multiple candidate diagnosis schemes including confidence scores.
10. A medical image automatic diagnosis system for implementing the method according to any one of claims 1-9, characterized in that, Including: Data preprocessing module: used to obtain multi-modal medical image data of the target object, including CT images, MRI images and ultrasound images, and perform format standardization processing on the medical image data; Image feature extraction module: constructing a dual-channel convolutional neural network, where the first channel uses a three-dimensional convolutional kernel to extract spatial features, and the second channel uses a temporal convolutional kernel to extract dynamic sequence features, and performing parallel feature extraction on the standardized medical image data; Cross-modal feature fusion layer: performing cross-dimensional splicing on the spatial features and dynamic sequence features extracted through the dual channels to generate a fused feature map; Adaptive attention weight allocation module: dynamically adjusting the feature weights of different regions in the fused feature map through an adaptive attention weight allocation mechanism to highlight the lesion-related features; Temporal dependence modeling module: using a cascaded bidirectional long short-term memory network to perform temporal dependence modeling on the adjusted feature sequence to capture the lesion evolution law; Anomaly detection module: designing an anomaly detection module based on a probabilistic graphical model, generating a potential lesion probability distribution map according to the temporal dependence modeling result, and marking the boundary of the abnormal area; Multi-scale context information enhancement module: integrating a multi-scale context information enhancement module, performing local fine segmentation on the boundary of the abnormal area, and generating a final lesion segmentation mask in combination with global semantic constraints; Diagnostic reasoning engine: Construct a knowledge graph-driven diagnostic reasoning engine, associate and map the lesion segmentation results with a preset medical entity relationship graph, and generate structured diagnostic suggestions.
Citation Information
Cited By
Liver disease image recognition processing method based on multi-modal fusion, medium and equipment
CN120580526A
Liver disease image recognition processing method, medium and device based on multi-modal fusion
CN120580526B
Medical image segmentation method based on adaptive anisotropic convolution
CN120726076A
Chest X-ray image processing method for pneumoconiosis based on AI small shadow detection rate
CN120747041A
Chest x-ray image processing method for ai small shadow detection rate for pneumoconiosis
CN120747041B