Automatic image segmentation and identification system for radiology department
By using multi-scale transformation and adaptive filtering technology, CNN and attention mechanism, recurrent neural network and graph neural network in the medical imaging diagnosis system, the problems of image noise, poor equipment compatibility and insufficient model interpretation are solved, and high-precision imaging diagnosis and personalized treatment plans are achieved.
Patent Information
- Application Number
- CN202510159518.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing medical imaging diagnosis systems have problems such as image noise, poor equipment compatibility, and insufficient model interpretation, resulting in insufficient diagnostic accuracy and reliability, making it difficult to meet personalized medical needs.
The radiological imaging automatic segmentation and recognition system is adopted, and image preprocessing is combined with multi-scale transformation and adaptive filtering technology, and feature extraction and fusion is integrated with CNN and attention mechanism. Recurrent neural networks and graph neural networks are used for disease diagnosis and reasoning, and personalized treatment plans are generated through reinforcement learning algorithms.
It significantly improves the quality of image data and diagnostic accuracy, improves the pertinence and effectiveness of personalized treatment plans, and enhances the system's interpretability and doctor's trust.
Smart Images

Figure CN119991690A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent medical technology, and in particular to an automatic segmentation and recognition system for radiology images. Background Art
[0002] In the current medical field, with the rapid development of science and technology, medical imaging technology has become a key means of disease diagnosis and treatment monitoring. Traditional imaging diagnosis mainly relies on doctors to observe X-rays, CT, MRI and other images with their naked eyes, which is highly subjective. In addition, when faced with massive imaging data and complex diseases, doctors have limited energy and experience, which can easily lead to misdiagnosis and missed diagnosis.
[0003] On the one hand, the raw data collected by imaging equipment is often interfered by many factors, such as the electronic noise of the equipment itself and the physiological movements of the patient during the examination (breathing, heartbeat, etc.). These noises and artifacts will blur the image details and reduce the image quality, which brings great challenges to accurate diagnosis. Moreover, imaging equipment of different brands and models has different imaging parameters and data formats, which makes it difficult to integrate and share imaging data, further hindering the realization of accurate diagnosis.
[0004] On the other hand, in the process of disease diagnosis and reasoning, it is difficult to fully and quickly associate the complex and ever-changing symptoms and imaging features simply by relying on the medical knowledge in the doctor's memory. Medical knowledge is growing explosively, and traditional diagnostic methods cannot efficiently utilize this massive amount of knowledge. The lack of systematic knowledge guidance has greatly reduced the accuracy and reliability of the diagnostic results.
[0005] Furthermore, as the demand for personalized medicine becomes increasingly prominent, the impact of individual differences in patients (age, gender, medical history, genes, etc.) on treatment plans becomes increasingly critical. However, in the past, the medical system was unable to fully consider these factors and customize the best treatment strategy for each patient. Often, only relatively broad conventional treatment methods could be adopted, and the treatment effect was unsatisfactory.
[0006] At the same time, the uneven distribution of medical resources is also a major problem. Large medical institutions have accumulated a wealth of case data, but the data circulation is limited; grassroots medical institutions have difficulty improving their diagnosis level due to the small number of cases and insufficient experience. Moreover, the poor compatibility of imaging equipment and computing platforms in different medical institutions has hindered the widespread promotion and coordinated development of technology.
[0007] At the same time, the currently disclosed patent "CN 116052847 B" is a chest X-ray multi-abnormality recognition system, device and method based on deep learning, which solves some key problems in traditional chest X-ray diagnosis, but still has some limitations.
[0008] First, in terms of the comprehensiveness and accuracy of abnormal sign recognition, although the comparative patent can identify common abnormalities such as atelectasis and calcification, the imaging feature learning of some rare diseases and difficult diseases is still not in-depth enough. For example, in the recognition of subtle signs of early lung tumors, due to the relatively small number of samples, its model is difficult to accurately capture those extremely obscure feature changes, resulting in the risk of missed diagnosis of such diseases, which cannot meet the strict clinical requirements for high-precision diagnosis.
[0009] Secondly, from the perspective of the complexity of clinical application scenarios, different hospitals use different models and parameters of chest X-ray imaging equipment, and compared with patents, they lack cross-device compatibility. The training model often works well for images collected by specific devices. Once encountering chest X-rays output by old equipment or new high-end equipment, the differences in image brightness, contrast, resolution, etc. will cause the recognition accuracy to drop significantly, hindering the implementation of the technology in a wider range of medical scenarios.
[0010] Furthermore, regarding the interpretability of the model, the patent focuses on the construction and recognition function of the model, but does not fully consider the trust building needs of doctors in actual use. The deep learning model is like a "black box", and it is difficult for doctors to understand why the model makes a certain diagnosis. Especially when facing complex cases and doubtful diagnostic results, it is impossible to obtain effective explanations from the model level, which is not conducive to improving the scientificity and reliability of clinical decision-making.
[0011] Therefore, there is an urgent need for an innovative intelligent medical image analysis system to solve the above problems. Summary of the invention
[0012] In order to solve the problems of the prior art, the present invention provides a radiology image automatic segmentation and recognition system.
[0013] In order to solve the above technical problems, the present invention is implemented by the following technical solutions: a radiology image automatic segmentation and recognition system, comprising:
[0014] The image preprocessing module combines multi-scale transformation and adaptive filtering technology to perform noise reduction, enhancement and normalization on various medical images, including X-rays, CT and MRI, and build a pixel-level image quality improvement model. This module uses wavelet transform for multi-scale decomposition, and its continuous wavelet transform formula is:
[0015]
[0016] In image processing, f(t) corresponds to a certain dimensional signal of the image. By adjusting the scale parameter a and the translation parameter b, the wavelet function ψ(t) is used to remove noise and extract features. At the same time, during adaptive filtering, the filtering parameters are automatically adjusted according to the local characteristics of the image, such as calculating the adaptive filter kernel size based on the local variance, to achieve accurate noise suppression and contrast enhancement.
[0017] The feature extraction and fusion unit integrates the convolutional neural network (CNN) and the attention mechanism to extract multi-scale and multi-modal key features from the preprocessed images and generate comprehensive feature representations through feature fusion strategies; in CNN, the convolution operation follows the formula:
[0018]
[0019] Where x is the input feature map, w is the convolution kernel weight, b is the bias, y is the output feature map, l represents the number of layers, and the feature pyramid is constructed by extracting local and global features in parallel through convolution kernels of different sizes. The attention mechanism uses the softmax function to assign weights to different features, focusing on the image areas that are closely related to disease diagnosis, and improving the effectiveness of feature fusion.
[0020] The disease diagnosis reasoning engine uses a recurrent neural network (RNN) combined with a medical knowledge base to infer the type, severity and development trend of the disease based on image features; for dynamic images (such as dynamic CT and MRI sequences), a long short-term memory network is used.
[0021] (LSTM), whose forget gate f t =σ(W f .[h t-1 ,x t ]+b f ), input gate i t =σ(W i .[h t-1 ,x t ]+b i ), output gate o t =σ(W o .[h t-1 ,x t ]+b o ), candidate memory cells Memory Unit and the hidden state h t =o t ⊙tanh(C t )(The meaning of each parameter: σ is the sigmoid function, which controls the degree of information passing; W is the weight matrix, b is the bias term, h t-1 is the hidden state at the previous moment, x tThe current moment input is used to effectively process time series information through these gating units). At the same time, the medical knowledge base is constructed in the form of a knowledge graph and integrated with the image features. The graph neural network is used to realize the diagnostic reasoning guided by knowledge. The graph convolution operation of the graph neural network follows the formula:
[0022]
[0023] (where H l is the feature matrix of the lth layer, is the adjacency matrix with self-connection added, Its degree matrix, W l is the weight matrix, σ is the activation function, which realizes the fusion propagation of knowledge graph and image features);
[0024] The personalized treatment plan recommendation module uses a reinforcement learning algorithm to generate the most optimized treatment plan recommendation based on the diagnosis results and individual patient information, including but not limited to age, gender, and medical history. The treatment effect evaluation index is used as the reward function R(s,a) (s is the state, a is the action, i.e., the treatment plan). Through the interaction between the agent and the environment (medical decision-making scenario), according to the policy gradient algorithm▽ θ J(θ)=Ε[▽ θ logπ θ (a|s)R(s,a)](where π θ (a{s) strategy function, θ is the strategy parameter, and the treatment plan is optimized by iteratively optimizing the strategy parameter θ) iteratively optimizing the treatment plan.
[0025] In a specific implementation of the first aspect, the image preprocessing module includes:
[0026] The multi-scale transformation component uses wavelet transform or Laplace pyramid algorithm to decompose images at different scales, remove noise while retaining key edge information; for example, in wavelet transform, in addition to the above continuous wavelet transform formula, similar calculations will be extended to two dimensions when processing two-dimensional images to better capture image features;
[0027] The adaptive filter submodule automatically adjusts the filter parameters according to the local features of the image, such as using a smaller filter kernel in texture-rich areas and a larger filter kernel in smooth areas, to achieve accurate noise suppression and contrast enhancement. Specifically, the regional texture features can be judged based on statistics such as the local grayscale standard deviation, and the filter kernel size can be adjusted. The calculation formula can be expressed as k = f (σ local )(where k is the filter kernel size, σ local is the local grayscale standard deviation, f is the adjustment function, and the specific form is determined based on experience or experiments).
[0028] In a specific implementation of the first aspect, the feature extraction and fusion unit implements:
[0029] Multi-scale feature extraction mechanism, by designing convolution kernels of different sizes to extract local and global features of images in parallel in CNN and construct a feature pyramid;
[0030] Attention-guided feature fusion uses the attention mechanism to assign weights to different features, focusing on the image areas that are closely related to disease diagnosis, and improving the effectiveness of feature fusion. Based on the channel attention mechanism, the channel descriptors are first obtained by global average pooling of each channel of the feature map, and then the channel weights are calculated through a multi-layer perceptron (MLP). The linear transformation in the MLP can be expressed as y = W x +b (where x is the input vector, W is the weight matrix, b is the bias, and y is the output vector), and finally the weight is multiplied by the original feature map to achieve feature re-weighting.
[0031] In a specific embodiment of the first aspect, the disease diagnosis reasoning engine comprises:
[0032] Time series feature encoder, for dynamic images, including but not limited to dynamic CT and MRI sequences, uses LSTM or GRU units to encode the image's changing features over time;
[0033] The knowledge graph fusion unit constructs the medical knowledge base in the form of a knowledge graph, integrates it with image features, and uses graph neural networks to achieve knowledge-guided diagnostic reasoning. In the process of knowledge graph construction, the relationship between entities such as diseases, symptoms, and test results can be represented by the relationship matrix R. If there is a certain relationship between entities e1 and e2, then R e1,e2 =1, otherwise R e1,e2 =0, combined with the above graph convolution formula of the graph neural network, the deep fusion of knowledge graph and image features is realized.
[0034] In a specific embodiment of the first aspect, the personalized treatment plan recommendation module includes:
[0035] The patient feature modeling component quantifies and encodes the individual information of the patient to construct a patient feature vector. For categorical information such as age and gender, one-hot encoding is used. If gender is divided into male and female, male is encoded as [1,0] and female is encoded as [0,1]. For text information such as medical history, the text can be converted into a vector representation through a word vector model such as Word2Vec or a pre-trained BERT model, and then concatenated or fused to obtain a patient feature vector.
[0036] The reinforcement learning decision model uses the treatment effect evaluation index as the reward function, and iteratively optimizes the treatment plan through the interaction between the agent and the environment (medical decision scenario). In actual training, the experience replay mechanism can be used to store the agent's experience (s, a, r, s′) (state, action, reward, next state) in the replay buffer, and randomly sample for training to improve training efficiency. The size of the replay buffer and the sampling strategy can be adjusted according to system performance and training requirements.
[0037] In a specific implementation of the first aspect, it also includes:
[0038] The virtual case generation unit uses the generative adversarial network (GAN) technology to generate a variety of virtual cases based on real case data for model training and verification, ensuring that the generated virtual cases are similar to real cases in terms of imaging features and disease type distribution. By adjusting the noise input of GAN or introducing conditional variables, virtual cases covering different conditions and different patient characteristics are generated; based on the adversarial training objective function
[0039]
[0040] (Unconditional case), which is extended to medical imaging
[0041]
[0042] (where G is the generator, \-D) is the discriminator, and y is a specific case condition, such as disease type, imaging modality, etc. The generator attempts to generate realistic images to deceive the discriminator, and the discriminator learns to distinguish between real and generated images, and improves the generation ability through adversarial training). At the same time, during the generation process, based on anatomical prior knowledge, it ensures that the generated virtual cases conform to the physiological structure of the human body in terms of organ morphology and position.
[0043] The model interpretability analysis module reveals the basis and process of model decision-making through feature visualization and sensitivity analysis methods, thereby improving the credibility of the system. In terms of feature visualization, the gradient ascent method can be used to calculate the gradient of the input image with respect to the model output, and then adjust the image pixel value according to the gradient size to highlight the areas that have a greater impact on the model decision. The calculation process involves the gradient calculation formula in the back-propagation algorithm. By visualizing these areas, doctors can understand the focus of the model.
[0044] In a specific implementation of the first aspect, the system integrates:
[0045] Self-supervised pre-training strategy, which uses massive unlabeled image data for pre-training, learns the common features of images, and accelerates the convergence speed of subsequent supervised learning;
[0046] The continuous learning framework uses elastic weight integration (EWC) or gradient-based sample selection methods to achieve rapid model updates when new diseases and new imaging technologies emerge.
[0047] In a specific implementation of the first aspect, the system implements:
[0048] The distributed collaborative training architecture uses federated learning technology to achieve joint model training of multi-center data while protecting the data privacy of medical institutions. It uses homomorphic encryption technology to protect data privacy, based on the additive homomorphic formula of the Paillier encryption algorithm:
[0049] (Enc(m_1)\cdotEnc(m_2)=Enc(m_1+m_2) is the encryption function, (m_1 is the plaintext data. Through this encryption method, the data is calculated in the ciphertext state to protect privacy). At the same time, combined with the model aggregation algorithm, such as the FedAvg algorithm, the model parameters trained by each center are weighted averaged to obtain the global model.
[0050] The cross-platform adaptation interface enables the system to run stably on medical imaging devices and computing platforms of different brands and models through the software abstraction layer and the hardware adaptation layer. In the software abstraction layer, unified interface specifications can be defined, such as the image data reading interface and the model calling interface, to ensure compatibility on different devices and platforms; in the hardware adaptation layer, the calculation process of the model is optimized according to the characteristics of different hardware architectures (such as CPU, GPU, TPU, etc.), such as reasonably allocating GPU memory and optimizing convolution calculation when GPU acceleration is used, so as to improve the system operation efficiency.
[0051] In a second aspect, a method for automatic segmentation and recognition of radiology images comprises the following steps:
[0052] S1: Collect the multimodal medical imaging data of the patient and perform standardization processing to unify the image resolution and grayscale range, using the linear transformation formula (y=ax+b is the original image pixel value, (y, (b$ are coefficients determined according to standardization requirements);
[0053] S2: Input into the system to obtain disease diagnosis results and personalized treatment plan recommendations;
[0054] S3: Display diagnostic reports, treatment plans and related explanatory information through a visual interface to assist doctors in making clinical decisions. Use charts and text to display information in multiple forms, use bar charts to display the possibility of different diseases, and use text to describe the steps and precautions of the treatment plan in detail to improve doctors' acceptance of information.
[0055] In a third aspect, a device for constructing an automatic segmentation and recognition system for radiology images includes: a processor and a memory;
[0056] The memory is used to store one or more program instructions;
[0057] The processor is used to run one or more program instructions to execute the steps of a method for automatic segmentation and recognition of radiology images.
[0058] The beneficial effects of the present invention are:
[0059] 1. Improved diagnostic accuracy: The image preprocessing module combines multi-scale transformation (such as wavelet transform or Laplace pyramid algorithm) with adaptive filtering technology to effectively remove noise from various medical images (CT, MRI, etc.), while accurately retaining key information such as the edges of tiny lesions in the lungs and subtle tissue structures in the brain. This greatly improves the quality of image data input into subsequent analysis processes, laying a solid foundation for accurate diagnosis and reducing misdiagnosis and missed diagnosis caused by image noise or blur;
[0060] The feature extraction and fusion unit integrates CNN and attention mechanism. CNN uses convolution kernels of different scales to extract local and global features of images in parallel, builds a feature pyramid, and fully captures the characteristics of lesions. The attention mechanism further focuses on the image areas that are closely related to disease diagnosis and assigns high weights to key features. The two work together to ensure that the system does not miss subtle but important signs of symptoms, significantly improve the accuracy and comprehensiveness of feature extraction, and thus improve the accuracy of disease diagnosis.
[0061] The disease diagnosis inference engine uses a recurrent neural network (RNN, such as LSTM or GRU) combined with a knowledge graph and graph neural network constructed from a medical knowledge base to fully explore the time evolution characteristics in dynamic image sequences and the logical associations guided by knowledge. For example, in the diagnosis of lung diseases, the type, severity and development trend of the disease can be more accurately inferred based on the changes in the reinforcement pattern of dynamic CT images over time, combined with information about symptoms, disease causality, etc. in the medical knowledge graph, providing a reliable basis for clinical treatment.
[0062] 2. Customization of personalized medical plans: The personalized treatment plan recommendation module uses a reinforcement learning algorithm to generate the optimal treatment plan based on the diagnosis results and individual patient information (age, gender, medical history, etc.). By quantitatively encoding patient characteristics, such as hierarchical classification encoding of medical history, using algorithms such as proximal policy optimization (PPO) or deep Q network (DQN), and using treatment effect evaluation indicators as reward functions, the effects of different treatment strategies in virtual medical decision-making scenarios are simulated. This makes the treatment plan tailored for each patient more in line with their specific conditions, improves the targetedness and effectiveness of treatment, and is expected to improve treatment prognosis;
[0063] The system continuously learns patient treatment feedback information. When new treatment effect data flows in, the self-supervised pre-training framework and continuous learning mechanism (such as the elastic weight integration EWC algorithm) can quickly update the model knowledge and further optimize the treatment plan recommendations for subsequent patients, forming a virtuous circle and constantly adapting to the changing needs in clinical practice.
[0064] 3. Optimization of medical resource utilization: The data transmission module uses high-speed fiber optic channels or 5G wireless transmission technology, combined with a reliable verification mechanism (such as CRC or hash function verification) and compliance with the DICOM standard to achieve real-time, lossless transmission of image data, greatly shortening the time interval between image acquisition and diagnostic analysis. On the one hand, it reduces the time patients spend waiting for diagnostic results and improves their medical experience; on the other hand, it accelerates the turnover of hospital diagnosis and treatment processes and improves the utilization efficiency of medical resources (such as imaging equipment, medical staff time, etc.);
[0065] The virtual case generation unit uses the generative adversarial network (GAN) technology to generate diverse virtual cases. These virtual cases are not only used for model training and verification, expanding sample diversity and making up for the lack of actual case data, but also play a role in doctor training, preoperative simulation and other scenarios. For example, before a neurosurgery operation, by generating virtual MRI images of specific brain diseases, doctors can plan the operation plan in advance, reduce surgical risks, reduce unnecessary exploratory operations, and save medical costs.
[0066] 4. Enhanced system versatility and adaptability: The multi-center federated learning architecture uses homomorphic encryption technology to protect the data privacy of each medical institution and realize the multi-center data joint training model. This enables the system to gather image data knowledge collected by different regions and different devices, break through the data limitations of a single institution, improve the generalization ability of the model, and adapt to a wider range of patient groups and disease types. Whether in large comprehensive hospitals or primary medical institutions, it can play a stable and efficient diagnostic role;
[0067] The cross-platform adaptation interface ensures that the system is compatible with medical imaging equipment and computing platforms of different brands and models through the software abstraction layer and the hardware adaptation layer. Whether it is an old X-ray machine or the latest 3.0T MRI equipment, as well as servers with different architectures (CPU, GPU, TPU, etc.), the system can be seamlessly connected and run normally, reducing the hardware update cost of medical institutions and improving the feasibility of technology promotion.
[0068] 5. Improved transparency of medical decision-making: The model interpretability analysis module presents the basis of model decision-making to doctors in the form of intuitive heat maps and charts through feature visualization (such as Grad-CAM or Layer-CAM methods) and sensitivity analysis. Doctors can clearly see the image areas and features that the model focuses on during the diagnosis process and understand why the model makes a specific diagnostic conclusion. This not only enhances doctors' trust in the system's diagnostic results, but also makes it easier for doctors to make secondary judgments based on their own experience in complex cases, thereby improving the scientificity and transparency of medical decision-making. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 It is a schematic diagram of the internal system of the present invention.
[0070] Figure 2 It is a schematic diagram of the external data connection of the system of the present invention.
[0071] Figure 3 It is a schematic diagram of the disease diagnosis reasoning engine architecture of the present invention. DETAILED DESCRIPTION
[0072] The following will be combined with the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0073] like Figures 1 to 3 An automatic segmentation and recognition system for radiological images is shown.
[0074] Embodiment 1: Hardware support of lung disease diagnosis and treatment auxiliary system based on chest CT images:
[0075] Image acquisition unit:
[0076] The detector based on cadmium telluride (CdTe) technology is selected, and its pixel size reaches 0.1 mm × 0.1 mm, which can clearly capture tiny lung lesions. The detector has automatic exposure control technology, which monitors the changes in chest thickness and tissue density in real time through the built-in density sensor, and automatically adjusts the X-ray tube current and voltage to ensure that the grayscale range of each CT image collected is stable in an appropriate range, such as the grayscale value range is controlled within 0-4095, to ensure the consistency of image quality.
[0077] Equipped with an adjustable scanning bed, its movement accuracy can reach 0.01 mm, which facilitates the precise positioning of the patient's chest, ensures full coverage of the lung area, and reduces scanning blind spots.
[0078] Data transmission module:
[0079] It uses high-speed fiber optic channels to transmit image data, with a transmission rate of up to 10Gbps, and can losslessly transmit a complete chest CT scan data (about 500MB) to the back-end processing system within a few seconds.
[0080] In compliance with the DICOM3.0 standard, each data block transmitted is verified using the CRC32 cyclic redundancy check algorithm to ensure data integrity. Once a verification error is found, the retransmission mechanism is automatically triggered to ensure reliable data transmission.
[0081] Segmentation and recognition core processor:
[0082] It is equipped with NVIDIA Tesla V100 GPU, has 32GB of video memory, and uses the TensorFlow deep learning framework to run pre-trained lung image segmentation and recognition models.
[0083] The server host is equipped with an Intel Xeon Gold 6248R processor with a main frequency of 3.0 GHz, 24 cores and 48 threads, providing powerful computing power to ensure that the model runs quickly and meets real-time diagnosis needs.
[0084] Software support:
[0085] Image preprocessing module:
[0086] The multi-scale transformation component uses a two-dimensional discrete wavelet transform and Daubechies4 (db4) wavelet basis function to perform a three-layer decomposition of the chest CT image, removing noise while highlighting key information such as lung edges and textures. In the adaptive filter submodule, the adaptive filter kernel size is calculated based on the grayscale standard deviation in the local 5×5 neighborhood of the image. When the standard deviation is less than 10, a 7×7 filter kernel is used for smoothing; when the standard deviation is greater than 20, it switches to a 3×3 filter kernel to retain more details, effectively suppress noise and enhance contrast.
[0087] Feature extraction and fusion unit:
[0088] In the CNN feature extraction part, three convolution kernels of different sizes (3×3, 5×5, and 7×7) were designed to extract local and global features of lung images in parallel and construct a feature pyramid. For the fusion of feature maps of different scales, the concatenate operation was used to splice the feature maps of each scale in the channel dimension and input them into the subsequent network layer. The attention mechanism is based on channel attention. The channel descriptor of 1×1×C (C is the number of channels) is obtained through global average pooling. The channel weight is calculated through two layers of MLP (the number of neurons in the first layer is C / 2, and the number of neurons in the second layer is C). The MLP weight matrix is initialized to Xavier uniform distribution, and the bias is initialized to 0. Finally, the weight is multiplied by the original feature map to achieve feature reweighting and focus on the suspicious lesion area of the lung.
[0089] Disease Diagnosis Reasoning Engine:
[0090] For the dynamic scanning sequence of chest CT, LSTM units are used to encode the characteristics of image changes over time. The input dimension of LSTM units is 128, the hidden state dimension is 256, the weight matrix of the forget gate, input gate, and output gate is initialized to Glorot normal distribution, and the bias is initialized to 0. At the same time, the lung disease knowledge base is constructed in the form of a knowledge graph, and the Neo4j graph database is used to store entities such as diseases, symptoms, and test results and their relationships. The diagnostic reasoning guided by knowledge is realized through a custom graph neural network. The number of graph neural network layers is 3, and the number of neurons in each layer is 512. The ReLU activation function is used to combine image features with knowledge graph information to infer the type, severity, and development trend of lung diseases.
[0091] Personalized treatment plan recommendation module:
[0092] The patient feature modeling component quantifies and encodes the patient's age, gender, smoking history, family history and other information. Age is coded directly by numerical value, gender is coded by one-hot (male [1,0], female [0,1]), smoking history is divided into three levels according to the number of years of smoking (0-5 years [1,0,0], 6-15 years [0,1,0], more than 15 years [0,0,1]), and family history is coded as 1 if there is lung disease, otherwise it is 0. These codes are spliced into an 8-dimensional patient feature vector. The reinforcement learning decision model uses the symptom relief rate within 30 days as the treatment effect evaluation indicator (reward function), and uses the proximal policy optimization (PPO) algorithm to iteratively optimize the treatment plan. The policy network is a two-layer fully connected neural network with 64 neurons in each layer and a learning rate of 0.001. Through the interaction between the agent and the simulated medical decision environment, personalized treatment recommendations are generated for patients, such as medication regimens and review cycles.
[0093] Virtual case generation unit:
[0094] An improved model based on WassersteinGAN (WGAN) was used to generate virtual chest CT cases. Both the generator and the discriminator used a 5-layer convolutional neural network structure. The generator input a 100-dimensional random noise vector and combined it with specific lung disease conditions (such as early lung cancer, pneumonia, etc.) to generate realistic images through adversarial training. In the generation process, based on the prior knowledge of lung anatomy, a constraint formula based on shape context was used to ensure that the lung morphology and bronchial branching structure of the generated virtual cases were consistent with the physiological structure of the human body. The Euclidean distance between the generated cases and the real cases in the feature space was calculated (the 2048-dimensional feature vector was extracted using the pre-trained InceptionV3 model). If the distance was greater than the threshold of 0.5, the generator parameters were adjusted to make the generated cases closer to the distribution of real cases for model training and verification, and to expand the diversity of case samples.
[0095] Model interpretability analysis module:
[0096] The Grad-CAM (gradient weighted class activation mapping) method is used for feature visualization. The gradient of the model output is calculated for the input chest CT image, and the gradient is back-propagated to the convolutional layer. A weight is assigned to each feature map channel according to the gradient size. The feature maps are then weighted and summed to obtain a heat map. This highlights the lung areas that have a greater impact on the model's decision, helps doctors understand the model's focus, and improves the system's credibility.
[0097] Implementation process steps:
[0098] The patient lies on an adjustable scanning bed, and the image acquisition unit automatically adjusts the scanning parameters, performs a CT scan on the chest, and obtains the original image data.
[0099] The data is transmitted to the back-end system in real time via high-speed fiber channels, and CRC verification is performed during the transmission process to ensure that the data is complete and correct.
[0100] The image preprocessing module performs wavelet transform multi-scale decomposition and adaptive filtering on the received CT images in turn to improve the image quality.
[0101] The preprocessed images enter the feature extraction and fusion unit, where CNN and the attention mechanism work together to extract and fuse key features.
[0102] For dynamic scanning sequences, the LSTM unit in the disease diagnosis reasoning engine encodes temporal features and combines the knowledge graph for diagnostic reasoning to obtain the diagnosis results of lung diseases.
[0103] The personalized treatment plan recommendation module generates personalized treatment plans based on the diagnosis results and patient feature vectors through a reinforcement learning algorithm.
[0104] The virtual case generation unit generates virtual chest CT cases on demand for model optimization training; the model interpretability analysis module generates feature visualization heat maps to assist doctors in interpreting diagnostic results.
[0105] Ultimately, the diagnostic report (including disease type, severity, development trend, etc.), treatment plan (detailed suggestions on medication, follow-up, etc.) and related explanatory information (feature visualization diagrams, case comparisons, etc.) are displayed to doctors through a visual interface to assist clinical decision-making.
[0106] Example 2: Neurological disease diagnosis auxiliary system based on brain MRI images
[0107] Hardware support:
[0108] Image acquisition unit:
[0109] Amorphous silicon (a-Si) detectors with a pixel size of 0.08 mm × 0.08 mm are used in conjunction with a 3.0T superconducting magnetic resonance imaging system to obtain high-resolution brain MRI images, clearly presenting subtle brain structures such as the hippocampus and basal ganglia. Automatic exposure control technology dynamically adjusts radio frequency pulse sequence parameters based on the proton density, T1 and T2 relaxation time characteristics of different brain regions to ensure good image contrast under different scanning sequences (T1WI, T2WI, FLAIR, etc.), for example, the contrast between gray matter and white matter under the T1WI sequence reaches more than 1.5.
[0110] Equipped with a high-precision head fixation device with a positioning accuracy of up to 0.005 mm, it reduces the impact of patient head movement on image quality and ensures the repeatability of each scan.
[0111] Data transmission module:
[0112] Utilizing 5G wireless transmission technology, real-time transmission of brain MRI image data is achieved, with a transmission bandwidth of up to 1Gbps, meeting the needs of high-definition image transmission. A hash function-based verification mechanism is used to perform integrity verification on the transmitted data. Once the verification fails, the data is immediately retransmitted from the source to ensure data accuracy. The transmission protocol complies with the DICOM3.0 standard and seamlessly connects to the hospital information system.
[0113] Segmentation and recognition core processor:
[0114] Equipped with Google CloudTPUv3, it provides up to 100TFLOPS of computing power and runs the brain image segmentation and recognition model pre-trained under the PyTorch deep learning framework.
[0115] The supporting server uses AMD EPYC 7742 processor, with a main frequency of 2.25 GHz, 64 cores and 128 threads, and large-capacity memory (256 GB), ensuring that the system runs efficiently and stably when processing large-scale brain MRI data.
[0116] Software support:
[0117] Image preprocessing module:
[0118] The multi-scale transformation component uses the Laplace pyramid algorithm to perform a 4-layer decomposition of brain MRI images to remove noise and retain edge information. The adaptive filter submodule calculates the filter kernel size based on the grayscale variance in the local 7×7 neighborhood of the image. When the variance is less than 8, a 9×9 filter kernel is used to enhance the smoothing effect; when the variance is greater than 15, it switches to a 5×5 filter kernel to protect details, effectively improving image quality.
[0119] Feature extraction and fusion unit:
[0120] In the CNN feature extraction stage, four convolution kernels of different sizes (3×3, 5×3, 7×3, and 9×3) were designed to extract brain image features in different directions and construct feature pyramids. Feature fusion uses a combination of point-by-point addition and concatenation. First, some highly correlated feature maps are added point by point, and then the results are concatenated with other feature maps in the channel dimension and input into the subsequent network layer. The attention mechanism is based on spatial attention. By performing maximum pooling and average pooling operations on the feature map, two 1×1×C descriptors are obtained. After concatenating them, a convolution layer (the convolution kernel size is 7×7 and the number of output channels is 1) is used to calculate the spatial attention map, which is multiplied with the original feature map to focus on the brain lesion area.
[0121] Disease Diagnosis Reasoning Engine:
[0122] For the brain MRI dynamic enhancement sequence, GRU units are used to encode the characteristics of image changes over time. The input dimension of the GRU unit is 100, the hidden state dimension is 200, the update gate, reset gate, and candidate hidden state weight matrix are initialized to He normal distribution, and the bias is initialized to 0. At the same time, the knowledge base of neurological diseases is constructed in the form of a knowledge graph, and the GraphDB graph database is used to store entity relationships such as diseases, symptoms, and image features. The diagnostic reasoning is combined with a custom graph neural network. The graph neural network has 4 layers and 256 neurons in each layer. The LeakyReLU activation function is used to infer the type, severity, and development trend of neurological diseases based on image features and knowledge graphs.
[0123] Personalized treatment plan recommendation module:
[0124] The patient feature modeling component quantifies the patient's age, gender, history of hypertension, history of diabetes, previous stroke, and other information. The age and gender are coded in the same way as in the previous example. If the patient has a history of hypertension or diabetes, it is coded as 1, otherwise it is 0. A history of major diseases such as previous stroke is coded 0-1, and these codes are spliced into a 6-dimensional patient feature vector. The reinforcement learning decision model uses the neurological function recovery score within 90 days as the reward function and uses the deep Q network (DQN) algorithm to iteratively optimize the treatment plan. The Q network is a three-layer fully connected neural network with 32 neurons in each layer and a learning rate of 0.0005. Through the interaction between the intelligent agent and the simulated medical decision-making scenario, personalized treatment recommendations are generated for patients, such as rehabilitation training plans and drug treatment plans.
[0125] Virtual case generation unit:
[0126] The improved model based on CycleGAN is used to generate virtual brain MRI cases. Both the generator and the discriminator are 6-layer convolutional neural network architectures. The generator inputs a 120-dimensional random noise vector and generates realistic images through adversarial training in combination with specific neurological disease conditions (such as brain tumors, multiple sclerosis, etc.). During the generation process, based on the prior knowledge of brain anatomy, a constraint formula based on topology preservation is used to ensure that the brain structure of the generated virtual cases is reasonable. The cosine similarity between the generated cases and the real cases in the feature space is calculated (using the pre-trained ResNet50 model to extract the 1000-dimensional feature vector). If the similarity is lower than the threshold of 0.6, the generator parameters are adjusted to make the generated cases closer to the real case distribution for model training and verification, enriching the diversity of case samples.
[0127] Model interpretability analysis module:
[0128] The Layer-CAM (layer class activation mapping) method is used for feature visualization. The gradient of the model output is calculated after each convolutional layer of the model, and weights are assigned to the feature map of that layer according to the gradient size. Then, the weighted feature maps of each layer are summed to obtain a heat map, highlighting the brain areas that have a greater impact on the model's decision-making, helping doctors understand the basis for model judgment and enhancing the credibility of the system.
[0129] Implementation process steps:
[0130] The patient wears a head fixation device and enters the 3.0T magnetic resonance imaging system. The image acquisition unit automatically adjusts the scanning parameters according to the characteristics of the brain, performs an MRI scan, and obtains the original image data.
[0131] 5G wireless transmission technology is used to transmit image data to the back end in real time, and hash function verification is performed during the transmission process to ensure data integrity.
[0132] The image preprocessing module performs Laplace pyramid multi-scale decomposition and adaptive filtering on the received brain MRI images to improve the image quality.
[0133] The preprocessed image enters the feature extraction and fusion unit, and CNN and attention mechanism jointly extract and fuse key features.
[0134] For dynamic enhancement sequences, the GRU unit in the disease diagnosis reasoning engine encodes temporal features and combines the knowledge graph to complete diagnostic reasoning and obtain the diagnosis results of neurological diseases.
[0135] The personalized treatment plan recommendation module combines the diagnosis results with the patient's feature vector and generates a personalized treatment plan with the help of a reinforcement learning algorithm.
[0136] The virtual case generation unit generates virtual brain MRI cases on demand for model optimization training; the model interpretability analysis module generates feature visualization heat maps to assist doctors in interpreting diagnostic results.
[0137] Finally, the diagnostic report (covering disease type, severity, development trend, etc.), treatment plan (detailed suggestions on rehabilitation training and medication, etc.) and related explanatory information (feature visualization graphs, case comparisons, etc.) are displayed to doctors through a visual interface to assist in clinical decision-making.
[0138] These two examples elaborate on the hardware selection, software algorithm configuration and specific implementation process steps, and build a system with practical significance according to the claims, which can provide accurate imaging diagnosis and treatment assistance functions for lung diseases and nervous system diseases respectively. You can further adjust the parameters, algorithms or processes according to actual needs to make them more suitable for application scenarios.
[0139] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A radiology image automatic segmentation and recognition system, characterized in that: include: The image preprocessing module combines multi-scale transformation and adaptive filtering technology to perform noise reduction, enhancement and normalization on various medical images, including X-rays, CT and MRI, and build a pixel-level image quality improvement model; The feature extraction and fusion unit integrates the convolutional neural network (CNN) and the attention mechanism to extract multi-scale and multi-modal key features from the pre-processed images and generate a comprehensive feature representation through feature fusion strategy; Disease diagnosis inference engine, which uses recurrent neural network (RNN) combined with medical knowledge base to infer disease type, severity and development trend based on image features; The personalized treatment plan recommendation module uses a reinforcement learning algorithm to generate the most optimized treatment plan recommendations based on the diagnosis results and individual patient information, including but not limited to age, gender, and medical history.
2. The automatic segmentation and recognition system for radiology images according to claim 1, characterized in that: The image preprocessing module comprises: Multi-scale transformation component, using wavelet transform or Laplace pyramid algorithm to decompose images at different scales, removing noise while retaining key edge information; The adaptive filtering submodule automatically adjusts the filtering parameters according to the local characteristics of the image, such as using a smaller filter kernel in texture-rich areas and a larger filter kernel in smooth areas, to achieve accurate noise suppression and contrast enhancement.
3. The automatic segmentation and recognition system for radiology images according to claim 1, characterized in that: The feature extraction and fusion unit realizes: Multi-scale feature extraction mechanism, by designing convolution kernels of different sizes to extract local and global features of images in parallel in CNN and construct a feature pyramid; Attention-guided feature fusion uses the attention mechanism to assign weights to different features, focusing on image areas that are closely related to disease diagnosis, and improving the effectiveness of feature fusion.
4. The automatic segmentation and recognition system for radiology images according to claim 1, characterized in that: The disease diagnosis reasoning engine comprises: Time series feature encoder, for dynamic images, including but not limited to dynamic CT and MRI sequences, uses LSTM or GRU units to encode the image's changing features over time; The knowledge graph fusion unit constructs the medical knowledge base in the form of a knowledge graph, integrates it with image features, and uses graph neural networks to achieve knowledge-guided diagnostic reasoning.
5. The automatic segmentation and recognition system for radiology images according to claim 1, characterized in that: The personalized treatment plan recommendation module includes: Patient feature modeling component, which quantifies and encodes the individual information of patients and constructs patient feature vectors; The reinforcement learning decision-making model uses the treatment effect evaluation index as the reward function and iteratively optimizes the treatment plan through the interaction between the intelligent agent and the environment-medical decision-making scenario.
6. The automatic segmentation and recognition system for radiology images according to claim 1, characterized in that: Also includes: The virtual case generation unit uses the generative adversarial network (GAN) technology to generate a variety of virtual cases based on real case data for model training and verification, ensuring that the generated virtual cases are similar to real cases in terms of imaging features and disease type distribution. By adjusting the noise input of GAN or introducing conditional variables, virtual cases covering different conditions and different patient characteristics are generated; The model interpretability analysis module reveals the basis and process of model decision-making through feature visualization and sensitivity analysis methods, thereby improving the credibility of the system.
7. The automatic segmentation and recognition system for radiology images according to claim 1, characterized in that: The system integrates: Self-supervised pre-training strategy, which uses massive unlabeled image data for pre-training, learns the common features of images, and accelerates the convergence speed of subsequent supervised learning; The continuous learning framework uses elastic weight integration (EWC) or gradient-based sample selection methods to achieve rapid model updates when new diseases and new imaging technologies emerge.
8. The automatic segmentation and recognition system for radiology images according to claim 1, characterized in that: The system implements: Distributed collaborative training architecture, using federated learning technology, to achieve joint model training of multi-center data while protecting the privacy of medical institution data; The cross-platform adaptation interface enables the system to run stably on medical imaging devices and computing platforms of different brands and models through the software abstraction layer and hardware adaptation layer.
9. A method for automatic segmentation and recognition of radiological images, characterized in that: The following steps are involved: S1: Collect multimodal medical imaging data of patients and perform standardized processing; S2: Input into the system to obtain disease diagnosis results and personalized treatment plan recommendations; S3: Display diagnostic reports, treatment plans and related explanatory information through a visual interface to assist doctors in making clinical decisions.
10. A device for constructing an automatic segmentation and recognition system for radiology images, characterized in that: include: Processor and memory; The memory is used to store one or more program instructions; The processor is used to run one or more program instructions to execute the steps of the method for automatic segmentation and recognition of radiological images as claimed in claim 9.
Citation Information
Cited By
Prediction system for chronic obstructive pulmonary disease and storage medium
CN120299689A
Medical training evaluation method and system based on intelligent simulated patient
CN120656734A
Medical image diagnosis auxiliary system based on deep learning
CN121601220A