Medical image processing method and device, electronic equipment and storage medium
By constructing a four-level progressive intelligent auxiliary diagnostic system, the problems of efficiency, accuracy, and consistency in mediastinal lymph node assessment have been solved, achieving efficient and accurate identification and quantitative analysis of mediastinal lymph nodes, and providing reliable clinical auxiliary diagnostic support.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENYANG NEUSOFT INTELLIGENT MEDICAL TECH RES INST
- Filing Date
- 2025-12-26
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies are inefficient, have limited accuracy, are highly subjective, and have poor repeatability and consistency in mediastinal lymph node assessment. Furthermore, automated methods suffer from scale sensitivity, low contrast, and ambiguous boundaries, making it impossible to simultaneously meet the high standards required in clinical practice.
A four-level progressive intelligent auxiliary diagnostic system is constructed, including guidance, focusing, refinement and evaluation modules. Through deep learning and anatomical rules, the system achieves full-process automation from raw images to structured diagnostic reports, combined with multimodal feature fusion and interpretability analysis.
It enables efficient and accurate identification and quantitative analysis of mediastinal lymph nodes, providing reliable clinical auxiliary diagnostic support and improving the standardization and personalization of diagnosis and treatment.
Smart Images

Figure CN121883392A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical image processing technology, and in particular to a method, apparatus, electronic device and storage medium for processing medical images. Background Technology
[0002] In routine clinical practice, the evaluation of CT images plays a crucial role in the diagnosis and treatment of diseases. Taking the evaluation of mediastinal lymph nodes as an example, this process typically involves two steps: identification and segmentation. In related technologies, these two steps almost entirely rely on manual operation by physicians, or, in the segmentation step, automatic or semi-automatic image segmentation methods are used. Manual operation by physicians suffers from problems such as low efficiency, high resource consumption, strong subjectivity, poor repeatability and consistency, limited accuracy, and susceptibility to systematic biases. Segmentation methods based on traditional image processing have unstable segmentation results or are prone to misjudgment, limiting their practicality. Segmentation methods based on deep learning face challenges such as scale sensitivity, low contrast and blurred boundaries, interference from complex backgrounds, and limited functionality.
[0003] Therefore, neither traditional methods that rely entirely on manual labor nor automated methods can simultaneously meet the high standards of clinical mediastinal lymph node analysis in terms of efficiency, accuracy, and functionality. Summary of the Invention
[0004] This application provides a method, apparatus, electronic device, and storage medium for processing medical images, which simultaneously meet the clinical needs for target tissue analysis in medical images in terms of efficiency, accuracy, and functionality.
[0005] In a first aspect, one embodiment of this application provides a method for processing medical images, comprising: The image to be processed is processed to obtain a first image; wherein the image to be processed is a three-dimensional medical image, and the first image is a discretized three-dimensional image; The first image is divided into multiple image blocks, each image block is processed to obtain the corresponding target image block, and the target image blocks are combined into the second image; The voxel values of some or all voxel points in the second image are updated to obtain the third image, and the candidate tissue partitions included in the third image are filtered to obtain the target tissue partitions; wherein, the tissue indicated by the candidate tissue partition is the target tissue or other tissue, and the tissue indicated by the target tissue partition is the target tissue. Determine the evaluation indicators corresponding to each target tissue partition, and generate descriptive information for the image to be processed based on each evaluation indicator.
[0006] Secondly, one embodiment of this application provides a medical image processing apparatus, comprising: The processing unit is used to: process the image to be processed to obtain a first image; wherein the image to be processed is a three-dimensional medical image, and the first image is a discretized three-dimensional image; The processing unit is further configured to: divide the first image into multiple image blocks, process each image block to obtain a corresponding target image block, and combine the target image blocks into a second image; The processing unit is further configured to: update the voxel values of some or all voxel points in the second image to obtain a third image, and filter each candidate tissue partition included in the third image to obtain each target tissue partition; wherein the tissue indicated by the candidate tissue partition is the target tissue or other tissue, and the tissue indicated by the target tissue partition is the target tissue. The update unit is used to: determine the evaluation index corresponding to each target tissue partition, and generate descriptive information of the image to be processed based on each evaluation index.
[0007] Thirdly, one embodiment of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of any of the above methods.
[0008] Fourthly, one embodiment of this application provides a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implement the steps of any of the above methods.
[0009] Fifthly, one embodiment of this application provides a computer program product, including a computer program that, when executed by a processor, implements the steps of any of the above methods.
[0010] In this embodiment, the three-dimensional medical image is first processed into a discretized three-dimensional image. Then, to achieve more refined segmentation, the first image is divided into multiple image blocks, and each image block is processed to obtain a corresponding target image block. These target image blocks are then combined to form a second image. Next, the voxel values of some or all voxel points in the second image are updated to obtain a third image. The target value tissue partitions from the candidate tissue partitions included in the third image are then filtered out and moved to other tissue partitions. Finally, evaluation indicators corresponding to each target tissue partition are determined, and descriptive information for the image to be processed is generated based on these indicators. Thus, through multi-level intelligent processing of the image to be processed, compared to image processing methods that rely on subjective identification by doctors or are automatic or semi-automatic, the clinical needs for target tissue analysis in medical images are simultaneously met in terms of efficiency, accuracy, and functionality. Attached Figure Description
[0011] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 A framework diagram of an intelligent assisted diagnostic system provided in one embodiment of this application; Figure 2 A schematic flowchart illustrating a medical image processing method provided in one embodiment of this application; Figure 3 A schematic diagram of a three-level progressive guidance and positioning system provided in an embodiment of this application; Figure 4 A schematic diagram of a dual-path fine segmentation provided in an embodiment of this application; Figure 5 A schematic diagram of a three-dimensional graph optimization and morphological constraint system provided in an embodiment of this application; Figure 6 A schematic diagram illustrating a multimodal fusion and intelligent evaluation method according to an embodiment of this application; Figure 7 A schematic diagram of the structure of a medical image processing device provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0013] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.
[0014] For ease of understanding, the terms used in the embodiments of this application are explained below: (1) The core principle of computed tomography (CT) is to use an X-ray beam to perform a tomographic scan around a part of the human body, and then use a computer to process the scan data to generate a tomographic image of that part, which can clearly show the structural details of tissues and organs. This technology is widely used in medical imaging diagnosis. In the embodiments of this application, the medical image can be a CT image.
[0015] (2) Mediastinal lymph nodes (MLN) are a group of lymph nodes distributed in the mediastinum and are an important part of the human thoracic lymphatic system. In the embodiments of this application, the target tissue involved can be a lymph node, such as a mediastinal lymph node.
[0016] (3) A volume pixel, also known as a voxel, is a combination of volume and pixel. A voxel is the basic unit of a three-dimensional digital image. It is the extension of the "pixel" in a two-dimensional image into three-dimensional space. The size of a voxel is determined by the slice thickness, interslice spacing, and planar resolution of the scan. It is a small cubic unit with fixed spatial dimensions (length × width × height). Each voxel corresponds to a specific location in three-dimensional space and carries the tissue feature information of that location (such as the CT value in CT and the signal intensity in magnetic resonance imaging (MRI)).
[0017] Taking medical imaging as an example, the spatial resolution of a 3D image is directly determined by the size of the voxels: the smaller the voxels, the more voxels per unit volume, and the clearer the image details; conversely, the larger the voxels, the more prone the image is to pixelation artifacts. In CT scans, the voxel size is determined by the scanning parameters, and the calculation formula is: voxel size = slice thickness × pixel size (in-plane). In CT images, the voxel value is the CT value, measured in Heinz units (HU), reflecting the X-ray attenuation coefficient of the tissue and distinguishing different tissues such as bone, muscle, fat, and air. In MRI images, the voxel value is the signal intensity value, determined by the tissue's proton density, T1 / T2 relaxation time, etc., and is used to distinguish between normal and diseased tissues.
[0018] (4) Volume data represents the three-dimensional spatial information of the scanned object and can completely restore the three-dimensional morphology of the tissue / object, such as the size, shape, and spatial relationship of the tumor with the surrounding tissue. Its basic unit is voxel. That is, volume data is an ordered collection and functional carrier of voxels, and voxel is the smallest independent unit that constitutes volume data. A single voxel has no practical application value. Only by arranging and integrating a large number of voxels according to their positional relationship in three-dimensional space to form a complete volume dataset can it be used for operations such as three-dimensional reconstruction, quantitative analysis, and virtual cutting.
[0019] The number of any elements in the accompanying drawings is for illustrative purposes only and not as a limitation, and any naming is for distinction only and has no limiting meaning.
[0020] Taking mediastinal lymph nodes as an example, their accurate assessment is particularly important in the diagnosis, clinical staging, treatment strategy formulation, and prognosis of thoracic diseases, especially in the diagnosis and treatment of malignant tumors such as lung cancer, esophageal cancer, and lymphoma. The assessment results directly determine the disease stage, thus influencing the choice of treatment plan, such as whether to use surgery, radiotherapy, chemotherapy, or targeted and immunotherapy, and how to accurately delineate the surgical scope and radiotherapy target area. Furthermore, changes in the size of mediastinal lymph nodes are also a key objective indicator for assessing treatment response during follow-up. Therefore, accurate and efficient identification and quantitative analysis of mediastinal lymph nodes are a core element in improving the standardization and personalization of tumor diagnosis and treatment.
[0021] In routine clinical practice using related technologies, the assessment of mediastinal lymph nodes in CT images relies almost entirely on manual operation by radiologists, oncologists, or radiation oncologists. This process typically involves two steps: identification and segmentation. The specific procedures for these two steps are as follows: During the identification phase, doctors need to browse through a complete chest CT sequence layer by layer to identify shadows with shapes and densities that match the characteristics of lymph nodes from the complex mediastinal anatomy (including major blood vessels, trachea, esophagus, fat, and connective tissue).
[0022] During the segmentation phase, for identified lymph nodes, doctors need to manually mark their boundaries on transverse, coronal, and sagittal images to obtain key quantitative parameters such as their maximum short axis, long axis, cross-sectional area, three-dimensional volume, and average CT value. In radiotherapy planning, this process is fundamental to accurately delineating the clinical target volume (CTV) and the planning target volume (PTV).
[0023] The above manual processing method has the following problems: Problem 1: Low efficiency and high resource consumption. For example, it usually takes 30 minutes to several hours to identify and delineate all mediastinal lymph nodes in a single patient.
[0024] Problem 2: High subjectivity, poor reproducibility and consistency. For example, the segmentation results of the same group of lymph nodes may vary significantly between different doctors, or even by the same doctor at different times, introducing uncertainty into clinical decision-making. Furthermore, in multicenter clinical studies, this variation severely affects the homogeneity and comparability of the data, posing a challenge to the scientific validity of the research findings.
[0025] Problem 3: Limited accuracy, prone to systematic bias. For example, the boundaries of mediastinal lymph nodes are easily blurred on images. When manually delineating, lymph nodes with a diameter <5mm are easily ignored due to partial volume effect, and other adjacent tissues are easily over-segmented.
[0026] Therefore, to overcome the various problems of manual segmentation, a variety of automatic or semi-automatic image segmentation methods have emerged, such as those based on traditional image processing methods and those based on deep learning. However, these technologies still face insurmountable technical bottlenecks when applied to the specific and complex task of mediastinal lymph nodes. Traditional image processing methods are highly unstable in segmenting mediastinal lymph nodes, which have low contrast with surrounding tissues. They are also sensitive to image noise, easily misidentifying structures such as blood vessel sections as lymph nodes, resulting in a high false positive rate and very limited practicality.
[0027] The main problems with deep learning-based methods are as follows: Problem 1: The scale sensitivity challenge. For example, mediastinal lymph nodes vary greatly in size, ranging from tiny lymph nodes a few millimeters to enlarged lymph nodes a few centimeters. Many models struggle to simultaneously capture the overall outline of large targets and the fine structure of tiny targets.
[0028] Problem 2: Low contrast and blurred boundaries. For example, convolutional neural networks have a weak feature response to regions with blurred boundaries during feature extraction, which can easily produce "melting" or "sticking" segmentation results, failing to solve the fundamental difficulties faced in manual delineation.
[0029] Question 3: Complex background interference. For example, the mediastinum has a dense network of blood vessels and overlapping structures, making it easy for the model to misidentify cross-sections of blood vessels with similar density and morphology as lymph nodes. How to enable the model to "learn" to focus on the target rather than the background is a key challenge.
[0030] Question 4: Limitations of Functionality. For example, the vast majority of existing research only focuses on segmentation, making it an "isolated" tool. Its output is merely pixel-level masks, unable to provide direct decision support for clinicians, resulting in an incomplete value chain.
[0031] In summary, neither traditional methods relying entirely on manual labor nor existing automated approaches can simultaneously meet the high standards of clinical mediastinal lymph node analysis in terms of efficiency, accuracy, and functionality. Therefore, there is an urgent need for a medical image processing method that can systematically solve the entire chain of problems from precise localization and pixel-level segmentation to intelligent diagnosis, thereby truly providing clinicians with reliable, efficient, and intelligent auxiliary diagnostic services.
[0032] To this end, this application provides a method and system for mediastinal lymph node segmentation and classification based on multi-level intelligent processing and multi-modal feature fusion. Its core lies in constructing a four-level progressive intelligent auxiliary diagnostic system integrating "guidance-focusing-refinement-evaluation." This system does not rely on any single algorithm model, but rather simulates the diagnostic thinking of a senior radiologist through the organic collaboration and data flow of multiple functional modules, achieving full automation from raw CT image input to the generation of a structured diagnostic report. This application is the first to deeply integrate precise segmentation and benign / malignant risk assessment in a coherent process, and introduces interpretability analysis and feedback learning mechanisms, ultimately realizing an intelligent auxiliary diagnostic system that can not only "see clearly" but also "understand" and "continuously learn."
[0033] Figure 1 This is a framework diagram of an intelligent assisted diagnostic system provided in an embodiment of this application. The overall workflow is a process in which a data flow passes through four core processing levels at once and is uniformly scheduled by a central engine.
[0034] See Figure 1 The overall processing flow is as follows: Step 1: Input the raw CT images into the system.
[0035] Step 2: The data first enters the first-level guidance and coarse localization module. Based on anatomical priors, this module quickly identifies suspicious lymph node candidate areas and marks their anatomical partitions.
[0036] Step 3: The candidate regions are fed into the second-level focusing and fine segmentation module, where each candidate target is precisely segmented at the pixel level through a dual-path deep learning network.
[0037] Step 4: All preliminary segmentation results undergo three-dimensional global optimization, false positive filtering, and result fusion in the third-level refinement and optimization module, and a high-quality three-dimensional lymph node model is output in this module.
[0038] Step 5: The above three-dimensional lymph node model is passed to the fourth-level assessment and decision support module, where benign and malignant risk assessment is performed through multimodal feature extraction, fusion and classification.
[0039] Step 6: Finally, an interpretability analysis can be combined to generate a structured diagnostic report.
[0040] Step 7: The entire process is driven by a central intelligent workflow and feedback learning engine. The doctor's review and correction of the report will form a feedback data stream, which is sent back to the engine for continuous system optimization, forming an intelligent closed loop.
[0041] To further illustrate the technical solutions provided in the embodiments of this application, a detailed description is provided below in conjunction with the accompanying drawings and specific implementation methods. Although the embodiments of this application provide method operation steps as shown in the following embodiments or drawings, the method may include more or fewer operation steps based on conventional or non-inventive methods. In steps where there is no logically necessary causal relationship, the execution order of these steps is not limited to the execution order provided in the embodiments of this application.
[0042] The following is combined with Figure 1 The technical solutions provided in the embodiments of this application will be described.
[0043] refer to Figure 2 This application provides a method for processing medical images, including the following steps: S201: Process the image to be processed to obtain the first image.
[0044] S202: Divide the first image into multiple image blocks, process each image block to obtain the corresponding target image block, and combine the target image blocks into the second image; S203: Update the voxel values of some or all voxel points in the second image to obtain the third image, and filter the candidate tissue partitions included in the third image to obtain the target tissue partitions.
[0045] Among them, the candidate organization partition indicates the target organization or other organizations, and the target organization partition indicates the target organization; S204: Determine the evaluation indicators corresponding to each target tissue partition, and generate descriptive information of the image to be processed based on each evaluation indicator.
[0046] In this embodiment, firstly, the three-dimensional medical image is processed into a discretized three-dimensional image. Secondly, to achieve more refined segmentation, the first image is divided into multiple image blocks, and each image block is processed to obtain a corresponding target image block. These target image blocks are then combined into a second image. Next, the voxel values of some or all voxel points in the second image are updated to obtain a third image. The target value tissue partitions from the candidate tissue partitions included in the third image are then filtered out and moved to other tissue partitions. Finally, evaluation indicators corresponding to each target tissue partition are determined, and descriptive information of the image to be processed is generated based on these evaluation indicators. Thus, through multi-level intelligent processing of the image to be processed, compared to image processing methods that rely on subjective identification by doctors or are automatic or semi-automatic, the clinical needs for target tissue analysis in medical images are simultaneously met in terms of efficiency, accuracy, and functionality.
[0047] In step S201, the image to be processed is a three-dimensional medical image, such as a CT image. In this embodiment, the image to be processed is processed to obtain a first image, which is a discretized three-dimensional image.
[0048] Figure 3 A schematic diagram of a three-level progressive guidance and positioning system provided in an embodiment of this application is shown below. Figure 3 The core idea of this process is to establish a precise spatial coordinate system in the complex mediastinal environment through the collaborative work of deep learning and anatomical rule engine, and to achieve the initial localization of lymph nodes on this basis. Its workflow begins with the deep analysis of the original CT images. Through a specially trained anatomical structure segmentation network, the system can identify key landmark structures in the mediastinum like an experienced radiologist.
[0049] Optionally, this process can be achieved through steps A1-A5: A1: Identify at least one anatomical structure included in the image to be processed.
[0050] The anatomical structures can include the trachea, ascending aorta, descending aorta, superior vena cava, esophagus, etc. In a specific example, this process can be achieved by executing steps A11-A13 through a high-precision anatomical structure segmentation subsystem: A11: Obtain volume data from the image to be processed.
[0051] This involves applying image processing methods from relevant technologies to process the image to be identified, obtaining volumetric data, which can then be used as the basis for subsequent model processing. For example, volumetric data characterizes the three-dimensional spatial information of the scanned object, fully reconstructing the three-dimensional morphology of the tissue / object, such as the size, shape, and spatial relationship of a tumor with surrounding tissues.
[0052] A12: Input the volume data into the first model to determine the anatomical structure to which each voxel in the image to be processed belongs.
[0053] The first model can be a deep fully convolutional network trained on large-scale labeled data (such as nnU-Net variants, which can be embedded with different specialized networks depending on the specific application scenario). This network is optimized for the complex anatomical environment of the mediastinum. The network input is the volumetric data of a complete chest CT scan. It uses an encoder-decoder structure with skip connections and a multi-resolution parallel processing strategy. In the encoding stage, the network uses consecutive convolution and pooling operations to expand the receptive field and capture global contextual information. In the decoding stage, it recovers spatial details through upsampling and feature fusion, and finally outputs the probability of each voxel location belonging to a specific anatomical structure (trachea, ascending aorta, descending aorta, superior vena cava, esophagus, etc.).
[0054] A13: Based on the anatomical structure to which each voxel belongs, determine at least one anatomical structure included in the image to be processed.
[0055] Once the anatomical structure to which each voxel belongs is determined, then all the anatomical structures included in the image to be processed can be determined.
[0056] A2: Based on the pre-defined correspondence between anatomical structures and anatomical regions, perform the operation of dividing at least one anatomical structure into anatomical regions to obtain at least one anatomical region.
[0057] Optionally, this process can be implemented by an intelligent dynamic partitioning mapping engine. The system incorporates a digital mediastinal anatomical atlas conforming to the International Association for the Study of Lung Cancer (IASLC) standards. This atlas defines the three-dimensional spatial relationships between 14 standard lymph node partitions and key anatomical landmarks in a computer-understandable form; that is, the correspondence between anatomical structures and anatomical partitions. Optionally, obtaining at least one anatomical partition can be achieved through the following process: A21: Locating anatomical landmarks.
[0058] Based on the segmentation results obtained in step A1, the characteristic geometric elements of each anatomical structure are automatically identified, such as the tracheal centerline, the highest point of the aortic arch, and the key coordinates of the bronchial carina.
[0059] A22: Analyze spatial relationships.
[0060] Spatial relationships can include pre-defined correspondences between anatomical structures and anatomical regions, or they can be understood as predefined topological rules. The spatial range of each region under each anatomical structure is dynamically calculated. For example, the definition of region 4R depends on the spatial relationship between the right lateral wall of the trachea, the posterior wall of the superior vena cava, and the upper edge of the azygos arch.
[0061] A23: Adjust adaptive boundary.
[0062] In addition, taking into account individual anatomical variations, the system uses fuzzy logic instead of rigid thresholds to handle boundary cases, ensuring the biological rationality of the partitioning results.
[0063] Therefore, the above embodiments achieve intelligent mapping from anatomical structures to functional zones.
[0064] A3: Identify each anatomical region separately and determine the three-dimensional bounding box of the candidate tissue regions included in the anatomical region.
[0065] After obtaining multiple anatomical regions, each region can be identified to determine the candidate tissue regions it includes, and then the 3D bounding box of each candidate tissue region (which could be a lymph node or other tissue) can be determined. Optionally, this process can be implemented using a region-adaptive candidate detection network, as follows: After completing the mediastinal partitioning, that is, after obtaining multiple anatomical partitions, a specially optimized 3D candidate detection network can be run in parallel within each partition. This network adopts a lightweight U-Net architecture and has the following characteristics in the selection of training strategies and loss functions: Feature 1: Region-specific training. For example, training targeted detectors for different regions, taking into full account the typical size and morphological characteristics of lymph nodes in each region.
[0066] Feature 2: High recall optimization objective. For example, using focal loss to balance positive and negative samples ensures high sensitivity for small lymph nodes and marginal cases.
[0067] Feature 3: Multi-task learning framework. For example, it can simultaneously predict the center point heatmap and bounding box parameters of lymph nodes, improving localization accuracy.
[0068] In summary, the 3D detection network outputs the 3D bounding boxes of candidate lymph node partitions within each anatomical region, along with detection confidence scores. In practical applications, the candidate tissue partitions and their respective anatomical region labels can together form the input for the next stage of processing.
[0069] A4: Based on the 3D bounding box of each candidate tissue partition, determine the voxel values of the voxel points included in the candidate tissue partition.
[0070] Optionally, the voxel values of the voxel points included in each candidate tissue partition can be determined based on the coordinates of the three-dimensional bounding box of each candidate tissue partition. For example, the voxel value of the voxel points included in candidate tissue partition 1 is 1, the voxel value of the voxel points included in candidate tissue partition 2 is 2, and the voxel value of the voxel points in the non-candidate tissue partition (background area) is 0.
[0071] A5: Generate the first image based on the voxel values of the voxel points included in each candidate tissue partition.
[0072] Among them, the voxel values of the voxel points included in each candidate tissue partition, as well as the positions of the voxel points, can be used to determine that the first image is a discretized three-dimensional image by referring to the aforementioned calculation process of voxel values.
[0073] Optionally, the discretization operation can be a binarization operation or a gray-level quantization operation. Binarization refers to mapping image pixel values to two discrete values (such as 0 and 255) to distinguish between the target and the background. Gray-level quantization, also known as gray-level layering, compresses an 8-bit grayscale image (256 levels) into 16-level or 8-level grayscale images. Essentially, it divides a continuous grayscale range into N discrete intervals, with each interval's pixel value uniformly mapped to a representative value. In this embodiment, if there is only one candidate tissue partition, the first image is a binarized image; when there are N candidate tissue partitions, the first image includes N+1 voxel values, with each candidate tissue partition corresponding to one voxel value and the background area corresponding to one voxel value.
[0074] In this embodiment of the application, compared with the traditional method of searching for a needle in a haystack in the whole CT image, an intelligent positioning guidance system is established by simulating the anatomical thinking of a senior radiologist. This system transforms abstract anatomical knowledge into calculable and executable spatial rules, laying the foundation for subsequent fine processing.
[0075] In S202, in order to achieve more refined processing, after obtaining the first image, the first image can be divided into multiple image blocks, and then each image block can be processed to obtain the corresponding target image block, and the target image blocks can be combined into the second image.
[0076] Each image block can be 64 64 64 voxels.
[0077] Figure 4 This is a schematic diagram of a dual-path fine segmentation provided in an embodiment of this application, combined with... Figure 4 The process of processing each image patch to obtain the corresponding target image patch can be achieved through steps B1-B3 using a dual-path fine segmentation module that fuses context and details: B1: For each image patch, the image patch is input into the first network layer of the second model to obtain the first feature map (indicating the global features of the image patch), and the image patch is input into the second network layer of the second model to obtain the second feature map (indicating the local features of the image patch).
[0078] The first network layer can be a global context path module. The second model consists of four layers. The first layer uses a large 7×7×7 convolutional kernel to capture a large range of contextual information. Subsequent layers use bottleneck structures of 1×1×1, 3×3×3, and 1×1×1, respectively, to reduce the number of parameters while maintaining feature expressiveness. Instance normalization, rather than batch normalization, is used after each residual block to accommodate contrast differences among different patients. This path introduces a non-local attention module in the high-level feature layer to explicitly model the long-range dependencies between different regions within the lymph node, which is crucial for understanding the overall morphological features of the lymph node.
[0079] The second network layer can be a local detail enhancement path module, which employs a different processing strategy. The input image patch first undergoes contrast-limited adaptive histogram equalization, performed within a local 32×32 window, with the contrast amplification factor limited to within 2.0 to avoid excessive noise enhancement. The enhanced image is then fed into a network based on 3DMobileNetV2, which uses depthwise separable convolutions to significantly reduce the number of parameters while maintaining strong feature representation through linear bottlenecks and inverse residual structures. After the second bottleneck layer of this path, the network introduces a coordinate attention module, which can simultaneously capture channel relationships and positional information, making it particularly effective for features with strong spatial correlations, such as boundaries.
[0080] B2: The first and second feature maps are fused, and the fused feature map is input into the third network layer of the second model to obtain the fused feature map.
[0081] In this model, the third network layer can be a dynamic feature fusion module. This fusion occurs at the decoder entry point, and the process is not a simple concatenation or addition, but rather intelligent fusion achieved through dynamic attention gating. This mechanism first performs global average pooling on the feature maps of both paths to obtain a global descriptor for each feature channel. Then, it learns the non-linear relationship between channels through two fully connected layers to generate channel attention weights. Simultaneously, a 1×1 convolution compresses the dual-path features into a single channel, which is then processed by a sigmoid function to generate a spatial attention map. The two attention mechanisms are combined through an outer product to produce a three-dimensional attention voxel that adaptively calibrates the importance of each spatial location and each feature channel.
[0082] In this way, the second model possesses both macroscopic understanding capabilities and microscopic observation and analysis capabilities.
[0083] B3: Input the fused feature map into the decoder of the second model to obtain the target image patch.
[0084] In practical applications, to improve the learning of features of different lymph nodes, partition prior information can be injected into the second model. This injection occurs before each upsampling layer of the decoder. Specifically, the partition labels are first encoded as 32-dimensional embedding vectors, and then mapped through fully connected layers to the scaling parameter γ and translation parameter β of the instance normalization layer. When the network processes lymph nodes from region 7, the normalization parameters are automatically adjusted to a distribution suitable for the typical lymph node features of that region; while when processing lymph nodes from region 4, the parameters change accordingly. This conditional normalization enables a single network to learn the specific features of lymph nodes from different regions, significantly improving the model's adaptability.
[0085] In this embodiment, the module is responsible for performing sub-millimeter-level fine segmentation on each candidate region obtained in step S201. Addressing the technical challenges of low contrast and blurred boundaries between mediastinal lymph nodes and surrounding tissues, this module innovatively proposes a dual-path collaborative segmentation architecture. Through complementary enhancement of macroscopic context and microscopic details, it achieves accurate boundary delineation in complex backgrounds.
[0086] In S203, after obtaining the second image, the voxel values of some or all voxel points in the second image are updated to obtain the third image.
[0087] Figure 5 This is a schematic diagram of a three-dimensional graph optimization and morphological constraint system provided in an embodiment of this application. Referring to the three-dimensional graph optimization section, the process transforms the segmentation problem into an energy minimization problem and seeks the global optimal solution by constructing a three-dimensional graph model. Specifically, this can be achieved through steps C1-C4: C1: For each voxel in the second image, construct a graph structure centered on the voxel.
[0088] Establish neighborhood connection edges for each voxel point as a node to ensure isotropic spatial relationship modeling.
[0089] C2: Construct the energy function.
[0090] The energy function includes a data term, a smoothing term, and a constraint term. The data term includes the probability that each voxel in the graph structure represents the target tissue (lymph node). For example, if the current voxel is voxel 1, and its neighboring edges include voxels 2-10, then the data term includes the probability that voxels 2-10 indicate a lymph node. The smoothing term is a set model determined based on the voxel distances between the voxel and other voxels in the graph structure (e.g., the distances between voxel 1 and voxels 2-10 respectively). This model (a contrast-sensitive Potts model, which reduces smoothing penalties at clear boundaries to maintain sharp boundaries) is used. The constraint term consists of pre-determined prior constraint values, such as soft constraints based on anatomical knowledge, like a prior reduction in the probability of lymph nodes within a vascular region.
[0091] C3: When the energy function takes a set function value, determine the target voxel value of the voxel point.
[0092] Optionally, a parallelized maximum flow / minimum cut algorithm can be used to ensure that the global optimum or strongly approximate solution, i.e., the target voxel value of the current voxel point, is obtained in a reasonable amount of time. The same method can then be used to obtain the target voxel values of other voxels.
[0093] C4: Generate a third image based on the target voxel values of each voxel point in the second image.
[0094] For example, a third image is generated based on the target voxel values of each voxel point in the second image and the position of each voxel point.
[0095] After obtaining the third image, the candidate tissue partitions included in the third image can be filtered to obtain the target tissue partitions. See [link to relevant documentation]. Figure 5 The morphological constraint part can rely on a morphological manifold constraint filtering system through steps D1-D4. This system is based on the assumption that "normal lymph nodes should conform to a certain statistical regularity in morphology," and constructs lymph node morphological manifolds through unsupervised learning. D1: For each candidate tissue partition included in the third graph, extract the multidimensional morphological features of the candidate tissue partition.
[0096] In this context, the candidate tissue partition indicates the target tissue (lymph node) or other tissues (non-lymph node), while the target tissue partition indicates the target tissue (lymph node).
[0097] Optional multidimensional morphological features include compactness, sphericity, surface area to volume ratio, Fourier descriptor, and moment of inertia ratio.
[0098] D2: Based on the relationship between the pre-defined standard features of the target tissue partition and the extracted multidimensional morphological features of the candidate partition, determine the tissue partition type of the candidate tissue partition.
[0099] Among them, the organization partition type is set organization partition and other organization partitions. That is, the organization partition type of the determined organization partition is set organization partition or other organization partition, such as lymph node partition or other partition.
[0100] Optionally, a variational autoencoder can be used to learn the latent spatial distribution of lymph node morphology on a large amount of labeled data, and the manifold captures the natural range of variation of lymph node morphology that can represent standard features.
[0101] In addition, when comparing the features of standard features and extracted candidate partitions, the reconstruction error and Mahalanobis distance of the candidate partitions in morphological prevalence can be calculated to comprehensively evaluate the normality of their morphology, and those exceeding the threshold can be filtered or marked. Among them, those exceeding the threshold are other tissue partitions.
[0102] D3: Determine the organization partition type as the target organization partition.
[0103] Taking lymph nodes as an example again, the target tissue region is determined to be the lymph node region.
[0104] In practical applications, there is also the problem of adhesion between lymph nodes and blood vessels. This problem can be solved in the following ways: Based on the anatomical structure obtained in S201, the distance between each candidate lymph node and the surrounding blood vessel surface can be calculated. When the distance is less than 2 pixels, the adhesion resolution procedure is initiated. This procedure first determines the adhesion region through morphological operations and topological continuity analysis, and then employs an intelligent separation strategy, using a watershed algorithm based on minimum surface combined with shape priors for intelligent separation. During the separation process, the system ensures that the segmentation results meet topological rationality constraints, such as that a single lymph node should be simply connected and its surface should be a continuous two-dimensional manifold. This process can be achieved through topological relationship analysis and correction components; this is only an example and does not constitute a specific limitation.
[0105] In addition, during the final result integration phase, the system performs three-dimensional connectivity analysis on all lymph nodes that pass the quality check, using a queue-based connected component labeling algorithm to ensure continuity in three-dimensional space. For the same lymph node repeatedly detected in multiple candidate boxes, the system arbitrates based on segmentation confidence and morphological integrity, selects the optimal result, and assigns a globally unique identifier to each finally confirmed lymph node.
[0106] In this embodiment, the preliminary analysis results are refined and verified from a three-dimensional global perspective. By introducing global optimization based on graph theory or graph structure and morphological filtering based on statistical learning, the inherent inconsistencies and false positives of two-dimensional segmentation are systematically solved.
[0107] Regarding S204: After selecting each target tissue partition from each candidate tissue partition, the evaluation index corresponding to each target tissue partition can be determined, and descriptive information of the image to be processed can be generated based on each evaluation index.
[0108] Figure 6 A schematic diagram illustrating multimodal fusion and intelligent evaluation provided in an embodiment of this application is shown below. Figure 6 This process can be achieved through steps E1-E2: E1: For each target organization partition, extract at least one relevant feature of the target organization partition.
[0109] At least one relevant feature may be part or all of morphological and density features, depth semantic features, and clinical context features. Optionally, morphological and density features include short diameter / long diameter, volume / CT value, and impact omics; clinical and context features include anatomical regions and patient age / medical history.
[0110] Specifically, the process of extracting at least one relevant feature of the target tissue is as follows: For morphological and density features, the system not only calculates conventional dimensional parameters but also uses principal component analysis to determine the main orientation of lymph nodes, ensuring that the measurements of the short and long diameters meet clinical standards. A multi-scale strategy is employed for morphological and density feature extraction, calculating the statistical characteristics of CT values at the original resolution, 2mm, and 5mm scales to capture density distribution information at different levels.
[0111] For radiomics features, the extraction strictly follows the Image Biomarker Standardization Initiative. The system first discretizes the CT values using a fixed width of 25 HU, and then calculates the eigenvalues of the gray-level co-occurrence matrix in three-dimensional space, including energy, entropy, contrast, and correlation. To ensure feature reproducibility, all calculations are performed in isotropically resampled space, and mirror-filled boundary voxels are used.
[0112] For deep semantic features, extraction is achieved through a pre-trained specialized network, which is pre-trained on a large lung CT dataset. The fully connected layers are removed at the end, and a 1024-dimensional feature vector is extracted from the penultimate convolutional layer. To capture multi-scale information, the system also extracts features from the network's intermediate layers, which are then concatenated with high-level features after global average pooling to form a rich deep feature representation.
[0113] Clinical and contextual features include anatomical context (zonal coding, distance from important blood vessels, mechanical stress environment of the mediastinal region, etc.), patient-specific factors (age, gender, immune status, history of primary tumor, etc.), and imaging context (special features of other lymph nodes in the same examination, imaging manifestations of the primary lung lesion, etc.).
[0114] E2: Weight at least one relevant feature to obtain a fusion feature, and determine the evaluation index of the target organizational partition based on the fusion feature.
[0115] The weighted processing can be achieved through an attention-weighted feature fusion layer, which learns adaptive weights for different feature modalities and then performs weighted processing on at least one related feature to obtain the fused feature.
[0116] Optionally, the feature fusion process is as follows: Multimodal feature fusion employs a hierarchical attention mechanism. The first level of attention operates within each modality, calculating importance weights for features and suppressing redundant features. The second level of attention operates between modalities, dynamically calculating the relative importance of morphological, depth, and clinical features through a gating mechanism. For example, for smaller lymph nodes, depth semantic features may be assigned higher weights; while for patients with a clear clinical history, the weight of clinical features will be correspondingly increased. The third level is cross-modal attention, modeling the interaction relationships between features from different modalities and capturing complementary information.
[0117] After feature fusion is complete, a gradient boosting tree classifier can be used to identify the fused features and obtain evaluation metrics for the target tissue partitions, such as the probabilities of benign, malignant, and suspected tissue. Additionally, confidence assessment results can be output, for example, based on probability distributions or model variance.
[0118] Optionally, the classifier is implemented based on XGBoost, and its hyperparameters are determined through Bayesian optimization. Five-fold cross-validation is used during training, and early stopping is performed based on validation set performance to prevent overfitting. To provide probability-calibrated prediction results, the system uses Platt scaling on the original XGBoost output for probability calibration, ensuring that the predicted probabilities accurately reflect the likelihood of a sample belonging to a particular class.
[0119] Furthermore, the confidence assessment module integrates multiple sources of uncertainty. For the uncertainty inherent in the model itself, the system employs a deep ensemble approach, training multiple XGBoost models with different initialization parameters, and measuring model uncertainty by the variance of the prediction results. For the inherent uncertainty of the data, the system estimates it by analyzing the density of samples in the feature space; samples located in sparse regions of the feature space are considered to have higher data uncertainty.
[0120] In practical applications, feasibility analysis and report generation can be performed. For example, for morphological and density characteristics, a contribution value (confidence level) of +0.3 indicates a positive effect. The structured analysis report can include two tables: ID / partition / size / volume / probability / confidence level. Visualization can be 2D / 3D views or heatmaps.
[0121] Optionally, the interpretive engine is built on the SHAP framework but optimized for the specific task. For tree models, the system uses the TreeSHAP algorithm to efficiently compute the Shapley value for each feature. To provide more clinically intuitive interpretations, the system groups related features, such as calculating the joint contribution of all texture features as a group. The interpretation results are presented in two forms: for individual predictions, a waterfall plot of feature contributions is provided; for group analysis, a summary plot of feature importance is provided.
[0122] Furthermore, while the report generation system employs a template-based design, the content is entirely data-driven. The system automatically selects the most discriminative features and the most representative image slices for display. For lymph nodes with a high probability of malignancy, the report highlights key factors that drive the malignancy assessment, such as a short diameter exceeding 10mm, irregular borders, and uneven internal density. Simultaneously, the report provides recommendations based on confidence level assessments, explicitly indicating the need for manual review of low-confidence predictions.
[0123] In summary, the embodiments of this application deeply integrate deep learning, image processing, 3D reconstruction and interpretable artificial intelligence technologies. By integrating radiomics, deep learning and clinical information, it realizes a full-process auxiliary diagnostic solution from automatic detection and accurate segmentation to intelligent assessment of benign and malignant diseases, as well as intelligent transformation from image features to clinical insights. It can be widely applied to clinical practice and scientific research in oncology, radiology and radiotherapy.
[0124] Specifically, significant technological advancements have been achieved in the field of mediastinal lymph node analysis by constructing an innovative four-level progressive intelligent diagnostic system. This system establishes a fully automated solution from CT image input to structured report generation, greatly improving clinical diagnostic efficiency. In terms of technical performance, this invention demonstrates superior segmentation accuracy, effectively identifying lymph nodes of different sizes and achieving precise boundary delineation, fully meeting the accuracy requirements of clinical applications. By integrating multimodal features and interpretable artificial intelligence technology, the system exhibits excellent discriminative ability in benign / malignant assessment tasks, while providing transparent decision-making basis, enhancing the credibility of the results. The system also demonstrates strong adaptability and robustness, being compatible with different image acquisition devices and technical parameters, and continuously improving performance through a self-optimization mechanism. Furthermore, the system's integrated 3D visualization, interactive correction, and automated report generation functions bring revolutionary improvements to the clinical diagnostic workflow, providing strong technical support for the diagnosis, treatment planning, and efficacy evaluation of related diseases.
[0125] like Figure 7 As shown, based on the same inventive concept as the above-described medical image processing method, this application also provides a medical image processing apparatus, including: The processing unit 71 is used to: process the image to be processed to obtain a first image; wherein the image to be processed is a three-dimensional medical image, and the first image is a discretized three-dimensional image; The processing unit 71 is further configured to: divide the first image into multiple image blocks, process each image block to obtain a corresponding target image block, and combine the target image blocks into a second image; The processing unit 71 is further configured to: update the voxel values of some or all voxel points in the second image to obtain a third image, and filter each candidate tissue partition included in the third image to obtain each target tissue partition; wherein the tissue indicated by the candidate tissue partition is the target tissue or other tissue, and the tissue indicated by the target tissue partition is the target tissue. The update unit 72 is used to: determine the evaluation index corresponding to each target tissue partition, and generate descriptive information of the image to be processed based on each evaluation index.
[0126] In one optional implementation, the processing unit 71 is specifically used for: Identify at least one anatomical structure included in the image to be processed; Based on the pre-defined correspondence between anatomical structures and anatomical regions, the operation of dividing anatomical regions is performed on at least one anatomical structure to obtain at least one anatomical region. Each anatomical region is identified separately, and the three-dimensional bounding boxes of the candidate tissue regions included in the anatomical region are determined. Based on the 3D bounding box of each candidate tissue partition, determine the voxel values of the voxel points included in the candidate tissue partition; The first image is generated based on the voxel values of the voxel points included in each candidate tissue partition.
[0127] In one optional implementation, the processing unit 71 is specifically used for: The volume data is obtained by identifying the image to be processed; the volume data represents the three-dimensional spatial information of the scanned object indicated by the object to be processed. The volume data is input into the first model to determine the anatomical structure to which each voxel in the image to be processed belongs; Based on the anatomical structure to which each voxel belongs, at least one anatomical structure is determined in the image to be processed.
[0128] In one optional implementation, the processing unit 71 is specifically used for: For each image patch, the image patch is input into the first network layer of the second model to obtain a first feature map, and the image patch is input into the second network layer of the second model to obtain a second feature map; wherein, the first feature map indicates the global features of the image patch, and the second feature map indicates the local features of the image patch; The first and second feature maps are fused, and the fused feature map is input into the third network layer of the second model to obtain the fused feature map. The fused feature map is input into the decoder of the second model to obtain the target image patch.
[0129] In one optional implementation, the processing unit 71 is specifically used for: For each voxel in the second image, construct a graph structure centered on the voxel. Construct an energy function; wherein the energy function includes a data term, a smoothing term, and a constraint term. The data term includes the probability that each voxel in the graph structure is the target organization. The smoothing term is a set model determined based on the voxel distance between the voxel and other voxels in the graph structure. The constraint term is a pre-determined prior constraint value. When the energy function takes a set function value, determine the target voxel value of the voxel point; A third image is generated based on the target voxel values of each voxel point in the second image.
[0130] In one optional implementation, the processing unit 71 is specifically used for: For each candidate tissue partition included in the third image, extract the multidimensional morphological features of the candidate tissue partition; Based on the relationship between the pre-defined standard features of the target organizational partition and the extracted multidimensional morphological features of the candidate partition, the organizational partition type of the candidate organizational partition is determined; where the organizational partition type is the set organizational partition and other organizational partitions. The organization partition type is set to the target organization partition.
[0131] In one alternative implementation, the updating unit 72 is specifically used for: For each target organization partition, extract at least one relevant feature of the target organization partition; At least one relevant feature is weighted to obtain a fused feature, and an evaluation index for the target tissue partition is determined based on the fused feature; wherein the evaluation index is used to indicate the characteristics of the target tissue partition.
[0132] The medical image processing apparatus proposed in this application embodiment adopts the same inventive concept as the above-described medical image processing method and can achieve the same beneficial effects, which will not be repeated here.
[0133] Based on the same inventive concept as the above-described medical image processing method, this application also provides an electronic device, which may specifically be a desktop computer, portable computer, smartphone, tablet computer, personal digital assistant (PDA), server, etc. Figure 8 As shown, the electronic device may include a processor 801 and a memory 802.
[0134] The processor 801 can be a general-purpose processor, such as a central processing unit (CPU), digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.
[0135] Memory 802, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memory, magnetic disk, optical disk, etc. Memory is any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. In the embodiments of this application, memory 802 may also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.
[0136] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned computer storage medium can be any available medium or data storage device that a computer can access, including but not limited to: mobile storage devices, random access memory (RAM), magnetic storage (e.g., floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical storage (e.g., CDs, DVDs, BDs, HVDs, etc.), and semiconductor storage (e.g., ROMs, EPROMs, EEPROMs, non-volatile memory (NAND flash), solid-state drives (SSDs)) and other media capable of storing program code.
[0137] Alternatively, if the integrated units described above in this application are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods of the various embodiments of this application. The aforementioned storage medium includes: mobile storage devices, random access memory (RAM), magnetic memory (e.g., floppy disks, hard disks, magnetic tapes, magneto-optical disks (MO), etc.), optical memory (e.g., CDs, DVDs, BDs, HVDs, etc.), and semiconductor memory (e.g., ROMs, EPROMs, EEPROMs, non-volatile memory (NAND flash), solid-state drives (SSDs), etc.) and other media capable of storing program code.
[0138] Based on the same inventive concept, this application also provides a computer program product, which includes computer program code. When the computer program code is run on a computer, it causes the computer to execute any of the medical image processing methods described above. Since the principle by which the above-described computer program product solves the problem is similar to that of the medical image processing method, the implementation of the above-described computer program product can be referred to the implementation of the method, and repeated details will not be repeated.
[0139] The above embodiments are only used to provide a detailed description of the technical solutions of this application. However, the description of the above embodiments is only for the purpose of helping to understand the methods of the embodiments of this application and should not be construed as a limitation on the embodiments of this application. Any changes or substitutions that can be easily conceived by those skilled in the art should be covered within the protection scope of the embodiments of this application.
Claims
1. A method of processing medical images, characterized in that, include: The image to be processed is processed to obtain a first image; wherein the image to be processed is a three-dimensional medical image, and the first image is a discretized three-dimensional image; The first image is divided into multiple image blocks, each image block is processed to obtain a corresponding target image block, and the target image blocks are combined into a second image; The voxel values of some or all voxel points in the second image are updated to obtain a third image, and the candidate tissue partitions included in the third image are filtered to obtain each target tissue partition; wherein, the tissue indicated by the candidate tissue partition is the target tissue or other tissue, and the tissue indicated by the target tissue partition is the target tissue; The evaluation indicators corresponding to each target tissue partition are determined, and the descriptive information of the image to be processed is generated based on each of the evaluation indicators.
2. The method of claim 1, wherein, The image to be processed is processed to obtain a first image, including: Identify at least one anatomical structure included in the image to be processed; Based on the pre-defined correspondence between anatomical structures and anatomical regions, the operation of dividing the at least one anatomical structure into anatomical regions is performed to obtain at least one anatomical region. Each anatomical region is identified separately, and the three-dimensional bounding box of the candidate tissue regions included in the anatomical region is determined. Based on the three-dimensional bounding box of each candidate tissue partition, the voxel values of the voxel points included in the candidate tissue partition are determined; The first image is generated based on the voxel values of the voxel points included in each candidate tissue partition.
3. The method of claim 2, wherein, Determining that the image to be processed includes at least one anatomical structure includes: The volume data is obtained by identifying the image to be processed; wherein the volume data represents the three-dimensional spatial information of the scanned object indicated by the object to be processed; The volume data is input into the first model to determine the anatomical structure to which each voxel in the image to be processed belongs; Based on the anatomical structure to which each voxel belongs, at least one anatomical structure is determined in the image to be processed.
4. The method of claim 1, wherein, The process of processing each image block to obtain the corresponding target image block includes: For each image patch, the image patch is input into the first network layer of the second model to obtain a first feature map, and the image patch is input into the second network layer of the second model to obtain a second feature map; wherein, the first feature map indicates the global features of the image patch, and the second feature map indicates the local features of the image patch; The first feature map and the second feature map are fused, and the fused feature map is input into the third network layer of the second model to obtain the fused feature map; The fused feature map is input into the decoder of the second model to obtain the target image patch.
5. The method of claim 1, wherein, The step of updating the voxel values of some or all voxel points in the second image to obtain the third image includes: For each voxel in the second image, a graph structure is constructed with the voxel as the center; Construct an energy function; wherein the energy function includes a data term, a smoothing term, and a constraint term, the data term including the probability that each voxel point included in the graph structure is the target organization, the smoothing term being a set model determined based on the voxel distance between the voxel point and other voxel points included in the graph structure, and the constraint term being a predetermined prior constraint value; When the energy function takes a set function value, the target voxel value of the voxel point is determined; The third image is generated based on the target voxel values of each voxel point in the second image.
6. The method of claim 1, wherein, The step of filtering the candidate tissue partitions included in the third image to obtain the target tissue partitions includes: For each candidate tissue partition included in the third graph, extract the multidimensional morphological features of the candidate tissue partition; Based on the relationship between the pre-defined standard features of the target organizational partition and the extracted multidimensional morphological features of the candidate partition, the organizational partition type of the candidate organizational partition is determined; wherein, the organizational partition type is a set organizational partition and other organizational partitions. The organization partition type is determined to be the target organization partition.
7. The method of claim 1, wherein, The determination of the evaluation indicators corresponding to each target organizational partition includes: For each target organization partition, extract at least one relevant feature of the target organization partition; The at least one relevant feature is weighted to obtain a fused feature, and an evaluation index for the target tissue partition is determined based on the fused feature; wherein the evaluation index is used to indicate the characteristics of the target tissue partition.
8. A medical image processing apparatus, characterized by comprising: include: The processing unit is configured to: process the image to be processed to obtain a first image; wherein the image to be processed is a three-dimensional medical image, and the first image is a discretized three-dimensional image; The processing unit is further configured to: divide the first image into multiple image blocks, process each image block to obtain a corresponding target image block, and combine the target image blocks into a second image; The processing unit is further configured to: update the voxel values of some or all voxel points in the second image to obtain a third image, and filter each candidate tissue partition included in the third image to obtain each target tissue partition; wherein the tissue indicated by the candidate tissue partition is the target tissue or other tissue, and the tissue indicated by the target tissue partition is the target tissue. The update unit is used to: determine the evaluation index corresponding to each target tissue partition, and generate descriptive information of the image to be processed based on each of the evaluation indices.
9. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When executed by a processor, the computer program instructions implement the steps of the method according to any one of claims 1 to 7.