Image Processing Method, Device, Equipment and Storage Medium for Medical Images
By training the image processing model, differential training of encoding network, decoding network and generating network is solved, the problem of neural network models ignoring weak expression characteristics in medical image segmentation is achieved, and the acquisition and segmentation effect of multimodal features is improved.
Patent Information
- Application Number
- CN202110938701.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-16
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2041-08-16
AI Technical Summary
In the prior art, neural network models focus on strong expression characteristics when segmenting medical images and ignore weak expression characteristics, resulting in incomplete segmentation results.
By calling the encoding network, decoding network and generating network in the image processing model, the model is trained based on the difference between the predicted segmented image and the label image and the predicted generated image and the second sample image to obtain multimodal medical image features.
It improves the comprehensiveness of medical image segmentation results, solves the problem of image missing in single-modal image analysis, and enhances the segmentation effect.
Smart Images

Figure CN114283151B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of medical technologies, and particularly to an image processing method, apparatus, device, and storage medium for medical images. Background Art
[0002] In the medical field, medical image segmentation by medical imaging technology has become a common technique for assisting doctors in case judgment.
[0003] In related technologies, a medical image is usually input into a neural network model, and medical image segmentation is performed based on the medical image features extracted by the neural network, so as to obtain a medical image segmentation result.
[0004] However, in the neural network model in the above technology, the medical image features it focuses on are often the strong expressive features of the input medical image, and less attention is paid to the weak expressive features in the medical image, resulting in incomplete information in the obtained medical image segmentation result and poor medical image segmentation effect. Summary of the Invention
[0005] Embodiments of this application provide an image processing method, apparatus, device, and storage medium for medical images, which can improve the obtained image processing effect. The technical solution is as follows:
[0006] On the one hand, an image processing method for medical images is provided, and the method includes:
[0007] Call the first encoding network in the image processing model to encode the first sample image to obtain a first feature map corresponding to the first sample image; the first sample image is a sample medical image of the first modality of the target medical object;
[0008] Call the decoding network in the image processing model to decode based on the first feature map to obtain a predicted segmentation image of the first sample image; the predicted segmentation image is used to indicate at least one specified type of region;
[0009] Call the generation network in the image processing model to generate a predicted generation image based on the first feature map; the predicted generation image is a predicted image of the second modality corresponding to the first sample image;
[0010] Train the image processing model based on the difference between the predicted segmentation image and the label image, and the difference between the predicted generation image and the second sample image; the second sample image is a sample medical image of the second modality of the target medical object; the label image is an image corresponding to the target medical object and used to indicate at least one specified type of region.
[0011] On the other hand, there is provided an image processing apparatus for medical images, the apparatus comprising:
[0012] A first encoding module, configured to call a first encoding network in an image processing model to encode a first sample image and obtain a first feature map corresponding to the first sample image; the first sample image is a sample medical image of a first modality of a target medical object;
[0013] A decoding module, configured to call a decoding network in the image processing model to decode based on the first feature map and obtain a predicted segmentation image of the first sample image; the predicted segmentation image is used to indicate at least one specified type of region;
[0014] A generation module, configured to call a generation network in the image processing model to generate a predicted generated image based on the first feature map; the predicted generated image is a predicted image of a second modality corresponding to the first sample image;
[0015] A model training module, configured to train the image processing model based on the difference between the predicted segmentation image and the label image, and the difference between the predicted generated image and a second sample image; the second sample image is a sample medical image of a second modality of the target medical object; the label image is an image corresponding to the target medical object and used to indicate at least one specified type of region.
[0016] In a possible implementation manner, the model training module includes:
[0017] A first determination sub-module, configured to determine a function value of a first loss function based on the difference between the predicted segmentation image and the label image;
[0018] A second determination sub-module, configured to determine a function value of a second loss function based on the difference between the predicted generated image and the second sample image;
[0019] A model training sub-module, configured to train the image processing model based on the function value of the first loss function and the function value of the second loss function.
[0020] In a possible implementation manner, the model training sub-module is configured to update parameters of the first encoding network and parameters of the decoding network based on the function value of the first loss function;
[0021] Update parameters of the first encoding network and parameters of the generation network based on the function value of the second loss function.
[0022] In a possible implementation manner, the first determination sub-module includes:
[0023] A first determination unit, configured to determine a function value of a first branch function of the first loss function based on a similarity between the predicted segmentation image and the label image;
[0024] A second determination unit, configured to determine a function value of a second branch function of the first loss function based on positions of at least one specified type region predicted in the predicted segmentation image and positions of at least one specified type region in the label image;
[0025] A third determination unit, configured to determine a function value of the first loss function based on the function value of the first branch function and the function value of the second branch function.
[0026] In a possible implementation manner, the first determination unit is configured to obtain weight values respectively corresponding to respective divided regions in the predicted segmentation image; the respective divided regions in the predicted segmentation image include the at least one specified type region;
[0027] Based on the weight values respectively corresponding to the respective divided regions in the predicted segmentation image and the similarity between the respective divided regions in the predicted segmentation image and the respective divided regions in the label image, determine the function value of the first branch function of the first loss function.
[0028] In a possible implementation manner, the apparatus further includes:
[0029] A discrimination module, configured to call a discriminator to discriminate the predicted generated image and obtain a discrimination result of the predicted generated image;
[0030] A third determination module, configured to determine a function value of a third loss function based on the discrimination result; the discrimination result is used to indicate whether the predicted generated image is a real image;
[0031] The model training module is configured to train the image processing model based on the function value of the first loss function, the function value of the second loss function, and the function value of the third loss function.
[0032] In a possible implementation manner, the first encoding network includes N encoding layers, and the N encoding layers are connected to each other in pairs, N≥2 and is a positive integer;
[0033] The first encoding module includes:
[0034] A set acquisition sub-module, configured to acquire a first image pyramid corresponding to the first sample image, where the first image pyramid is a set of images obtained by downsampling the first sample image according to a specified gradient, and the first image pyramid contains N first images to be processed;
[0035] An encoding sub-module, configured to respectively input the N first images to be processed into corresponding encoding layers, encode the N first images to be processed, and obtain N first feature maps corresponding to the first sample image;
[0036] Wherein, in response to the target encoding layer being a non-first encoding layer among the N encoding layers, the input of the target encoding layer further includes the first feature map output by the previous encoding layer.
[0037] In a possible implementation, the decoding network in the image processing model includes N decoding layers, and the N decoding layers are connected pairwise, and the N decoding layers correspond to the N encoding layers one by one;
[0038] The decoding module includes:
[0039] A decoding sub-module, configured to respectively input the N first feature maps into the corresponding decoding layers of the decoding network, decode the N first feature maps, and obtain N decoding results; the N decoding results have the same resolution;
[0040] A merging sub-module, configured to merge the N decoding results to obtain a predicted segmentation image of the first sample image;
[0041] Wherein, in response to the target decoding layer being a non-first decoding layer among the N decoding layers, the input of the target decoding layer further includes the decoding result output by the previous decoding layer.
[0042] In a possible implementation, the apparatus further includes:
[0043] An image acquisition module, configured to acquire a prior constraint image of the image processing model based on a third sample image; the third sample image is a sample medical image of a third modality of the target medical object; the prior constraint image is used to indicate the position of the target medical object in the third sample image;
[0044] A second encoding module, configured to call a second encoding network in the image processing model, and encode based on the prior constraint image to obtain a second feature map corresponding to the third sample image;
[0045] A merging module, configured to merge the first feature map and the second feature map to obtain a comprehensive feature map;
[0046] The decoding module is configured to call the decoding network in the image processing module to perform decoding based on the comprehensive feature map to obtain the predicted segmentation image of the first sample image;
[0047] The generation module is configured to call the generation network in the image processing model to generate the predicted generated image based on the comprehensive feature map.
[0048] In a possible implementation manner, the apparatus further includes:
[0049] The cropping module is configured to crop the prior constraint image based on the position of the target medical object;
[0050] The second encoding module is configured to call the second encoding network in the image processing model to encode the cropped prior constraint image to obtain the second feature map corresponding to the third sample image.
[0051] In a possible implementation manner, the image acquisition module is configured to call a semantic segmentation network to process the third sample image to obtain the prior constraint image of the image processing model.
[0052] In a possible implementation manner, the parameters in the second encoding network share the weight with the parameters in the first encoding network.
[0053] On the other hand, a computer device is provided, which includes a processor and a memory. The memory stores at least one instruction, at least one program, a code set or an instruction set. The at least one instruction, the at least one program, the code set or the instruction set is loaded and executed by the processor to implement the above-mentioned image processing method for medical images.
[0054] On the other hand, a computer-readable storage medium is provided, in which at least one computer program is stored. The computer program is loaded and executed by a processor to implement the above-mentioned image processing method for medical images.
[0055] On the other hand, a computer program product or a computer program is provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned image processing method for medical images provided in various optional implementation manners.
[0056] The technical solution provided by this application may include the following beneficial effects:
[0057] The image processing method for medical images provided by the embodiments of the present application obtains multi-modal sample medical images corresponding to a target medical object and a label image corresponding to the target medical image and containing region labels of a specified type. A predicted segmentation image and a predicted generated image are generated based on a first sample image in the multi-modal sample medical images. Based on the difference between the predicted segmentation image and the label image and the difference between the predicted generated image and a second sample image corresponding to the target medical object, an image processing model including a first encoding network, a decoding network, and a generation network is trained, so that the trained image processing model can obtain the features of multi-modal medical images based on a single-modal medical image, making the information contained in the obtained medical image segmentation result more comprehensive and improving the segmentation effect of medical images;
[0058] Further, based on the trained image processing model, other modal medical images can be generated based on a single-modal medical image, thereby solving the problem of missing images in the process of medical image analysis. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present application and, together with the specification, are used to explain the principles of the present application.
[0060] Figure 1 It shows a schematic diagram of the system architecture of the image processing method for medical images provided by an exemplary embodiment of the present application;
[0061] Figure 2 It shows a flowchart of the image processing method for medical images provided by an exemplary embodiment of the present application;
[0062] Figure 3 It is a framework diagram of the generation of an image processing model and image processing shown according to an exemplary embodiment;
[0063] Figure 4 It shows a flowchart of the image processing method for medical images provided by an exemplary embodiment of the present application;
[0064] Figure 5 It shows a flowchart of the image processing method for medical images shown by an exemplary embodiment of the present application;
[0065] Figure 6 It shows a schematic diagram of the synthesis of approximate markings shown by an exemplary embodiment of the present application;
[0066] Figure 7 It shows a schematic diagram of the structure of the image processing model shown by an exemplary embodiment of the present application;
[0067] Figure 8The figure shows a schematic structural diagram of an encoding layer shown in an exemplary embodiment of the present application;
[0068] Figure 9 The figure shows a schematic structural diagram of a decoding layer shown in an exemplary embodiment of the present application;
[0069] Figure 10 The figure shows a schematic diagram of an application process of an image processing model shown in an exemplary embodiment of the present application;
[0070] Figure 11 The figure shows a block diagram of an image processing device for medical images shown in an exemplary embodiment of the present application;
[0071] Figure 12 The figure shows a block diagram of a computer device shown in an exemplary embodiment of the present application;
[0072] Figure 13 The figure shows a block diagram of a computer device shown in an exemplary embodiment of the present application. Detailed implementation manners
[0073] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0074] It should be understood that "a plurality" mentioned herein refers to two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0075] The embodiments of the present application provide an image processing method for medical images, which can improve the accuracy of image segmentation. The present application relates to artificial intelligence technology and machine learning technology;
[0076] Among them, Artificial Intelligence (AI) uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0077] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields involved, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning. The display device including an image acquisition component shown in this application mainly involves the directions of computer vision technology and machine learning / deep learning among them.
[0078] Machine Learning (ML) is an interdisciplinary subject across multiple fields, involving multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0079] Figure 1 The schematic diagram of the system architecture of the image processing method for medical images provided by an exemplary embodiment of this application is shown, as Figure 1 shown, the system includes: a computer device 110 and a medical image acquisition device 120.
[0080] Among them, the above computer device 110 can be implemented as a terminal or a server. When the computer device 110 is implemented as a server, the computer device 110 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. When the computer device 110 is implemented as a terminal, the computer device 110 can be a smart phone, a tablet computer, a laptop computer, a desktop computer, and so on.
[0081] The above medical image acquisition device 120 is a device with a medical image acquisition function. For example, the medical image acquisition device can be a CT (Computed Tomography) detector for medical detection, a nuclear magnetic resonance instrument, a positron emission computed tomography scanner, a cardiac magnetic resonance instrument, and other devices with an image acquisition device. Schematically, taking the cardiac magnetic resonance instrument as an example, Cardiac Magnetic Resonance (CMR) refers to a method of diagnosing heart and large blood vessel diseases using magnetic resonance imaging technology. Magnetic resonance is a non-invasive imaging technology; the CMR images obtained based on cardiac magnetic resonance imaging can provide anatomical and functional information of the heart to assist in the clinical diagnosis and treatment of heart diseases. For example, CMR images can assist in the clinical diagnosis and treatment of myocardial infarction.
[0082] Cardiac magnetic resonance imaging is a multi-modal imaging method. Different CMR imaging sequences correspond to different imaging focuses, which are used to provide different cardiac feature information. Schematically, the imaging sequences of CMR can include: balanced-Steady State Free Precession (bSSFP) sequence. This sequence can capture cardiac motion so that the corresponding bSSFP images can present a complete and clear myocardial boundary; T2-weighted imaging. Its corresponding T2-weighted images can clearly show myocardial edema or myocardial ischemic injury. For example, the T2-weighted images show the myocardial edema site or myocardial ischemic injury site in a highlighted form; Late Gadolinium Enhancement (LGE) technique. Its corresponding LGE images can prominently show myocardial scars or myocardial infarction regions. By combining multiple image sequences, rich and reliable information about myocardial pathology and morphology can be obtained to assist in clinical diagnosis and the setting of treatment plans. It should be noted that the above description of the CMR imaging sequences is only schematic. Relevant personnel can set different imaging sequences according to actual needs to obtain different CMR images, and this application does not limit this. Further, the multi-modal medical images shown in this application can be medical images corresponding to the same medical object obtained based on different medical image acquisition devices. For example, the multi-modal medical images can include medical images such as T1-weighted images, T2-weighted images, and CT images.
[0083] Optionally, one or more computer devices 110 and one or more medical image acquisition devices 120 are included in the above system. The number of computer devices 110 and medical image acquisition devices 120 is not limited in the embodiments of this application.
[0084] The medical image acquisition device 120 and the computer device 110 are connected through a communication network. Optionally, the communication network is a wired network or a wireless network.
[0085] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is usually the Internet, but can also be any network, including but not limited to any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or a virtual private network. In some embodiments, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent the data exchanged through the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can be used to encrypt all or some of the links. In other embodiments, customized and / or dedicated data communication technologies can be used to replace or supplement the above data communication technologies. This application does not make any restrictions here.
[0086] Figure 2 The flowchart of the image processing method for medical images provided by an exemplary embodiment of the present application is shown. This method is executed by a computing device, and the computer device can implement Figure 1 the server shown, such as Figure 2 shown, and the image processing method for medical images includes the following steps:
[0087] Step 210, call the first encoding network in the image processing model to encode the first sample image, and obtain a first feature map corresponding to the first sample image; the first sample image is a sample medical image of the first modality of the target medical object.
[0088] In the embodiments of the present application, for medical images, the specified type of region label can be used to represent the lesion information in the first sample image, and the lesion information can include information such as the position and shape of the lesion in the first sample image.
[0089] The sample images (including the first sample image and the second sample image) in the embodiments of the present application can be medical images obtained by a medical image acquisition device, such as Figure 1The medical image acquisition device as shown; alternatively, the sample image can be obtained based on the medical image data stored in the database. Schematically, the sample image involved in the embodiments of the present application can be obtained based on the image data in the publicly available dataset MyoPS20, which consists of multi-sequence myocardial cases CMR, including 45 cases of bSSFP images, T2-weighted images, and LGE images, among which 25 cases are labeled. For the original CMR sequences of each patient, the bSSFP images consist of 8-12 slices, with an in-plane resolution of 1.25×1.25 mm and a slice thickness of 8 to 13 mm. The T2-weighted images consist of 3-7 slices, with an in-plane resolution of 1.35×1.35 mm and a slice thickness of 12-20 mm. The LGE images have 10-18 slices, with an in-plane resolution of 0.75×0.75 mm and a slice thickness of 5 mm. Align the above images to a common space and resample them to the same spatial resolution to obtain the sample image in the present application.
[0090] The label image corresponding to the first sample image may include other lesion regions with lower resolution in the first sample image. Schematically, taking the first sample image as the T2-weighted image of the target medical object as an example, the lesion region with higher resolution is the myocardial edema region. If the target medical object corresponding to the T2-weighted image has myocardial scar, then in the label image corresponding to the T2-weighted image, in addition to including the label of the myocardial edema region, it may also include the label of the myocardial scar region. Correspondingly, when the first sample image is the LGE image of the target medical object, the lesion region with higher resolution is the myocardial scar region. In the label image corresponding to the LGE image, in addition to including the label of the myocardial scar region, it may also include the label of the myocardial edema region; that is to say, the label images corresponding to different modalities of medical images of the same target medical object are the same.
[0091] The modality of the sample medical image is used to indicate the acquisition method of the medical image. Schematically, the sample medical image of the first modality can be a T2-weighted image, or the sample medical image of the first modality can also be a LEG image, or the sample medical image of the first modality can also be a medical image acquired by any other medical image acquisition method.
[0092] Step 220, call the decoding network in the image processing model, decode based on the first feature map, and obtain the predicted segmentation image of the first sample image; the predicted segmentation image is used to indicate at least one specified type region predicted.
[0093] Optionally, the number of predicted specified type regions in the predicted segmentation image is equal to the number of specified label regions in the label image. Schematically, the predicted specified type region may be a lesion region in the first sample image predicted after processing by the first encoding network and the decoding network.
[0094] Step 230: Invoke the generation network in the image processing model to generate a predicted generated image based on the first feature map; the predicted generated image is a predicted image of the second modality corresponding to the first sample image.
[0095] Among them, the first modality to which the first sample image belongs is different from the second modality to which the predicted generated image belongs.
[0096] Step 240: Train the image processing model based on the difference between the predicted segmentation image and the label image, and the difference between the predicted generated image and the second sample image; the second sample image is a sample medical image of the second modality of the target medical object; the label image is an image corresponding to the target medical object and used to indicate at least one specified type region.
[0097] The computer device obtains sample medical images of different first modalities as the first sample image, iteratively executes the above steps 210 to 240, and iteratively updates the parameters in the image processing model based on the difference between the predicted segmentation image corresponding to each first sample image and the label image, and the difference between the predicted generated image and the second sample image, until the training completion condition is reached. The training completion condition includes: the image processing model converges, the number of iterations reaches the number threshold, and so on.
[0098] The image processing model after training completion can be used to perform medical image segmentation on the input target medical image of the first modality to obtain the specified type region in the target medical image, and / or generate a medical image of the second modality corresponding to the target medical image.
[0099] In summary, the image processing method for medical images provided in the embodiments of the present application obtains sample medical images of multiple modalities corresponding to a target medical object, and a label image including labels of specified type regions corresponding to the target medical image. Based on the first sample image in the sample medical images of multiple modalities, a predicted segmentation image and a predicted generated image are generated. Based on the difference between the predicted segmentation image and the label image, and the difference between the predicted generated image and the second sample image corresponding to the target medical object, the image processing model including the first encoding network, the decoding network, and the generation network is trained, so that the trained image processing model can obtain the characteristics of medical images of multiple modalities based on a single modality medical image, making the information included in the obtained medical image segmentation result more comprehensive and improving the segmentation effect of medical images;
[0100] Furthermore, the image processing model obtained based on training can generate medical images of other modalities based on medical images of a single modality, thereby solving the problem of missing images in the process of medical image analysis.
[0101] In the solution described in the embodiments of the present application, by using multi-modal medical sample images of the same target medical object and training an image processing model with the label image corresponding to the target medical object, the medical image segmentation effect of the image processing model can be improved, and / or the problem of missing images in the process of medical image analysis can be solved. The application scenarios of the above solution include, but are not limited to, the following scenarios:
[0102] 1) Myocardial infarction diagnosis and treatment scenario:
[0103] The assessment of myocardial viability is crucial for the diagnosis and treatment management of patients with myocardial infarction. In practical applications, cardiac magnetic resonance (CMR) imaging technology can be used to obtain CMR images of the heart corresponding to the imaging sequence to provide anatomical and functional information of the heart. Different imaging sequences can image and provide different characteristic information of the heart, including late gadolinium enhancement (LGE) images showing myocardial infarction regions, T2-weighted images highlighting myocardial edema or myocardial ischemic injury, and balanced steady-state free precession sequence (bSSFP) sequence images that can capture heart movement and present clear boundaries. These multi-sequence CMR images can provide rich and reliable information on myocardial pathology and morphology, helping doctors with diagnosis and treatment planning. However, in a single-modal scenario, the information about the heart obtained based on medical images of a single modality is limited. For example, when only T2-weighted images are available, only relatively clear myocardial edema or myocardial ischemic injury can be obtained based on the T2-weighted images, and it is difficult to obtain information about myocardial infarction (myocardial scar) regions; when only LGE images are available, only relatively clear myocardial infarction regions can be obtained, and it is difficult to obtain information about myocardial edema or myocardial ischemic injury. In this case, the image processing models corresponding to the T2-weighted images and LGE images respectively obtained by the image processing method for medical images provided in the embodiments of the present application can be used, that is, the image processing model obtained by using the sample medical images of the T2-weighted modality as the first sample images for training, and the image processing model obtained by using the sample medical images of the LGE modality as the first sample images for training. The T2-weighted image is input into the image processing model corresponding to the T2-weighted modality to obtain a segmentation image including myocardial scar regions and myocardial edema (or myocardial ischemic injury), and / or the T2-weighted image corresponding to the LGE image; or, the LGE image is input into the image processing model corresponding to the LGE modality to obtain a segmentation image including myocardial scar regions and myocardial edema (or myocardial ischemic injury), and / or the T2-weighted image corresponding to the LGE image.
[0104] 2) Medical image lesion judgment scenario:
[0105] In the medical field, medical staff often judge the lesion area of an organ through medical images obtained by medical image acquisition devices. For example, for lesion examination of the stomach, gastric ulcer is confirmed; lung tumor is confirmed; and brain tumor is confirmed, etc. In the above scenarios, the image processing method for medical images provided by this application can be used to obtain an image processing model corresponding to each of the above scenarios, so as to determine information such as the position and shape of the lesion in the organ. For example, determine the lesion position, shape, and size of gastric ulcer in the stomach, so that medical staff can allocate medical resources based on the position, shape, and size of the existing lesion, etc. Therefore, the image processing model obtained based on the image processing method for medical images provided by this application can improve the accuracy of medical image segmentation, and can further improve the accuracy of lesion judgment, thereby realizing the reasonable allocation of medical resources.
[0106] The solution involved in this application includes an image processing model generation stage and an image processing stage. Figure 3 is a framework diagram of image processing model generation and image processing shown according to an exemplary embodiment, as Figure 3 shown. In the image processing model generation stage, the image processing model generation device 310 obtains an image processing model through a pre-set training sample data set (including sample medical images in the first modality and label images of the target medical object corresponding to the sample medical images); then, an image processing model is generated based on this image processing model. In the image processing stage, the image processing device 320 processes the input target medical image in the first modality based on this image processing model to obtain an image segmentation result of the target medical image in the first modality. The image segmentation result may include the lesion area annotation that can be obtained in medical images in multiple modalities of the medical object corresponding to the target medical image. For example, determine at least one lesion position, shape, etc. information in the medical object corresponding to the target medical image; and / or, process the input target medical image in the first modality to obtain an image generation result of the target medical image in the first modality, and generate a second modality medical image of the medical object corresponding to the target medical image in the first modality to solve the problem of missing images, and at the same time make the image segmentation result interpretable.
[0107] In a possible implementation, when applying the image processing model, if image segmentation of a target medical image is required, the first encoding network and the decoding network in the image processing model can be used. Alternatively, based on the first encoding network and the decoding network in the image processing model, an image segmentation model can be reconstructed, and the parameters in this image segmentation model are consistent with those of the first encoding network and the decoding network in the image processing model. In another possible implementation, when applying the image processing model, if a corresponding medical image of the second modality needs to be generated based on the target medical image of the first modality as input, the first encoding network and the generation network in the image processing model can be used. Alternatively, based on the first encoding network and the generation network in the image processing model, an image generation model can be reconstructed, and the parameters in this image generation model are consistent with those of the first encoding network and the generation network in the image processing model.
[0108] Among them, the above-mentioned image processing model generation device 310 and the image processing device 320 can be computer devices. For example, the computer device can be a fixed computer device such as a personal computer or a server, or the computer device can also be a mobile computer device such as a tablet computer or an e-book reader.
[0109] Optionally, the above-mentioned image processing model generation device 310 and the image processing device 320 can be the same device, or the image processing model generation device 310 and the image processing device 320 can also be different devices. Moreover, when the image processing model generation device 310 and the image processing device 320 are different devices, the image processing model generation device 310 and the image processing device 320 can be devices of the same type. For example, the image processing model generation device 310 and the image processing device 320 can both be servers; or the image processing model generation device 310 and the image processing device 320 can also be devices of different types. For example, the image processing device 320 can be a personal computer or a terminal, while the image processing model generation device 310 can be a server, etc. The specific types of the image processing model generation device 310 and the image processing device 320 are not limited in the embodiments of the present application.
[0110] Figure 4 The flowchart of the image processing method for medical images provided by an exemplary embodiment of the present application is shown. This method is executed by a computing device, and the computer device can be implemented as Figure 1 the server shown, as Figure 4 shown, and the image processing method for medical images includes the following steps:
[0111] Step 410: Invoke the first encoding network in the image processing model to encode the first sample image and obtain the first feature map corresponding to the first sample image. The first sample image is a sample medical image of the first modality of the target medical object.
[0112] Step 420: Invoke the decoding network in the image processing model to decode based on the first feature map and obtain the predicted segmentation image of the first sample image. The predicted segmentation image is used to indicate at least one specified type of region predicted.
[0113] Step 430: Invoke the generation network in the image processing model to generate a predicted generated image based on the first feature map. The predicted generated image is a predicted image of the second modality corresponding to the first sample image.
[0114] Step 440: Determine the function value of the first loss function based on the difference between the predicted segmentation image and the label image.
[0115] In the embodiments of the present application, based on the similarity between the predicted segmentation image and the label image, determine the function value of the first branch function of the first loss function;
[0116] Based on at least one specified type of region in the predicted segmentation image and at least one specified type of region label in the label image, determine the function value of the second branch function of the first loss function;
[0117] Based on the function value of the first branch function and the function value of the second branch function, determine the function value of the first loss function.
[0118] Since in the same medical image, there are large differences in the areas of different divided regions in the medical image, that is, there is a state of extremely unbalanced positive and negative samples. Therefore, in order to balance the importance of positive and negative samples, the Focal Dice loss L FDL can be used as the first branch function in the first loss function; in this first branch function, different divided regions are set with different weights, so that the difficult-to-segment divided regions can obtain higher weights during the segmentation process, enabling the network to focus on learning more difficult categories. The first branch function of the first loss function can be expressed as:
[0119]
[0120] where ω represents the weight of the divided region t, and the parameter represents the power of the Dice t of the divided region t, and schematically β = 2;
[0121] Among them, the Dice coefficient is a metric function used to evaluate the similarity between two samples, with a value range between 0 and 1. The larger the value, the more similar. In the embodiments of the present application, the similarity between these two samples is reflected in the similarity between the predicted segmentation image and the label image, and further reflected in the similarity between each divided region in the predicted segmentation image and the corresponding divided region in the label image. Schematically, the divided regions in the predicted image may include a lesion region, a normal region, and a background region. Generally speaking, the area of the background region accounts for a relatively large proportion of the area of the medical image, the area of the normal region accounts for the second largest proportion of the area of the medical image, and the area of the lesion region accounts for the smallest proportion of the area of the medical image. To make the decoding network pay more attention to the lesion region, in the first branch function, the weight value corresponding to the lesion region can be set to the largest, the weight value corresponding to the normal region is the second largest, and the weight value corresponding to the background region is the smallest. Optionally, the value of the weight is inversely proportional to the area of each divided region in the medical image. Or, the value of the weight corresponds to the type of each divided region. For example, taking this medical image as a myocardial image, its corresponding lesions include myocardial scar and myocardial edema. Then, the weight set corresponding to each divided region in the predicted segmentation image can be set as ω = {1, 1, 1, 0.5}, where the weight of the divided region corresponding to the myocardial scar is 1, the weight of the divided region corresponding to the myocardial edema is 1, the weight of the divided region corresponding to the normal myocardium is 1, and the weight of the divided region corresponding to the background is 0.5. It should be noted that the above setting of weights is only schematic, and the present application does not limit the weight values corresponding to each divided region and the relationship between each weight value.
[0122] In the embodiments of the present application, the mean square error loss function can be used to quantify the difference between the position of at least one specified position region predicted in the predicted segmentation image and the position of at least one specified type region in the label image. The second branch function of this first loss function can be expressed as:
[0123]
[0124] Among them, H and W respectively represent the width and height of the predicted segmentation image (label image), P t and G t respectively represent the predicted position and the label image position of the specified type region t. Optionally, the specified type region in the label image may include at least one of a lesion region, a normal region, and a background region.
[0125] In an embodiment of the present application, the sum of the function values of the first branch function and the function values of the second branch function is obtained as the function value of the first loss function; further, to balance the roles of the first branch function and the second branch function, different weight values can be set for the first branch function and the second branch function. Schematically, the first loss function can be expressed as:
[0126] L seg = L FDL + λL mse
[0127] where λ represents the weight value of the second branch function relative to the first branch function; schematically, the value of λ can be 100.
[0128] Step 450: Determine the function value of the second loss function based on the difference between the predicted generated image and the second sample image; the second sample image is a sample medical image of the second modality of the target medical object.
[0129] Schematically, the second loss function can be expressed as:
[0130]
[0131] where x identifies the predicted generated image, x' represents the sample medical image of the second modality of the target medical object, and H and W respectively represent the width and height of the predicted generated image (sample medical image of the second modality).
[0132] Step 460: Train the image processing model based on the function value of the first loss function and the function value of the second loss function.
[0133] In an embodiment of the present application, the first loss function and the second loss function update the parameters of different network combinations in the image processing model;
[0134] Optionally, update the parameters of the first encoding network and the parameters of the decoding network based on the function value of the first loss function;
[0135] Update the parameters of the first encoding network and the parameters of the generation network based on the function value of the second loss function.
[0136] That is to say, the function value of the first loss function and the function value of the second loss function both guide the parameter update of the first encoding network in the image processing model; therefore, the generation network is used to assist the generation of the image segmentation model (a model including an encoding network and a decoding network); or rather, the decoding network is used to assist the generation of the image generation model (a model including an encoding network and a generation network).
[0137] In an embodiment of the present application, a third loss function may also be introduced to train the image processing model. The third loss function is used to indicate the authenticity of the predicted generated image. This process can be implemented as follows:
[0138] Call the discriminator to discriminate the predicted generated image to obtain the discrimination result of the predicted generated image.
[0139] Based on the discrimination result, determine the function value of the third loss function. The discrimination result is used to indicate whether the predicted generated image is a real image.
[0140] Schematically, the third loss function can be expressed as:
[0141]
[0142] Where G represents the generation network, D represents the discrimination network (discriminator), represents the real image distribution, represents the false image distribution.
[0143] In the above case, training the image processing model includes: training the image processing model based on the function value of the first loss function, the function value of the second loss function, and the function value of the third loss function.
[0144] In an embodiment of the present application, based on the function value of the third loss function, update the parameters of the first encoding network and the parameters of the generation network.
[0145] Among them, the discriminator can be pre-trained; or, the parameters in the discriminator can be updated based on the function value of the third loss function. In this case, the input of the discriminator also includes the sample medical image of the second modality corresponding to the target medical object to train the discriminator. The discriminator plays an auxiliary role in training the generation network to improve the quality of the images generated by the generation network.
[0146] In summary, the image processing method for medical images provided by the embodiments of the present application obtains the multi-modal sample medical images corresponding to the target medical object, and the label image including the region labels of the specified type corresponding to the target medical image. Based on the first sample image in the multi-modal sample medical images, generate a predicted segmentation image and a predicted generated image. Based on the difference between the predicted segmentation image and the label image, and the difference between the predicted generated image and the second sample image corresponding to the target medical object, train the image processing model including the first encoding network, the decoding network, and the generation network, so that the trained image processing model can obtain the features of the multi-modal medical images based on a single-modal medical image, making the information included in the obtained medical image segmentation result more comprehensive and improving the segmentation effect of the medical images.
[0147] Furthermore, the image processing model obtained through training can generate medical images of other modalities based on medical images of a single modality, thereby solving the problem of missing images in the process of medical image analysis.
[0148] Optionally, in order to improve the accuracy of model training and reduce the error caused by the class imbalance problem, a prior constraint image of the image processing model can be obtained based on a third sample image, and the prior constraint image is used to indicate the predicted position of the target medical object in the sample image. Among them, the third sample image can be one of the first sample image and the second sample image, or the third sample image can also be a sample medical image of the third modality of the target medical image; for the CMR image obtained by cardiac magnetic resonance, the third sample image can be a bSSFP image that can capture cardiac motion and present clear boundaries. Compared with T2-weighted images and LGE images, more accurate myocardial position and shape can be obtained based on the bSSFP image. Therefore, the prior constraint image obtained based on the bSSFP image is more accurate in predicting the position of the target medical object in the sample image.
[0149] In the above situation, based on Figure 4 the image processing method for medical images shown in the above embodiments Figure 5 shows a flowchart of the image processing method for medical images shown in an exemplary embodiment of the present application, as Figure 5 shown, the method includes the following steps:
[0150] Step 510, obtain a prior constraint image of the image processing model based on a third sample image; the third sample image is a sample medical image of the third modality of the target medical object; the prior constraint image is used to indicate the position of the target medical object in the third sample image.
[0151] The position of the target medical object indicated by the prior constraint image in the third sample image can indicate the position of the target medical image in other sample images (including the first sample image and the second sample image).
[0152] In an embodiment of the present application, a semantic segmentation network (U-Net) can be called to process the third sample image to obtain a prior constraint object of the image processing model;
[0153] Among them, the semantic segmentation network is obtained through training based on a sample image set; the sample image set includes a fourth sample image and an approximate label corresponding to the fourth sample image, and the fourth sample image refers to a sample medical image of the third modality corresponding to other medical objects; the approximate label is used to indicate the position of other medical objects in the fourth sample image.
[0154] Taking the type of other medical objects as the myocardium as an example, in the myocardial medical image, the myocardial edema area and the myocardial scar area account for a relatively low proportion of the medical image, and the corresponding areas do not overlap with each other. Therefore, the image combining the normal myocardial area, the myocardial edema area, and the myocardial scar area can be obtained as the approximate label corresponding to the medical image; Figure 6 shows a schematic diagram of the synthesis of the approximate label shown in an exemplary embodiment of the present application, as Figure 6 shown, extract the myocardial edema area 610 and the myocardial scar area 620 from the label image, and combine them with the normal myocardial area 630 to generate an approximate label 640; obtain this approximate label as the label of the first sample object, and train the semantic segmentation image so that the trained semantic segmentation image can process the input third sample image to obtain the prior constraint image corresponding to the third sample image.
[0155] Step 520, call the second encoding network in the image processing model, and encode based on the prior constraint image to obtain the second feature map corresponding to the third sample image.
[0156] In the embodiment of the present application, the image processing model may further include a second encoding network. In order to alleviate the overfitting problem caused by too large network parameters, in the embodiment of the present application, the parameters of the second encoding network can be set to be the same as those of the first encoding network, that is, the parameter weight values in the second encoding network and the first encoding network are shared.
[0157] Optionally, in order to further reduce the influence of the background area on model training, after obtaining the prior constraint image, the prior constraint image can be cropped based on the position of the target medical object; then, call the second encoding network in the image processing model to encode the cropped prior constraint image to obtain the second feature map corresponding to the third sample image.
[0158] Adapted to the size of the prior constraint image, preprocess other sample images (including the first sample image and the second sample image), that is, crop other sample images, and ensure that the position of the target medical object in other sample images is similar to the position of the target medical object in the prior constraint image within a specified error range. Optionally, it can be ensured that the position of the target medical object in other image samples is at the center of the image within a specified error range with the position of the target medical object in the prior constraint image. Taking the target medical object as the myocardium as an example, for the prior constraint image, since the myocardium is a circular symmetric tissue, the prior constraint image is cropped according to the center of the approximate label obtained above; for other sample images, it can be cropped according to the position of the specified type area in the label image, or it can also be cropped according to the center of the approximate label above. The present application does not limit the cropping basis of other sample images.
[0159] Optionally, since the data ranges corresponding to different cases vary greatly, the cropped prior image can be further processed. For example, histogram equalization and random gamma methods can be applied to further balance the data distribution after uniformly setting the window level and window width.
[0160] In addition, before processing the sample image, the sample image can be first subjected to data augmentation processing, and the data augmentation processing method includes methods such as random rotation, random cropping, and random scaling.
[0161] Step 530: Invoke the first encoding network in the image processing model to encode the first sample image to obtain a first feature map corresponding to the first sample image; the first sample image is a sample medical image of the first modality of the target medical object.
[0162] If the image input to the second encoding network is a prior constraint image, the first sample image is the original first sample image; if the image input to the second encoding network is a cropped prior constraint image, the first sample image is the cropped first sample image; that is, the sizes of the images input to each encoding network are kept consistent.
[0163] In the embodiment of the present application, to make the generated predicted segmentation image and / or predicted generation image more accurate, the image processing model can be built based on a butterfly network architecture. The first encoding network in the butterfly network includes N encoding layers, and the N encoding layers are connected in pairs; the decoding network in the butterfly network includes N decoding layers, and the N decoding layers are connected in pairs. The N decoding layers in the decoding network correspond one-to-one with the N encoding layers in the first encoding network. Figure 7 shows a schematic structural diagram of an image processing model shown in an exemplary embodiment of the present application, as Figure 7 shown, the first encoding network 710 includes N encoding layers, and the decoding network 730 includes N decoding layers. Among them, the first encoding layer 711 in the first encoding network 710 corresponds one-to-one with the Nth decoding layer 733 in the decoding network 730, and the second encoding layer 712 in the first encoding network 710 corresponds one-to-one with the (N - 1)th decoding layer 732 in the decoding network 730; and so on, the Nth encoding layer 713 in the first encoding network 710 corresponds one-to-one with the first decoding layer 731 in the decoding network 730. Optionally, as Figure 7 shown, the image processing model includes a second encoding layer 720, and the second encoding network 720 can also include N encoding layers, and the N encoding layers in the second encoding layer are connected in pairs.
[0164] When the image processing model is Figure 7When building the model based on the butterfly network architecture as shown, the process of calling the first encoding network in the image processing model to encode the first sample image and obtain the first feature map corresponding to the first sample image can be implemented as follows:
[0165] Obtain the first image pyramid corresponding to the first sample image. The first image pyramid is a set of images obtained by downsampling the first sample image according to a specified gradient. The first image pyramid contains N first images to be processed;
[0166] Input the N first images to be processed into the corresponding encoding layers respectively, and encode the N first images to be processed to obtain N first feature maps corresponding to the first sample image;
[0167] Among them, in response to the target encoding layer being a non-first encoding layer among the N encoding layers, the input of the target encoding layer further includes the first feature map output by the previous encoding layer.
[0168] There are differences in the resolutions of the first images to be processed in the first image pyramid. Each image in the first image pyramid corresponds to a side input path, and each side input path is used to input the corresponding first image to be processed into the corresponding encoding layer in the first encoding network. As Figure 7 shown, the first image pyramid 750 contains N first images to be processed, and each first image to be processed corresponds to a side input path. For a non-first encoding layer in the first encoding network 710, its input includes the first image to be processed input by the corresponding side input path and the first feature map output by the previous encoding layer of this encoding layer.
[0169] Correspondingly, taking the image input to the second encoding network 720 as the cropped prior constraint image as an example, for the second encoding network 720, the process of obtaining the second feature map includes: obtaining the second image pyramid corresponding to the cropped prior constraint image. The second image pyramid is a set of images obtained by downsampling the cropped prior constraint image according to the specified gradient. The second image pyramid contains N second images to be processed;
[0170] Input the N second images to be processed into the corresponding encoding layers in the second encoding network respectively, and encode the N second images to be processed to obtain N encoding results corresponding to the cropped prior constraint image;
[0171] Merge the N encoding results to obtain the second feature map corresponding to the cropped prior constraint image;
[0172] Among them, when the encoding layer in the second encoding network is a non-first encoding layer, the input of this encoding layer further includes the encoding result output by the previous encoding layer.
[0173] The resolutions of the respective second images to be processed in the second image pyramid are different. Each second image to be processed in the second image pyramid corresponds to a side input path, and each side input path is used to input the corresponding second image to be processed into the corresponding encoding layer in the second encoding network. As Figure 7 shown, the second image pyramid 760 includes N second images to be processed, and each second image to be processed corresponds to a side input path. For non-first encoding layers in the second encoding network 720, its input includes the second image to be processed input by the corresponding side input path and the encoding result output by the previous encoding layer of this encoding layer.
[0174] In the embodiments of the present application, the structure of the encoding layer in the encoding network (the first encoding network / the second encoding network) can adopt a convolutional layer structure of two layers of "3x3 separable convolution + ReLU activation function + Dropout operation". Figure 8 shows a schematic structural diagram of the encoding layer shown in an exemplary embodiment of the present application. As Figure 8 shown, the encoding layer in the encoding network includes a convolutional layer 810 and a convolutional layer 820. A channel attention module 830 is added between the two convolutional layers in a residual connection manner; in the channel attention module 830, the feature map obtained after passing through the convolutional layer 810 is compressed in the spatial dimension by using max pooling and average pooling; the shared network consists of a multi-layer perceptron (MLP). By perceiving, concatenating and merging, and activation function processing of the compressed feature map, a channel attention feature map is obtained; the channel attention feature map is multiplied by the input of the channel attention module and added to the output of the convolutional layer 810 of the encoding network to form a residual structure, and an intermediate feature map is obtained, which is then followed by a convolutional layer 820 for downsampling the intermediate feature map using a convolutional layer with a specified stride to obtain the feature map (the first feature map / the encoding result) output by the convolutional layer 820. Schematically, the specified stride can be 2; optionally, to better extract the features of the input image, the convolutional layer in the encoding network can be replaced with a depthwise separable convolutional layer.
[0175] Step 540, merge the first feature map and the second feature map to obtain a comprehensive feature map.
[0176] When the image processing model is Figure 7 the model built based on the butterfly network architecture shown, this comprehensive feature map is the merged result of the first feature map output by the Nth encoding layer of the first encoding network 710 and the second feature map output by the second encoding network 720.
[0177] Step 550, call the decoding network in the image processing module, and decode based on the comprehensive feature map to obtain the predicted segmentation image of the first sample image.
[0178] When the image processing model is Figure 7 the model built based on the butterfly network architecture as shown, the process of calling the decoding network in the image processing module and decoding based on the comprehensive feature map to obtain the predicted segmentation image of the first sample image can be implemented as follows:
[0179] Input the N first feature maps into the corresponding decoding layers of the decoding network respectively, decode the N first feature maps to obtain N decoding results; the N decoding results have the same resolution;
[0180] Merge the N decoding results to obtain the predicted segmentation image of the first sample image;
[0181] Among them, in response to the target decoding layer being a non-first decoding layer among the N decoding layers, the input of the target decoding layer further includes the decoding result output by the previous decoding layer.
[0182] In the embodiments of the present application, the structure of the decoding layer in the decoding network can adopt the structure of a two-layer convolutional layer of "3x3 separable convolution + ReLU activation function + Dropout operation", Figure 9 shows the structural schematic diagram of the decoding layer shown in an exemplary embodiment of the present application, as Figure 9 shown, the decoding layer in the decoding network includes a convolutional layer 910 and a convolutional layer 920. A spatial attention module 930 is added between the two convolutional layers in a residual connection manner. The spatial attention module 930 mainly focuses on position information; in the spatial attention module 930, through using max pooling and average pooling to process in the channel dimension, a feature map is obtained, then cascaded and convolved through a convolutional layer, and then processed through an activation function to obtain a spatial attention feature map; multiply the spatial attention feature map by the input of the spatial attention module and add it to the output of the convolutional layer 910 of the decoding network to form a residual structure to obtain an intermediate feature map, and then connect a convolutional layer 920 behind it, which is used to downsample the intermediate feature map using a convolutional layer with a specified stride to obtain the decoding result output by the convolutional layer 920.
[0183] Step 560, call the generation network in the image processing model to generate a predicted generated image based on the comprehensive feature map.
[0184] As Figure 7 shown, the image processing model may include a generation network 740 for generating a predicted generated image 741 from the comprehensive feature map.
[0185] Step 570: Train the image processing model based on the differences between the predicted segmentation image and the label image, and between the predicted generated image and the second sample image. The second sample image is a sample medical image of the second modality of the target medical object, and the label image corresponds to the target medical object and is used to indicate at least one specified type of region.
[0186] The image processing model with the butterfly-shaped network architecture provided in this application can combine deep semantic information and ground layer position information, thereby reducing the vanishing gradient while ensuring the network width. On the other hand, through the supervision of multi-scale and multi-resolution input images, more image features can be obtained, and thus better image segmentation effects and / or image generation effects can be achieved.
[0187] In summary, the image processing method for medical images provided in the embodiments of this application obtains sample medical images of multiple modalities corresponding to the target medical object, and label images corresponding to the target medical images and containing labels of specified type regions. Generate a predicted segmentation image and a predicted generated image based on the first sample image in the multi-modal sample medical images, and based on the differences between the predicted segmentation image and the label image, and between the predicted generated image and the second sample image corresponding to the target medical object, train the image processing model including the first encoding network, the decoding network, and the generation network, so that the trained image processing model can obtain the features of multi-modal medical images based on a single-modal medical image, making the information contained in the obtained medical image segmentation result more comprehensive and improving the segmentation effect of medical images;
[0188] Furthermore, based on the trained image processing model, medical images of other modalities can be generated based on a single-modal medical image, thereby solving the problem of missing images in the medical image analysis process.
[0189] In a possible implementation, when training an image processing model, the training results of two image processing models can be combined to obtain a final image processing model. Schematically, the input of the first image processing model is a first sample image, which is a sample medical image of the first modality of a target medical object. Using the label image corresponding to the target medical object and the second sample image as labels, the first image processing model is trained to obtain a trained first image processing model. The second sample image is a sample medical image of the second modality of the target medical object. The first image processing model is used to generate a predicted segmentation image corresponding to the input medical image of the first modality, and / or generate a medical generation image of the second modality corresponding to the input medical object of the first modality. The input of the second image processing model is the second sample image. Using the label image corresponding to the target medical object and the first sample image as labels, the second image processing model is trained to obtain a trained second image processing model. The second image processing model is used to generate a predicted segmentation image corresponding to the input medical image of the second modality, and / or generate a medical generation image of the first modality corresponding to the input medical object of the second modality. Among them, if the input images of the two image processing models are medical images of different modalities of the same medical object, the predicted segmentation images obtained after processing by the first image processing model and the second image repair model are the same, or the error is within a specified threshold range.
[0190] Optionally, to reduce network parameters, the parameters of the encoding network and the decoding network of the first image processing model can share weights with the parameters of the encoding network and the decoding network in the second image processing model. This process can be carried out during the model training process or after the model training is completed. Schematically, weight sharing can be implemented as follows: replacing the parameters of the encoding network and the decoding network in one of the image processing models with the parameters of the encoding network and the decoding network in the other image processing model, or taking the average of the parameters of the encoding networks in the two image processing models and the average of the parameters of the decoding networks, and respectively replacing them in the encoding networks and the decoding networks of the two image processing models. The way of weight sharing in this application is not limited.
[0191] Schematically, taking the segmentation of myocardial scar and myocardial edema as an example, the application process of the image processing model generated based on this application will be described. Figure 10 FIG. shows a schematic diagram of the application process of the image processing model shown in an exemplary embodiment of this application. This process can be implemented on a terminal or a server deployed with the image processing model, or on a terminal or a server deployed with an image segmentation model constructed based on the image processing model, such as Figure 10 As shown, based on cardiac magnetic resonance technology, CMR images corresponding to the same medical object are obtained. Figure 10Among them are bSSFP images, T2-weighted images, and LGE images; in the first stage, the bSSFP image 1010 is input into the U-Net network 1020 to obtain the prior constraint image 1030 output by the U-Net network, which is used to indicate the position information of the medical object in the CMR image; based on the central position of the medical object in the prior constraint image, the prior constraint image and the T2-weighted image are cropped, and the cropped T2-weighted image and the cropped prior constraint image are input into the first image processing model 1040 corresponding to the T2 mode to obtain the first predicted segmentation image 1050 output by the first image processing model, which contains the position information of myocardial scar and the position information of myocardial edema; the cropped LGE image and the cropped prior constraint image are input into the second image processing model 1060 corresponding to the LGE mode to obtain the second predicted segmentation image 1070 output by the second image processing model; to further improve the accuracy of the predicted segmentation image, the first predicted segmentation image and the second predicted segmentation image are merged to obtain the segmentation image 1080 of myocardial scar and myocardial edema corresponding to the medical object; in addition, when this process is implemented in a terminal or server equipped with an image processing model, a corresponding LGE image can be generated based on the T2-weighted image, and a corresponding T2-weighted image can be generated based on the LGE image.
[0192] Figure 11 FIG. shows a block diagram of an image processing device for medical images according to an exemplary embodiment of the present application, as Figure 11 shown, the device includes:
[0193] A first encoding module 1110, configured to call a first encoding network in the image processing model to encode a first sample image and obtain a first feature map corresponding to the first sample image; the first sample image is a sample medical image of a first modality of a target medical object;
[0194] A decoding module 1120, configured to call a decoding network in the image processing model to decode based on the first feature map and obtain a predicted segmentation image of the first sample image; the predicted segmentation image is used to indicate at least one specified type of region predicted;
[0195] A generation module 1130, configured to call a generation network in the image processing model to generate a predicted generation image based on the first feature map; the predicted generation image is a predicted image of a second modality corresponding to the first sample image;
[0196] A model training module 1140 is configured to train the image processing model based on the differences between the predicted segmentation image and the label image, and between the predicted generation image and the second sample image; the second sample image is a sample medical image of a second modality of the target medical object; the label image corresponds to the target medical object and is used to indicate at least one specified type of region.
[0197] In a possible implementation, the model training module 1140 includes:
[0198] A first determination sub-module is configured to determine the function value of a first loss function based on the difference between the predicted segmentation image and the label image;
[0199] A second determination sub-module is configured to determine the function value of a second loss function based on the difference between the predicted generation image and the second sample image;
[0200] A model training sub-module is configured to train the image processing model based on the function value of the first loss function and the function value of the second loss function.
[0201] In a possible implementation, the model training sub-module is configured to update the parameters of the first encoding network and the parameters of the decoding network based on the function value of the first loss function;
[0202] Update the parameters of the first encoding network and the parameters of the generation network based on the function value of the second loss function.
[0203] In a possible implementation, the first determination sub-module includes:
[0204] A first determination unit is configured to determine the function value of a first branch function of the first loss function based on the similarity between the predicted segmentation image and the label image;
[0205] A second determination unit is configured to determine the function value of a second branch function of the first loss function based on the positions of at least one specified type of region predicted in the predicted segmentation image and the positions of at least one specified type of region in the label image;
[0206] A third determination unit is configured to determine the function value of the first loss function based on the function value of the first branch function and the function value of the second branch function.
[0207] In a possible implementation, the first determining unit is configured to obtain the weight values respectively corresponding to the respective divided regions in the predicted segmentation image; each divided region in the predicted segmentation image includes the at least one specified type region;
[0208] Based on the weight values respectively corresponding to the respective divided regions in the predicted segmentation image, and the similarity between the respective divided regions in the predicted segmentation image and the respective divided regions in the label image, determine the function value of the first branch function of the first loss function.
[0209] In a possible implementation, the apparatus further includes:
[0210] A discrimination module, configured to call a discriminator to discriminate the predicted generated image, and obtain the discrimination result of the predicted generated image;
[0211] A third determining module, configured to determine the function value of the third loss function based on the discrimination result; the discrimination result is used to indicate whether the predicted generated image is a real image;
[0212] The model training module 1140 is configured to train the image processing model based on the function value of the first loss function, the function value of the second loss function, and the function value of the third loss function.
[0213] In a possible implementation, the first encoding network includes N encoding layers, and the N encoding layers are connected in pairs, N≥2 and is a positive integer;
[0214] The first encoding module 1110 includes:
[0215] A set obtaining sub-module, configured to obtain a first image pyramid corresponding to the first sample image, the first image pyramid being an image set obtained by downsampling the first sample image according to a specified gradient, and the first image pyramid includes N first images to be processed;
[0216] An encoding sub-module, configured to respectively input the N first images to be processed into corresponding encoding layers, and encode the N first images to be processed to obtain N first feature maps corresponding to the first sample image;
[0217] Wherein, in response to the target encoding layer being a non-first encoding layer among the N encoding layers, the input of the target encoding layer further includes the first feature map output by the previous encoding layer.
[0218] In a possible implementation, the decoding network in the image processing model includes N decoding layers, and the N decoding layers are connected in pairs, and the N decoding layers correspond to the N encoding layers one by one;
[0219] The decoding module 1120 includes:
[0220] A decoding sub-module, configured to respectively input the N first feature maps into corresponding decoding layers of the decoding network, decode the N first feature maps, and obtain N decoding results; the N decoding results have the same resolution;
[0221] A merging sub-module, configured to merge the N decoding results to obtain a predicted segmentation image of the first sample image;
[0222] Wherein, in response to the target decoding layer being a non-first decoding layer among the N decoding layers, the input of the target decoding layer further includes a decoding result output by the previous decoding layer.
[0223] In a possible implementation manner, the apparatus further includes:
[0224] An image acquisition module, configured to acquire a prior constraint image of the image processing model based on a third sample image; the third sample image is a sample medical image of a third modality of the target medical object; the prior constraint image is used to indicate the position of the target medical object in the third sample image;
[0225] A second encoding module, configured to call a second encoding network in the image processing model, and encode based on the prior constraint image to obtain a second feature map corresponding to the third sample image;
[0226] A merging module, configured to merge the first feature map and the second feature map to obtain a comprehensive feature map;
[0227] The decoding module 1120 is configured to call a decoding network in the image processing module, and decode based on the comprehensive feature map to obtain the predicted segmentation image of the first sample image;
[0228] The generation module 1130 is configured to call a generation network in the image processing model, and generate a predicted generation image based on the comprehensive feature map.
[0229] In a possible implementation manner, the apparatus further includes:
[0230] A cropping module, configured to crop the prior constraint image based on the position of the target medical object;
[0231] The second encoding module is configured to call a second encoding network in the image processing model, and encode the cropped prior constraint image to obtain a second feature map corresponding to the third sample image.
[0232] In a possible implementation, the image acquisition module is configured to call a semantic segmentation network to process the third sample image and obtain a prior constraint image of the image processing model.
[0233] In a possible implementation, the parameters in the second encoding network share the same weight as the parameters in the first encoding network.
[0234] In summary, the image processing device for medical images provided by the embodiments of the present application obtains multi-modal sample medical images corresponding to a target medical object and a label image including region labels of a specified type corresponding to the target medical image, generates a predicted segmentation image and a predicted generated image based on a first sample image in the multi-modal sample medical images, and trains an image processing model including a first encoding network, a decoding network, and a generation network based on the differences between the predicted segmentation image and the label image and between the predicted generated image and a second sample image corresponding to the target medical object. As a result, the trained image processing model can obtain the features of multi-modal medical images based on a single-modal medical image, making the information included in the obtained medical image segmentation result more comprehensive and improving the segmentation effect of medical images.
[0235] Furthermore, based on the trained image processing model, other modal medical images can be generated based on a single-modal medical image, thereby solving the problem of missing images in the process of medical image analysis.
[0236] Figure 12 FIG. shows a block diagram of the structure of a computer device 1200 shown in an exemplary embodiment of the present application. The computer device may be implemented as the server in the above solution of the present application. The computer device 1200 includes a central processing unit (CPU) 1201, a system memory 1204 including a random access memory (RAM) 1202 and a read-only memory (ROM) 1203, and a system bus 1205 connecting the system memory 1204 and the central processing unit 1201. The computer device 1200 also includes a mass storage device 1206 for storing an operating system 1209, application programs 1210, and other program modules 1211.
[0237] The large-capacity storage device 1206 is connected to the central processing unit 1201 through a large-capacity storage controller (not shown) connected to the system bus 1205. The large-capacity storage device 1206 and its associated computer-readable medium provide non-volatile storage for the computer device 1200. That is to say, the large-capacity storage device 1206 may include computer-readable media (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM) drive.
[0238] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, erasable programmable read-only memory (EPROM), electrically-erasable programmable read-only memory (EEPROM) flash memory or other solid-state storage technologies, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic tape cartridges, tapes, disk storage or other magnetic storage devices. Of course, those skilled in the art will know that the computer storage media is not limited to the above several types. The above system memory 1204 and large-capacity storage device 1206 may be collectively referred to as memory.
[0239] According to various embodiments of the present disclosure, the computer device 1200 may also be run by connecting to a remote computer on the network through a network such as the Internet. That is, the computer device 1200 may be connected to the network 1208 through a network interface unit 1207 connected to the system bus 1205, or in other words, the network interface unit 1207 may also be used to connect to other types of networks or remote computer systems (not shown).
[0240] The memory further includes at least one instruction, at least one program, a code set or an instruction set, and the at least one instruction, at least one program, code set or instruction set is stored in the memory. The central processing unit 1201 implements all or part of the steps in the image processing method for medical images shown in the above various embodiments by executing the at least one instruction, at least one program, code set or instruction set.
[0241] Figure 13The block diagram of a computer device 1300 provided by an exemplary embodiment of the present application is shown. The computer device 1300 may be implemented as the above-mentioned terminal, such as: a smart phone, a tablet computer, a notebook computer or a desktop computer. The computer device 1300 may also be referred to by other names such as user equipment, portable terminal, laptop terminal, desktop terminal, etc.
[0242] Generally, the computer device 1300 includes: a processor 1301 and a memory 1302.
[0243] The processor 1301 may include one or more processing cores, such as a 4-core processor, a 13-core processor, etc. The processor 1301 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), PLA (Programmable Logic Array). The processor 1301 may also include a main processor and a co-processor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the co-processor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1301 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1301 may also include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.
[0244] The memory 1302 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1302 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices, flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1302 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 1301 to implement all or part of the steps of the image processing method for medical images provided in the method embodiments of the present application.
[0245] In some embodiments, the computer device 1300 may further optionally include: a peripheral device interface 1303 and at least one peripheral device. The processor 1301, the memory 1302, and the peripheral device interface 1303 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1303 through a bus, signal lines, or a circuit board. Specifically, the peripheral devices include at least one of a radio frequency circuit 1304, a display screen 1305, a camera assembly 1306, an audio circuit 1307, a positioning component 1308, and a power supply 1309.
[0246] The peripheral device interface 1303 can be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 1301 and the memory 1302. In some embodiments, the processor 1301, the memory 1302, and the peripheral device interface 1303 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1301, the memory 1302, and the peripheral device interface 1303 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.
[0247] In some embodiments, the computer device 1300 further includes one or more sensors 1310. The one or more sensors 1310 include but are not limited to: an acceleration sensor 1311, a gyroscope sensor 1312, a pressure sensor 1313, a fingerprint sensor 1314, an optical sensor 1315, and a proximity sensor 1316.
[0248] Those skilled in the art can understand that Figure 13 the structure shown in does not constitute a limitation on the computer device 1300, and it may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component arrangement.
[0249] In an exemplary embodiment, there is also provided a computer-readable storage medium for storing at least one instruction, at least one segment of program, a code set, or an instruction set, and the at least one instruction, the at least one segment of program, the code set, or the instruction set is loaded and executed by the processor to implement all or part of the steps in the above-mentioned image processing method for medical images. For example, the computer-readable storage medium may be a Read-Only Memory (ROM), a Random Access Memory (RAM), a Compact Disc Read-Only Memory (CD-ROM), magnetic tape, a floppy disk, and an optical data storage device, etc.
[0250] In an exemplary embodiment, a computer program product or a computer program is further provided. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes all or part of the steps of the method shown in any of the above Figure 2 、 Figure 4 or Figure 5 embodiments.
[0251] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and embodiments are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0252] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. An image processing method for medical images, characterized in that, The method includes: Invoking a first encoding network in an image processing model to encode a first sample image, and obtaining a first feature map corresponding to the first sample image; the first sample image is a sample medical image of a first modality of a target medical object; the modality of the sample medical image is used to indicate the acquisition method of the sample medical image; Invoking a decoding network in the image processing model to decode based on the first feature map, and obtaining a predicted segmentation image of the first sample image; the predicted segmentation image is used to indicate at least one specified type of region predicted; Invoking a generation network in the image processing model to generate a predicted generation image based on the first feature map; the predicted generation image is a predicted image of a second modality corresponding to the first sample image; Training the image processing model based on the difference between the predicted segmentation image and a label image, and the difference between the predicted generation image and a second sample image; the second sample image is a sample medical image of a second modality of the target medical object; the label image corresponds to the target medical object and is used to indicate an image of at least one specified type of region.
2. The method according to claim 1, characterized in that, The training of the image processing model based on the difference between the predicted segmentation image and the label image, and the difference between the predicted generation image and the second sample image includes: Determining the function value of a first loss function based on the difference between the predicted segmentation image and the label image; Determining the function value of a second loss function based on the difference between the predicted generation image and the second sample image; Training the image processing model based on the function value of the first loss function and the function value of the second loss function.
3. The method according to claim 2, wherein The training of the image processing model based on the function value of the first loss function and the function value of the second loss function includes: Updating the parameters of the first encoding network and the parameters of the decoding network based on the function value of the first loss function; Updating the parameters of the first encoding network and the parameters of the generation network based on the function value of the second loss function.
4. The method according to claim 2, wherein The determining the function value of the first loss function based on the difference between the predicted segmentation image and the label image includes: Determining the function value of a first branch function of the first loss function based on the similarity between the predicted segmentation image and the label image; Determining the function value of a second branch function of the first loss function based on the positions of at least one specified type of region predicted in the predicted segmentation image and the positions of at least one specified type of region in the label image; Determining the function value of the first loss function based on the function value of the first branch function and the function value of the second branch function.
5. The method according to claim 4, wherein The determining the function value of the first branch function of the first loss function based on the similarity between the predicted segmentation image and the label image includes: Obtaining the weight values respectively corresponding to the respective divided regions in the predicted segmentation image; each divided region in the predicted segmentation image includes the at least one specified type of region; Determine the function value of the first branch function of the first loss function based on the weight values corresponding to the respective divided regions in the predicted segmentation image and the similarity between the respective divided regions in the predicted segmentation image and the respective divided regions in the label image.
6. The method according to claim 2, wherein The method further includes: Invoking a discriminator to discriminate the predicted generated image to obtain the discrimination result of the predicted generated image; Determine the function value of the third loss function based on the discrimination result; the discrimination result is used to indicate whether the predicted generated image is a real image; The training of the image processing model based on the function value of the first loss function and the function value of the second loss function includes: Train the image processing model based on the function value of the first loss function, the function value of the second loss function, and the function value of the third loss function.
7. The method according to claim 1, wherein The first encoding network includes N encoding layers, and the N encoding layers are connected pairwise, N≥2 and is a positive integer; The invoking the first encoding network in the image processing model to encode the first sample image to obtain the first feature map corresponding to the first sample image includes: Obtain the first image pyramid corresponding to the first sample image, where the first image pyramid is a set of images obtained by downsampling the first sample image according to a specified gradient, and the first image pyramid includes N first images to be processed; Input the N first images to be processed into the corresponding encoding layers respectively to encode the N first images to be processed to obtain N first feature maps corresponding to the first sample image; Wherein, in response to the target encoding layer being a non-first encoding layer among the N encoding layers, the input of the target encoding layer further includes the first feature map output by the previous encoding layer.
8. The method according to claim 7, wherein The decoding network in the image processing model includes N decoding layers, and the N decoding layers are connected pairwise, and the N decoding layers correspond to the N encoding layers one by one; The invoking the decoding network in the image processing model to decode based on the first feature map to obtain the predicted segmentation image of the first sample image includes: Input the N first feature maps into the corresponding decoding layers of the decoding network respectively to decode the N first feature maps to obtain N decoding results; the N decoding results have the same resolution; Merge the N decoding results to obtain the predicted segmentation image of the first sample image; Wherein, in response to the target decoding layer being a non-first decoding layer among the N decoding layers, the input of the target decoding layer further includes the decoding result output by the previous decoding layer.
9. The method according to claim 1, characterized in that, The method further includes: Based on the third sample image, obtain the prior constraint image of the image processing model; the third sample image is the sample medical image of the third modality of the target medical object; the prior constraint image is used to indicate the position of the target medical object in the third sample image; Invoke the second encoding network in the image processing model to encode based on the prior constraint image to obtain the second feature map corresponding to the third sample image; Merge the first feature map and the second feature map to obtain a comprehensive feature map; The step of calling the decoding network in the image processing model to decode based on the first feature map to obtain the predicted segmentation image of the first sample image includes: Call the decoding network in the image processing model to decode based on the comprehensive feature map to obtain the predicted segmentation image of the first sample image; The step of calling the generation network in the image processing model to generate a predicted generated image based on the first feature map includes: Call the generation network in the image processing model to generate the predicted generated image based on the comprehensive feature map.
10. The method according to claim 9, wherein Before calling the second encoding network in the image processing model to encode the prior constraint image to obtain the second feature map corresponding to the third sample image, it further includes: Crop the prior constraint image based on the position of the target medical object; The step of calling the second encoding network in the image processing model to encode the prior constraint image to obtain the second feature map corresponding to the third sample image includes: Call the second encoding network in the image processing model to encode the cropped prior constraint image to obtain the second feature map corresponding to the third sample image.
11. The method according to claim 9, wherein The step of obtaining the prior constraint image of the image processing model based on the third sample image includes: Call the semantic segmentation network to process the third sample image to obtain the prior constraint image of the image processing model.
12. The method according to claim 9, characterized in that The parameters in the second encoding network share the weight with the parameters in the first encoding network.
13. An image processing apparatus for medical images, characterized in that, The device includes: A first encoding module, configured to call the first encoding network in the image processing model to encode the first sample image to obtain the first feature map corresponding to the first sample image; the first sample image is a sample medical image of the first modality of the target medical object; the modality of the sample medical image is used to indicate the acquisition method of the sample medical image; A decoding module, configured to call the decoding network in the image processing model to decode based on the first feature map to obtain the predicted segmentation image of the first sample image; the predicted segmentation image is used to indicate at least one specified type of region; A generation module, configured to call the generation network in the image processing model to generate a predicted generated image based on the first feature map; the predicted generated image is a predicted image of the second modality corresponding to the first sample image; A model training module, configured to train the image processing model based on the difference between the predicted segmentation image and the label image, and the difference between the predicted generated image and the second sample image; the second sample image is a sample medical image of the second modality of the target medical object; the label image is an image corresponding to the target medical object and used to indicate at least one specified type of region.
14. A computer device, characterized in that, The computer device includes a processor and a memory. The memory stores at least one instruction, at least one program, a code set, or an instruction set. The at least one instruction, the at least one program, the code set, or the instruction set is loaded and executed by the processor to implement the image processing method for medical images as described in any one of claims 1 to 12.
15. A computer-readable storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium. The computer program is loaded and executed by the processor to implement the image processing method for medical images as described in any one of claims 1 to 12.
16. A computer program product, characterized in that, The computer program product includes computer instructions. The computer instructions are executed by the processor of the computer device so that the computer device executes the image processing method for medical images as described in any one of claims 1 to 12.
Citation Information
Patent Citations
Neural network training and image processing method and device
CN112749801A