Image processing method and image processing model training method
By combining PET/CT image feature extraction and fusion segmentation image processing methods, lesion areas are automatically identified, solving the error and burden problems caused by relying on manual identification, and achieving more efficient and accurate lesion detection.
Patent Information
- Application Number
- CN202511120647.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2025-12-30
AI Technical Summary
In existing technologies, the identification of lesion areas in PET/CT images relies on the personal experience of clinicians, which introduces subjective errors and increases workload.
By using image processing methods and models, and combining feature extraction and fusion segmentation of PET and CT images, lesion areas can be automatically identified, and abnormal mask images can be generated, reducing manual intervention.
It improves the accuracy and efficiency of lesion identification, reduces the workload of clinicians, and decreases human error and variability.
Smart Images

Figure CN121237369A_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the field of image processing technology, and in particular to image processing methods and image processing model training methods. Background Technology
[0002] As an indispensable tool in tumor imaging, PET / CT (Positron emission tomography / computed tomography) plays a crucial role in tumor delineation, treatment planning, and prognostic assessment. PET / CT can simultaneously acquire tissue metabolic and anatomical information of the lesion site in the human body, thereby enabling the identification of the lesion area.
[0003] However, in current clinical practice, lesions in PET / CT images are generally identified through visual inspection. This method relies on the clinician's personal experience, is subject to subjective error, and increases the clinician's workload. Therefore, there is an urgent need for an effective technical solution to address these issues. Summary of the Invention
[0004] In view of the above, embodiments of this specification provide image processing methods. One or more embodiments of this specification also relate to image processing apparatus, image processing model training methods, image processing model training devices, CT image processing methods, computer-aided diagnosis methods for tumors, computer-aided diagnosis systems for tumors, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0005] According to a first aspect of the embodiments of this specification, an image processing method is provided, comprising:
[0006] Determine the first and second images corresponding to the target object;
[0007] The first image and the second image are input into the first image processing model to obtain the anomaly mask image corresponding to the target object;
[0008] The abnormal mask image is determined based on the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, the fusion segmentation result and fusion image features between the first image and the second image, and the image segmentation result corresponding to the preset image index. The image segmentation result is obtained by segmenting the second image based on the preset image index.
[0009] According to a second aspect of the embodiments of this specification, an image processing apparatus is provided, comprising:
[0010] The determination module is configured to determine the first and second images corresponding to the target object;
[0011] The input module is configured to input the first image and the second image into a first image processing model to obtain an anomaly mask image corresponding to the target object;
[0012] The abnormal mask image is determined based on the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, the fusion segmentation result and fusion image features between the first image and the second image, and the image segmentation result corresponding to the preset image index. The image segmentation result is obtained by segmenting the second image based on the preset image index.
[0013] According to a third aspect of the embodiments of this specification, another image processing method is provided, comprising:
[0014] Determine the first and second images corresponding to the target object;
[0015] The first image and the second image are input into the second image processing model to obtain the anomaly score corresponding to the target object;
[0016] The anomaly score is determined based on the region fusion features corresponding to at least one anomaly region of the target object. The region fusion features corresponding to at least one anomaly region are determined based on the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, the fusion segmentation result and fusion image features between the first image and the second image, and the target part mask image corresponding to the target object. The target part mask image is obtained by masking the first image.
[0017] According to a fourth aspect of the embodiments of this specification, another image processing apparatus is provided, comprising:
[0018] The determination module is configured to determine the first and second images corresponding to the target object;
[0019] The input module is configured to input the first image and the second image into a second image processing model to obtain the anomaly score corresponding to the target object;
[0020] The anomaly score is determined based on the region fusion features corresponding to at least one anomaly region of the target object. The region fusion features corresponding to at least one anomaly region are determined based on the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, the fusion segmentation result and fusion image features between the first image and the second image, and the target part mask image corresponding to the target object. The target part mask image is obtained by masking the first image.
[0021] According to a fifth aspect of the embodiments of this specification, an image processing model training method is provided, comprising:
[0022] Determine the first sample image, the second sample image, and the label anomaly mask image corresponding to the target object;
[0023] The first sample image and the second sample image are input into a first image processing model to obtain a predicted anomaly mask image corresponding to the target object. The predicted anomaly mask image is determined based on a first segmentation result and first image features corresponding to the first sample image, a second segmentation result and second image features corresponding to the second sample image, fused image features and fused segmentation results between the first sample image and the second sample image, and image segmentation results corresponding to preset image indicators. The image segmentation results are obtained by segmenting the second sample image based on the preset image indicators.
[0024] The first image processing model is trained based on the predicted anomaly mask image, the labeled anomaly mask image, and the reference part mask image corresponding to the target object until a trained first image processing model is obtained. The reference part mask image is obtained by masking the sample first image.
[0025] According to a sixth aspect of the embodiments of this specification, an image processing model training apparatus is provided, comprising:
[0026] The determination module is configured to determine the first sample image, the second sample image, and the label anomaly mask image corresponding to the target object;
[0027] The input module is configured to input the first sample image and the second sample image into a first image processing model to obtain a predicted anomaly mask image corresponding to the target object. The predicted anomaly mask image is determined based on a first segmentation result and first image features corresponding to the first sample image, a second segmentation result and second image features corresponding to the second sample image, fused image features and fused segmentation results between the first sample image and the second sample image, and an image segmentation result corresponding to a preset image index. The image segmentation result is obtained by segmenting the second sample image based on the preset image index.
[0028] The training module is configured to train the first image processing model based on the predicted anomaly mask image, the labeled anomaly mask image, and the reference part mask image corresponding to the target object, until a trained first image processing model is obtained, wherein the reference part mask image is obtained by masking the sample first image.
[0029] According to a seventh aspect of the embodiments of this specification, another image processing model training method is provided, comprising:
[0030] Determine the first sample image, the second sample image, and the label anomaly score corresponding to the target object;
[0031] The first sample image and the second sample image are input into a second image processing model to obtain a predicted anomaly score corresponding to the target object. The predicted anomaly score is determined based on the region fusion features corresponding to at least one anomaly region of the target object. The region fusion features corresponding to at least one anomaly region are determined based on the first segmentation result and first image features corresponding to the first sample image, the second segmentation result and second image features corresponding to the second sample image, the fused image features and fused segmentation result between the first sample image and the second sample image, and the target part mask image corresponding to the target object. The target part mask image is obtained by masking the first sample image.
[0032] The second image processing model is trained based on the predicted anomaly score and the labeled anomaly score until a fully trained second image processing model is obtained.
[0033] According to an eighth aspect of the embodiments of this specification, another image processing model training apparatus is provided, comprising:
[0034] The determination module is configured to determine the first sample image, the second sample image, and the label anomaly score corresponding to the target object;
[0035] The input module is configured to input the first sample image and the second sample image into a second image processing model to obtain a predicted anomaly score corresponding to the target object. The predicted anomaly score is determined based on the region fusion features corresponding to at least one anomaly region of the target object. The region fusion features corresponding to at least one anomaly region are determined based on the first segmentation result and first image features corresponding to the first sample image, the second segmentation result and second image features corresponding to the second sample image, the fused image features and fused segmentation result between the first sample image and the second sample image, and the target part mask image corresponding to the target object. The target part mask image is obtained by masking the first sample image.
[0036] The training module is configured to train the second image processing model based on the predicted anomaly score and the labeled anomaly score until a trained second image processing model is obtained.
[0037] According to a ninth aspect of the embodiments of this specification, a CT image processing method is provided, comprising:
[0038] Receive a CT image processing task, wherein the CT image processing task carries a CT image and a PET image corresponding to the target object, and the CT image processing task is used to detect abnormal information of the target object;
[0039] The CT image and the PET image are input into a first image processing model to obtain an anomaly mask image for the target object output by the first image processing model, wherein the first image processing model is trained according to the above-described image processing model training method; and / or
[0040] The CT image and the PET image are input into the second image processing model to obtain the anomaly score for the target object output by the second image processing model, wherein the second image processing model is trained according to the above-described image processing model training method.
[0041] According to a tenth aspect of the embodiments of this specification, an image processing method is provided, applied to a cloud-side device, comprising:
[0042] The receiving end device sends an image processing task, wherein the image processing task carries a first image and a second image corresponding to the target object, and the image processing task is used to detect abnormal information of the target object;
[0043] The first image and the second image are input into a first image processing model to obtain an anomaly mask image for the target object output by the first image processing model, wherein the first image processing model is trained according to the above-described image processing model training method; and / or
[0044] The first image and the second image are input into the second image processing model to obtain the anomaly score for the target object output by the second image processing model, wherein the second image processing model is trained according to the above image processing model training method;
[0045] Send the anomaly mask image and / or anomaly score to the end-side device.
[0046] According to the eleventh aspect of the embodiments of this specification, a computer-aided diagnosis method for tumors is provided, comprising:
[0047] Receive a tumor screening task, wherein the tumor screening task carries CT images and PET images corresponding to the target object, and the CT image processing task is used to detect tumor information of the target object;
[0048] The CT image and the PET image are input into a first image processing model to obtain an anomaly mask image for the target object output by the first image processing model, wherein the first image processing model is trained according to the above-described image processing model training method; and / or
[0049] The CT image and the PET image are input into the second image processing model to obtain the anomaly score for the target object output by the second image processing model, wherein the second image processing model is trained according to the above-described image processing model training method.
[0050] According to a twelfth aspect of the embodiments of this specification, a computer-aided diagnostic system for tumors is provided, comprising a client and a server, wherein,
[0051] The client is used to send a CT image processing task to the server, wherein the CT image processing task carries a CT image and a PET image corresponding to the target object, and the CT image processing task is used to detect abnormal information of the target object;
[0052] The server is configured to input the CT image and the PET image into a first image processing model to obtain an anomaly mask image for the target object output by the first image processing model, wherein the first image processing model is trained according to the aforementioned image processing model training method; and / or
[0053] The CT image and the PET image are input into the second image processing model to obtain the anomaly score for the target object output by the second image processing model, wherein the second image processing model is trained according to the above image processing model training method;
[0054] The server is also used to send the anomaly mask image and / or the anomaly score to the client.
[0055] According to a thirteenth aspect of the embodiments of this specification, a computing device is provided, comprising:
[0056] Memory and processor;
[0057] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions, which, when executed by the processor, implement the steps of the above method.
[0058] According to a fourteenth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0059] According to a fifteenth aspect of an embodiment of this specification, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the steps of the method described above.
[0060] This specification provides an image processing method according to one embodiment, comprising: determining a first image and a second image corresponding to a target object; inputting the first image and the second image into a first image processing model to obtain an anomaly mask image corresponding to the target object; wherein the anomaly mask image is determined based on a first segmentation result and a first image feature corresponding to the first image, a second segmentation result and a second image feature corresponding to the second image, a fusion segmentation result and fusion image features between the first image and the second image, and an image segmentation result corresponding to a preset image index, wherein the image segmentation result is obtained by segmenting the second image based on the preset image index.
[0061] In the above method, after determining the first and second images of the target object, the first and second images are input into a first image processing model. The first image processing model predicts abnormal information of the target object based on the first and second images, achieving intelligent recognition of the first and second images without requiring manual identification of abnormal regions by clinicians. Furthermore, the first image processing model can perform feature processing and segmentation on the first and second images separately, and then fuse these processes to obtain a first segmentation result and first image features corresponding to the first image, a second segmentation result and second image features corresponding to the second image, and a fused segmentation result and fused image features between the first and second images. Based on preset image metrics, the second image is segmented to obtain the image segmentation result corresponding to the preset image metrics, thereby obtaining an abnormal mask image of the target object. This further ensures accurate identification of abnormal regions of the target object, reduces the workload of clinicians, and improves the accuracy and efficiency of abnormal information recognition of the target object. Attached Figure Description
[0062] Figure 1 This is a schematic diagram illustrating an application scenario of an image processing method provided in one embodiment of this specification;
[0063] Figure 2 This is a flowchart illustrating an image processing method provided in one embodiment of this specification;
[0064] Figure 3 This is a flowchart of the training process of the first image processing model in an image processing method provided in one embodiment of this specification;
[0065] Figure 4 This is a flowchart of the training process of the second image processing model in an image processing method provided in one embodiment of this specification;
[0066] Figure 5 This is a schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification;
[0067] Figure 6 This is a flowchart of another image processing method provided in one embodiment of this specification;
[0068] Figure 7 This is a schematic diagram of the structure of another image processing apparatus provided in one embodiment of this specification;
[0069] Figure 8 This is a flowchart illustrating an image processing model training method provided in one embodiment of this specification;
[0070] Figure 9This is a schematic diagram of the structure of an image processing model training device provided in one embodiment of this specification;
[0071] Figure 10 This is a flowchart of another image processing model training method provided in one embodiment of this specification;
[0072] Figure 11 This is a schematic diagram of the structure of another image processing model training device provided in one embodiment of this specification;
[0073] Figure 12 This is a flowchart of a CT image processing method provided in one embodiment of this specification;
[0074] Figure 13 This is a flowchart of another image processing method provided in one embodiment of this specification;
[0075] Figure 14 This is a flowchart illustrating a computer-aided diagnostic method for tumors provided in one embodiment of this specification;
[0076] Figure 15 This is a schematic diagram of the structure of a computer-aided diagnosis system for tumors provided in one embodiment of this specification;
[0077] Figure 16 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0078] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0079] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0080] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0081] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0082] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of model parameters. A large model can also be called a foundation model. It is pre-trained using large-scale unlabeled corpora to produce a pre-trained model with hundreds of millions of parameters. Such models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and multi-modal pre-training models.
[0083] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as Natural Language Processing (NLP) and Computer Vision. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as NLP tasks such as text-based sentiment classification, text summarization, and machine translation. The main application scenarios for large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0084] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0085] CT stands for Computed Tomography, a technique that uses X-rays to take images of the body's interior from multiple angles, generating detailed cross-sectional views of the body.
[0086] PET: Positron Emission Tomography, a nuclear medicine imaging technique that creates three-dimensional images of organs and tissues inside the body by detecting gamma rays emitted by radioactive materials (usually radioactive tracers injected into the body).
[0087] Softmax: An activation function that transforms a K-dimensional vector containing arbitrary real numbers into another K-dimensional real vector, where each element of the output vector lies in the interval (0,1), and the sum of all elements is 1. It is commonly used in multi-class classification problems as the output layer activation function to represent the probability distribution of belonging to each class.
[0088] SUV: Standardized Uptake Value, is an important parameter in PET imaging used to quantify the extent of tracer uptake in the body. It helps assess tumor metabolic activity by comparing the ratio of the lesion area to the average active concentration throughout the body.
[0089] DLBCL: Diffuse Large B-Cell Lymphoma, is one of the most common types of non-Hodgkin lymphoma, characterized by abnormal proliferation of B cells and the formation of large tumor masses.
[0090] nnUNet: No New-Net, is a deep learning model based on the U-Net architecture, specifically designed for medical image segmentation tasks. nnUNet emphasizes the importance of data preprocessing, network architecture selection, and post-processing steps, aiming to automate and adapt to different medical image segmentation challenges.
[0091] U-Net is a convolutional neural network architecture specifically designed for biomedical image segmentation. The U-Net architecture consists of a shrinking path (downsampling path) and an expanding path (upsampling path). The shrinking path captures contextual information in the image through a series of convolutional and pooling layers, while the expanding path uses upsampling operations to recover the precise location of objects.
[0092] IPI: International Prognostic Index, is mainly used to predict the survival rate of patients with aggressive non-Hodgkin lymphoma. Factors considered include age, disease stage, serum lactate dehydrogenase level, physical status, and number of extranodal lesions.
[0093] mask: In medical image analysis, a mask can be used to identify specific anatomical structures or areas of pathological changes.
[0094] MRT: Magnetic Resonance Tomography, also known as MRI in some countries, is a technique that uses strong magnetic fields and radio frequency waves to image the human body. It can provide very detailed soft tissue images and is widely used in the diagnosis of a variety of diseases.
[0095] Hybrid expert model: This is a machine learning technique that uses a gating model to divide a single task space into multiple sub-tasks, and then multiple expert networks (sub-models) handle specific sub-tasks respectively.
[0096] This specification provides an image processing method, and also relates to an image processing apparatus, an image processing model training method, an image processing model training apparatus, a CT image processing method, a computer-aided diagnosis method for tumors, a computer-aided diagnosis system for tumors, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0097] See Figure 1 , Figure 1 The illustration shows an application scenario of an image processing method according to an embodiment of this specification, which includes the following steps.
[0098] Determine the first and second images corresponding to the target object.
[0099] The first image and the second image are input into the first image processing model to obtain the anomaly mask image corresponding to the target object.
[0100] The abnormal mask image is determined based on the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, the fusion segmentation result and fusion image features between the first image and the second image, and the image segmentation result corresponding to the preset image index. The image segmentation result is obtained by segmenting the second image based on the preset image index.
[0101] Figure 1 It includes end-side device 102 and cloud-side device 104.
[0102] In practical applications, in the field of assisted medical care, users can send a first image and a second image corresponding to a target object to a cloud-based device 104 via the edge device 102. The cloud-based device 104 can input the first image and the second image into a first image processing model, and perform feature extraction and image segmentation on the first image and the second image based on the first image processing model to finally obtain an anomaly mask image corresponding to the target object. The cloud-based device 104 can send the anomaly mask image to the edge device 102 and display it to the user through the display interface of the edge device 102, thereby realizing automated analysis of medical images of the target object and achieving lesion segmentation and tumor identification.
[0103] The edge device 102 may include a browser, an app (application), or a web application such as an H5 (Hypertext Markup Language 5) application, a lightweight application (also known as a mini-program), or a cloud application. The edge device can be developed based on a software development kit (SDK) provided by the server, such as a real-time communication (RTC) SDK. The edge device can be deployed in an electronic device and depends on the device's operation or certain apps within the device to run. The electronic device may have a display screen and support information browsing, such as a personal mobile terminal like a mobile phone, tablet, or personal computer. Various other types of applications can also be configured in the electronic device, such as human-computer interaction applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email clients, and social media platform software.
[0104] Cloud-side device 104 can be understood as a server providing various services, including physical servers and cloud servers. Examples include servers providing communication services to multiple clients, servers supporting backend training of models used on clients, and servers processing data sent by clients. It should be noted that cloud-side device 104 can be implemented as a distributed server cluster composed of multiple servers, or as a single server. Cloud-side device 104 can also be a server for a distributed system, or a server integrated with blockchain. Cloud-side device 104 can also be a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology.
[0105] It is worth noting that the image processing method provided in the embodiments of this specification can be executed by the cloud-side device 104. In other embodiments of this specification, the first image processing model can be deployed in the edge device 102, so that the edge device 102 can also have similar functions to the cloud-side device 104, thereby executing the image processing method provided in the embodiments of this specification. In other embodiments, the image processing method provided in the embodiments of this specification can also be jointly executed by the edge device 102 and the cloud-side device 104.
[0106] See Figure 2 , Figure 2 A flowchart of an image processing method according to an embodiment of this specification is shown, which specifically includes the following steps.
[0107] Step 202: Determine the first image and the second image corresponding to the target object.
[0108] Specifically, the image processing method provided in the embodiments of this specification can be applied to the field of auxiliary medical care. Specifically, it can be used to determine whether there are abnormalities in the target object, and further, it can be used to determine whether there are abnormalities in the target detection area of the target object.
[0109] In the field of assistive medical imaging, the target object can be understood as the person to be examined, and the target detection area can be understood as the organ or part to be examined, such as the liver, lungs, or stomach. The target detection area can also be understood as the entire body of the target object. Therefore, this image processing method can be used to determine whether lesions exist in the target detection area, specifically to determine whether tumors or other lesions exist in the target object's organs. The first image can be understood as a CT image obtained from a medical scan of the target object, and the second image can be understood as a PET image obtained from a medical scan of the target object.
[0110] In practical applications, PET / CT plays a crucial role in tumor delineation, treatment planning, and prognostic assessment. The complementarity of PET and CT images is a primary reason for its widespread adoption. Typically, PET images provide sensitive information about metabolic activity by identifying areas of increased glucose consumption, often indicating abnormal activity; while CT images provide anatomical specificity, helping to distinguish malignant tumors from normal organs that may also show high glucose uptake. For example, normal organs such as the bladder, muscles, intestines, and bones may appear as high intensities on PET images, making it difficult to distinguish lesions from these organs using PET images alone. Furthermore, due to tumor type or size, some lesions may have low glucose uptake, further complicating lesion detection and leading to the misidentification of normal metabolic areas as lesions or the overlooking of true lesions with subtle intensities, largely due to a lack of prior clinical knowledge of metabolic activity and anatomical structures. By integrating these two imaging techniques, PET / CT provides a comprehensive view, enhancing the accuracy of tumor assessment. Segmented lesion regions obtained from PET / CT contain both the metabolic intensity dominated by PET images and the structural details dominated by CT images, providing a comprehensive characterization of the lesion. However, manually delineating lesion volumes from PET / CT images is highly labor-intensive and exhibits significant inter-observer variability, posing a challenge to its consistent application in routine clinical practice. Therefore, this paper proposes a first image processing model that intelligently processes PET and CT images, thereby improving the efficiency and reproducibility of tumor burden assessment and reducing human error and variability.
[0111] In practical applications, the image processing method provided in the embodiments of this specification can be applied to a server. The server can then receive a first image and a second image of the target detection region of the target object sent by the client. Furthermore, the server can receive multiple first images and multiple second images of the target detection region sent by the client. Subsequently, similar processing can be performed on each first image and each second image to further achieve anomaly detection of the target detection region.
[0112] Step 204: Input the first image and the second image into the first image processing model to obtain the anomaly mask image corresponding to the target object;
[0113] The abnormal mask image is determined based on the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, the fusion segmentation result and fusion image features between the first image and the second image, and the image segmentation result corresponding to the preset image index. The image segmentation result is obtained by segmenting the second image based on the preset image index.
[0114] The first image processing model can be used to predict abnormal information of a target object. The abnormality mask image can be understood as the abnormal information of the target object predicted by the first image processing model. The abnormality mask image can include the abnormality mask information and the location mask information of the target object. In the field of assisted medical care, the abnormality mask image can be used to represent the abnormal location information of the target detection area of the target object, such as the location of lesions and tumors in the target detection area of the target object. The abnormality mask information can be understood as the lesion mask information of the lesion area of the target object, and the location mask information can be understood as the mask information of the organ of the target object.
[0115] Furthermore, after inputting the first image and the second image into the first image processing model, the first image processing model can also output target detection results for the target object. The target detection results can be understood as the detection results of whether there are abnormalities in the target detection area of the target object and the classification of the existing abnormalities. The target detection results can include abnormal location information in the target detection area and / or information on whether there are abnormalities in the target detection area and the classification of abnormalities (such as the number of abnormal classifications). For example, when the target detection area is the liver area, the target detection results can be used to determine whether the liver area has lesions and the classification of the lesions. The lesion classification can include benign lesions and malignant lesions.
[0116] In practical applications, preset image metrics can be understood as quantitative indicators for PET images in the medical field. For example, this could be the SUV threshold, which is a standardized uptake value used to measure the degree of uptake of radiolabeled glucose analogs at specific sites in the body. Since tumor cells typically have higher metabolic activity than normal cells, they take up more glucose analogs, thus displaying a higher SUV value on PET images. Based on this, PET images can be segmented using this SUV threshold to obtain the corresponding SUV threshold map.
[0117] In specific implementation, the first image processing model includes a first feature extraction layer, a first segmentation layer, and a first feature fusion layer;
[0118] The step of inputting the first image and the second image into the first image processing model to obtain the anomaly mask image corresponding to the target object includes:
[0119] The first image and the second image are input into the first image processing model. The first feature extraction layer is used to extract and fuse features of the first image and the second image to obtain the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, and the fused segmentation result and fused image features between the first image and the second image.
[0120] Using the first segmentation layer, the second image is segmented based on the preset image index to obtain the image segmentation result;
[0121] Using the first feature fusion layer, the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, the fused segmentation result and fused image features between the first image and the second image, and the image segmentation result are fused to obtain the abnormal mask image corresponding to the target object.
[0122] Specifically, after inputting the CT image and PET image into the first image processing model, the CT image and PET image are input into the first feature extraction layer to obtain the first segmentation result and first image feature corresponding to the CT image, the second segmentation result and second image feature corresponding to the PET image, and the fused image feature and fused segmentation result between the CT image and PET image. The PET image is then input into the first segmentation layer to obtain the image segmentation result obtained by segmenting the PET image based on preset image indicators. The first segmentation result and first image feature, the second segmentation result and second image feature, the fused segmentation result and fused image feature, and the image segmentation result are then input into the first feature fusion layer to obtain the anomaly mask image corresponding to the target object output by the first feature fusion layer.
[0123] In practical applications, the first segmentation layer may include an SUV thresholding layer and a convolutional block. During the processing of the second image based on preset image metrics in the first segmentation layer, the second image can be processed using an empirical thresholding method to obtain the SUV threshold map corresponding to the second image. The convolutional block is then used to perform feature processing on the SUV threshold map to obtain the image features corresponding to the SUV threshold map, which are then used as the image segmentation result. The empirical thresholding method can be understood as a method for determining one or more thresholds based on experience or experimental data in data analysis, image processing, and other fields that require distinguishing different data categories or features. These thresholds are used to convert continuous data into discrete categories or to separate objects or signals of interest from the background. In the field of assisted medical care, SUV thresholds can be set based on the past experience of medical personnel and a large amount of case data, serving as a standard for judging whether a lesion may be malignant. The convolutional block can be used to extract image features from the SUV threshold map. This convolutional block may include convolutional layers, activation function layers, batch normalization layers, and pooling layers; however, this specification does not limit the specific implementation of these components.
[0124] In summary, by using the first image processing model to automatically process CT and PET images, anomaly mask images of the target object are obtained. This combines the anatomical structure information provided by CT images and the metabolic activity information provided by PET images, ensuring the accuracy of the anomaly mask images and improving the efficiency and accuracy of lesion detection.
[0125] In practical applications, the first feature extraction layer includes a first encoder, a second encoder, a first decoder, a second decoder, and a fusion decoder;
[0126] The step of using the first feature extraction layer to extract and fuse features in the first image and the second image to obtain a first segmentation result and first image features corresponding to the first image, a second segmentation result and second image features corresponding to the second image, and a fused segmentation result and fused image features between the first image and the second image includes:
[0127] The first image is input into the first encoder to obtain the first encoded feature corresponding to the first image;
[0128] The second image is input into the second encoder to obtain the second encoded feature corresponding to the second image;
[0129] The first encoded feature is input into the first decoder to obtain the first segmentation result and the first image feature corresponding to the first image;
[0130] The second encoded feature is input into the second decoder to obtain the second segmentation result and the second image feature corresponding to the second image;
[0131] The first encoded feature and the second encoded feature are input into the fusion decoder to obtain the fusion segmentation result between the first image and the second image, as well as the fusion image features.
[0132] Wherein, the first encoder can be a CT encoder, the second encoder can be a PET encoder, the first decoder can be a CT decoder, the second decoder can be a PET decoder, the fusion decoder can be a PET / CT decoder, the first segmentation result can be a CT-based segmented image, the second segmentation result can be a PET-based segmented image, the first coded feature can be a coded feature obtained by multi-scale and multi-level feature extraction of CT images, the first image feature can be a specific CT image feature, the second coded feature can be a coded feature obtained by multi-scale and multi-level feature extraction of PET images, and the second image feature can be a specific PET image feature.
[0133] Taking a CT encoder as an example, a CT encoder can include multiple CT coding layers. These multiple CT coding layers can be used to extract CT coding features at different levels or scales. Similarly, a CT decoder can also include multiple CT decoding layers. The CT coding features at different levels or scales output by the multiple CT coding layers are input into the corresponding CT decoding layer to obtain the first image features. Correspondingly, PET encoders, PET decoders, and PET / CT decoders all have similar structures for multi-scale, multi-level feature encoding and decoding. Multi-scale, multi-level coding features can include fine-grained anatomical details, as well as the overall layout and interrelationships of an entire organ or lesion area. Specific CT image features can be understood as imaging features captured by CT images that are closely related to a specific disease or pathological state, used to improve diagnostic accuracy, predict disease progression, or assess treatment response. Specifically, a CT decoder can perform quantitative analysis, texture analysis, morphological feature analysis, and functional information analysis on the first coding features to obtain specific CT image features. Correspondingly, a PET decoder can also perform similar processing on the second coding features; however, this specification does not limit this aspect in the embodiments.
[0134] Specifically, CT images can be input into a CT encoder to obtain the first coded features corresponding to the CT image, and PET images can be input into a PET encoder to obtain the second coded features corresponding to the PET image. The first coded features are then input into a CT decoder to obtain the first segmentation result and first image features corresponding to the CT image, and the second coded features are input into a PET decoder to obtain the second segmentation result and second image features corresponding to the PET image. The first and second coded features are then concatenated to obtain concatenated coded features, which are then input into a PET / CT decoder to obtain fused image features and fused segmentation results.
[0135] In practical applications, the first feature extraction layer can be a multi-branch feature fusion network used to extract features from the PET and CT modalities respectively. A shared-parameter encoder (i.e., a CT encoder and a PET encoder) is designed, taking PET and CT images as inputs to obtain multi-level encoded features (i.e., first and second encoded features). Next are three decoders: a CT decoder, a PET decoder, and a PET / CT decoder. The CT decoder uses the encoded features output from the CT encoder as input to obtain specific CT image features and CT-based segmentation results. The PET decoder uses the encoded features output from the PET encoder as input to obtain specific PET image features and PET-based segmentation results. The PET / CT decoder concatenates the encoded features output from the CT encoder and the PET encoder as input to obtain fused image features from both modalities and the fused segmentation results.
[0136] Optionally, parameters are shared between the first encoder and the second encoder. That is, the first encoder and the second encoder use the same weights and parameters, which can reduce the number of model parameters of the first image processing model. When different inputs pass through encoders that share the same parameters, the model can learn more generalized feature representations instead of customizing feature extractors for each input. This helps to improve the model's performance when faced with unseen data, enhances the model's generalization ability, and promotes knowledge transfer between multiple tasks and multiple modalities, thereby enhancing the model's robustness.
[0137] In summary, by extracting and fusing features from PET and CT images, clinical knowledge of anatomy and metabolism is effectively integrated to improve the accuracy of PET / CT lesion segmentation.
[0138] Further, the step of using the first feature fusion layer to fuse the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, the fused segmentation result and fused image features between the first image and the second image, and the image segmentation result to obtain the anomaly mask image corresponding to the target object includes:
[0139] Using the first feature fusion layer, the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, the fused segmentation result and fused image features between the first image and the second image, and the image segmentation result are fused to obtain the abnormal segmentation result corresponding to the target object and the probability corresponding to the abnormal segmentation result.
[0140] Based on the anomaly segmentation results and the probabilities corresponding to the anomaly segmentation results, the anomaly mask image corresponding to the target object is determined.
[0141] Among them, the abnormal segmentation result can be understood as the segmentation result that predicts whether the target object has a lesion and the location of the lesion. The probability corresponding to the abnormal segmentation result can be understood as the probability of the existence of a lesion.
[0142] Specifically, the first image features and the first segmentation result, the second image features and the second segmentation result, the fused image features and the fused segmentation result, and the image segmentation result can be input into the first feature fusion layer. In the first feature fusion layer, the abnormal segmentation result corresponding to the target object and the probability corresponding to the abnormal segmentation result are obtained. Based on the abnormal segmentation result and the probability corresponding to the first segmentation result, the abnormal mask image corresponding to the target object is generated and output.
[0143] In practical applications, in order to integrate the prediction results of multi-branch encoders and clinical prior knowledge (SUV threshold map), an interpretable fusion module, namely the first feature fusion layer, was designed based on the idea of hybrid expert model. The first feature fusion layer takes the image features and segmentation results output by each decoder and the SUV threshold map as input, and achieves the purpose of interpretable fusion by assigning voxel-level prediction probabilities to each segmentation result.
[0144] The first feature fusion layer includes a log-unit extraction block, a probability extraction block, and a softmax activation function. Specifically, logits can be understood as the output of the last layer (e.g., a fully connected layer) in the first feature fusion layer, the state before applying the softmax activation function. This can be an unnormalized score or predicted value. The log-unit extraction block can be used to obtain the untransformed, raw output of the first feature fusion layer. After obtaining the logits, the activation function can be used to transform them into a probability distribution. This means that for each input sample, a set of values will be generated, indicating the probability that the sample belongs to each category. The probability extraction block can be used to perform this transformation, mapping the logits to a probability space, making the results easier to interpret and understand. The softmax function is a commonly used activation function in multi-class classification problems, used to transform a K-dimensional vector containing arbitrary real numbers into another K-dimensional vector. The output of the softmax function can be interpreted as a probability distribution, representing the probability that the input belongs to each category.
[0145] Based on this, in the first feature fusion layer, the first image features and the first segmentation result, the second image features and the second segmentation result, the fused image features and the fused segmentation result, and the image segmentation result can be input into the first feature fusion layer. In the first feature fusion layer, logarithmic probability maps (i.e., logits maps) for the CT modality and the PET modality are obtained through the aforementioned logarithmic unit extraction block, and probability maps (i.e., probability maps) for the CT modality and the PET modality are obtained through the aforementioned probability extraction block. The logarithmic probability maps and the probability maps are fused and an activation function is used to obtain the anomaly mask image corresponding to the target object. Logits are the raw numerical values output by the last layer of the neural network. Logits Maps represent the raw predicted values (logits) output by the model for the inputs of the PET and CT modalities, respectively, and are usually used for classification tasks (such as segmentation or classification) for each pixel / voxel. The logits are converted into probability values corresponding to each category through the softmax activation function or other normalization methods. Probability maps represent the probability distribution of each category output by the model for the inputs of the CT and PET modalities on each pixel / voxel. Furthermore, the input to the first image processing model may also include images of other modalities. Modality can refer to data acquired by different imaging techniques, such as MRT images and ultrasound images corresponding to the target object. This specification does not limit this aspect in the embodiments.
[0146] In summary, by considering the characteristics of different modalities in PET and CT, and using the SUV threshold as clinical prior knowledge, SUV mapping can consider more candidate tumor regions, achieving rapid convergence and recall of more small tumors. This first feature fusion layer, as an interpretative fusion module, improves segmentation accuracy and clearly identifies the contribution of each modality to the final segmentation of each pixel, clarifying the different contributions of metabolic and anatomical information to model decision-making.
[0147] In practical applications, before inputting the first image and the second image into the first image processing model, the following steps are also included:
[0148] Determine the first sample image, the second sample image, and the label anomaly mask image corresponding to the target object;
[0149] The first sample image and the second sample image are input into a first image processing model to obtain a predicted anomaly mask image corresponding to the target object. The predicted anomaly mask image is determined based on a first segmentation result and first image features corresponding to the first sample image, a second segmentation result and second image features corresponding to the second sample image, fused image features and fused segmentation results between the first sample image and the second sample image, and image segmentation results corresponding to preset image indicators. The image segmentation results are obtained by segmenting the second sample image based on the preset image indicators.
[0150] The first image processing model is trained based on the predicted anomaly mask image, the labeled anomaly mask image, and the reference part mask image corresponding to the target object until a trained first image processing model is obtained. The reference part mask image is obtained by masking the sample first image.
[0151] Among them, the first sample image, the second sample image, and the label anomaly mask image can be used as training pairs for model training. The first sample image can be understood as a sample CT image, and the second sample image can be understood as a sample PET image.
[0152] Specifically, sample CT images, sample PET images, and label anomaly mask images corresponding to the target object can be identified. The sample CT and sample PET images are then input into a first image processing model to obtain a predicted anomaly mask image output by the target object. Based on this predicted anomaly mask image, the label anomaly mask image, and the reference part mask image corresponding to the target object, the first image processing model is trained until a fully trained first image processing model is obtained. Subsequently, during actual lesion detection, anomaly mask image prediction can be performed based on this trained first image processing model.
[0153] It is understandable that the process of obtaining the predicted anomaly mask image corresponding to the target object by processing the first sample image and the second sample image based on the first image processing model is similar to the process of obtaining the anomaly mask image corresponding to the target object by processing the first image and the second image based on the first image processing model in the above application process. The embodiments in this specification will not be described again here.
[0154] In summary, by training the first image processing model through supervised learning, the first image processing model can be adapted to PET / CT tumor segmentation tasks, which facilitates the automation of tumor identification.
[0155] In specific implementation, the step of training the first image processing model based on the predicted anomaly mask image, the labeled anomaly mask image, and the reference part mask image corresponding to the target object until a trained first image processing model is obtained includes:
[0156] The first model loss value is calculated based on the anomaly mask information contained in the labeled anomaly mask image and the predicted anomaly mask image.
[0157] The second model loss value is calculated based on the part mask information contained in the reference part mask image and the predicted anomaly mask image.
[0158] The first image processing model is trained based on the first model loss value and the second model loss value until a fully trained first image processing model is obtained.
[0159] The location mask information can be understood as the mask information of the organ part of the target object. The label anomaly mask image can be the mask image of the actual lesion of the target object.
[0160] Specifically, a mask image containing anomaly mask information and a mask image containing part mask information can be determined from the predicted anomaly mask image. Based on the mask image containing anomaly mask information and the label anomaly mask image, a first model loss value is calculated. Based on the mask image containing part mask information and the reference part mask image, a second model loss value is calculated. The first model loss value and the second model loss value are added or weighted to obtain the target model loss value. The first image processing model is trained based on the target model loss value until a first image processing model that meets the training stopping condition is obtained as the first image processing model that has been trained. The training stopping condition may be that the model loss value reaches a preset loss value threshold and / or the number of model training times reaches a preset number threshold.
[0161] In summary, by introducing the reference site mask image as prior clinical information during model training, the lesion segmentation performance of the first image processing model can be further improved.
[0162] In specific implementation, before training the first image processing model based on the predicted anomaly mask image, the labeled anomaly mask image, and the reference part mask image corresponding to the target object, until the trained first image processing model is obtained, the method further includes:
[0163] Based on the reference region corresponding to the target object, the first image of the sample is segmented to obtain a reference region mask image corresponding to the reference region.
[0164] The reference site corresponding to the target object can be understood as the organ part related to the abnormal information among the organs included in the target object. For example, in practical applications, the reference site can be the organ part with high physiological uptake among the organs included in the target object, as well as the organ with high false positive lesion prediction, that is, the organ part that is prone to error in lesion prediction. For example, for tumor recognition based on PET images, the bladder, heart and kidney, three key organs of the target object that are prone to misidentification, can be used as the reference site corresponding to the target object.
[0165] Specifically, based on a deep learning model, CT images can be segmented to obtain a full-body mask image of the target object. Then, based on the corresponding reference body part of the target object, a reference body part mask image can be determined from the full-body mask image. This full-body mask image can include organ mask information for all organs of the target object. The deep learning model can be understood as a deep learning model designed for medical image segmentation, which can be used to automatically segment various anatomical structures from CT images.
[0166] In summary, by introducing a multi-organ anatomical mask (i.e., a reference site mask image) as an additional supervisory signal to integrate anatomical knowledge during the training of the first image processing model, and by using error-prone physiological uptake organs (i.e., reference sites) as supervisory signals to enhance the model's ability to distinguish between physiological and pathological uptake, the first image processing model can learn the knowledge of segmenting key erroneous organs. Introducing this prior clinical information enables the first image processing model to have a basic understanding of human anatomy, facilitating accurate lesion segmentation based on the first image processing model in the future.
[0167] Furthermore, after determining the first and second images corresponding to the target object, the method further includes:
[0168] The first image and the second image are input into the second image processing model to obtain the anomaly score corresponding to the target object;
[0169] The anomaly score is determined based on the region fusion features corresponding to at least one anomaly region of the target object. The region fusion features corresponding to at least one anomaly region are determined based on the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, the fusion segmentation result and fusion image features between the first image and the second image, and the target part mask image corresponding to the target object. The target part mask image is obtained by masking the first image.
[0170] In the field of assistive medicine, abnormal scores can be understood as prognostic assessment results for a target individual. In practical applications, lymphoma is a malignant tumor affecting the lymphatic system. Unlike solid tumors, lymphoma typically spreads throughout the body. Diffuse large B-cell lymphoma (DLBCL) is the most common type of non-Hodgkin lymphoma, accounting for 30% of all non-Hodgkin lymphomas. After initial treatment, approximately 60% to 80% of patients achieve complete remission. However, 20% to 40% of patients may experience relapse or disease progression after initial treatment. Therefore, early identification of patients with poor prognoses is crucial, allowing clinicians to avoid potentially ineffective treatments and adjust treatment strategies in a timely manner.
[0171] Currently, the main prognostic markers for DLBCL are the International Prognostic Index (IPI) and its variants. However, the IPI is primarily based on clinical factors and does not adequately consider the heterogeneity of tumors among individuals. As an important clinical imaging tool for DLBCL prognosis, PET / CT provides a non-invasive method to assess tumor metabolism and capture tumor heterogeneity from a macroscopic perspective. Prognostic biomarkers derived from PET / CT, such as metabolic tumor volume and total lesion glycolysis, have shown promising results. However, these metabolic parameters mainly rely on simple features, such as standard uptake values (SUV) and lesion volume, and fail to adequately capture the complexity of tumor features. To effectively address the multiple challenges of PET / CT in DLBCL prognostic assessment, including the heterogeneity of lesion number and location, insufficient lesion regional characterization, and lack of modeling of lesion anatomical background, one embodiment of this specification proposes a second image processing model that utilizes the anatomical background of lesions and the aggregation process of multiple lesions for lymphoma prognostic analysis. By effectively extracting and integrating features of multiple lesions while considering the anatomical situation of the lesions, the aim is to improve the accuracy of prognosis.
[0172] Understandably, the CT and PET images input into the second image processing model, and the CT and PET images input into the first image processing model mentioned above, can be CT and PET images of the target object at different treatment stages. For example, CT and PET images of the target object during the diagnosis and treatment stages can be input into the first image processing model for lesion segmentation, and CT and PET images of the target object during the prognostic stage can be input into the second image processing model for prognostic assessment.
[0173] Furthermore, the second image processing model includes the first feature extraction layer, the first feature fusion layer, the second segmentation layer, the second feature fusion layer, and the score prediction layer;
[0174] The step of inputting the first image and the second image into the second image processing model to obtain the anomaly score corresponding to the target object includes:
[0175] The first image and the second image are input into the second image processing model. The first feature extraction layer is used to extract and fuse features of the first image and the second image to obtain the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, and the fused segmentation result and fused image features between the first image and the second image.
[0176] Using the first feature fusion layer, the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, and the fused segmentation result and fused image features between the first image and the second image are fused to obtain the reference anomaly mask image corresponding to the target object.
[0177] Using the second segmentation layer, the first image is segmented according to the target part corresponding to the target object to obtain the target part mask image corresponding to the target part;
[0178] Using the second feature fusion layer, feature fusion processing is performed on the first segmentation result and the first image features, the second segmentation result and the second image features, the fused segmentation result and the fused image features, the reference anomaly mask image and the target part mask image to obtain the region fusion features corresponding to at least one anomaly region of the target object;
[0179] Using the score prediction layer, score prediction is performed on the region fusion features corresponding to the at least one abnormal region to obtain the abnormal score corresponding to the target object.
[0180] The first feature extraction layer and the first feature fusion layer have similar structures to the first feature extraction layer and the first feature fusion layer in the aforementioned first image processing model, and will not be repeated here in the embodiments of this specification.
[0181] In practical applications, the second image processing model can comprehensively utilize multi-lesion region aggregation and lesion background fusion. By effectively extracting and integrating features from multiple lesions while considering the anatomical structure of the lesions, it improves the accuracy of prognosis. For PET / CT-based lymphoma prognosis prediction, this second image processing model implements a three-stage computational process. First, a multimodal segmentation network (i.e., the first feature extraction layer and the first feature fusion layer) is used to achieve voxel-level lesion segmentation, while generating high-dimensional feature maps (first image features, second image features, and fused image features) that maintain spatial correspondence with the original imaging data (first image and second image). Subsequently, whole-body anatomical segmentation images are obtained from the CT images to generate organ mask images (i.e., target site mask images), and lesion anatomical background fusion is achieved through an attention-based feature interaction mechanism. Finally, multi-lesion region aggregation is used to integrate the features of distributed lesion regions, combining local lesion features and contextual organ information to perform accurate prognostic stratification.
[0182] In practical applications, for the first feature extraction layer and the first feature fusion layer, a multimodal segmentation network integrating PET and CT images can be designed based on the nnUNet framework. To accurately distinguish between lesions and normal tissue in PET / CT, this multimodal segmentation network can comprehensively encode the local appearance and global anatomical information of each lesion region. Therefore, the feature maps output by the multimodal segmentation network should be valuable for downstream tasks (such as prognosis). Based on this assumption, feature maps with the same spatial dimension as the input image can be extracted from the last block of each decoder branch in the aforementioned multimodal segmentation network. Then, lesion features (i.e., regional features corresponding to abnormal regions) can be aggregated from these features. For example, first image features with the same spatial dimension as the input CT image can be extracted from the CT decoder, and second image features with the same spatial dimension as the input PET image can be extracted from the PET decoder.
[0183] Furthermore, DLBCL can occur not only in lymph nodes but also in extranodal organs. Extranodal involvement is often associated with poor prognosis in DLBCL. Therefore, a lesion anatomical context fusion module (i.e., the second segmentation layer and the second feature fusion layer) can be used to characterize the relationship between prognosis and extranodal organ involvement. For the second segmentation layer, the target site can be an organ part related to the abnormal information within the target object, or it can be any organ within the target object, such as an organ where lymphoma may occur. Based on this, the second segmentation layer can be used to segment the CT image corresponding to the target object based on the target site, obtaining a target site mask image.
[0184] In practical applications, the second segmentation layer can be a deep learning model. This model can be used to segment CT images, obtaining a full-body mask image of the target object, which can then be used as a mask image for the target body part. Alternatively, the same deep learning model can be used to segment CT images, obtaining a full-body mask image of the target object. Based on the target body part corresponding to the target object, a target body part mask image can be determined from the full-body mask image. This full-body mask image can include organ mask information for all organs of the target object. The deep learning model can be understood as a deep learning model designed for medical image segmentation, capable of automatically segmenting various anatomical structures from CT images.
[0185] In summary, multimodal fusion networks can be used to extract specific features from different modalities of PET and CT images to enrich the characterization of various lesion areas, and by incorporating relevant information from extranodal organs, more accurate prognostic assessment can be achieved.
[0186] Further, the step of using the second feature fusion layer to perform feature fusion processing on the first segmentation result and the first image features, the second segmentation result and the second image features, the fused segmentation result and the fused image features, the reference anomaly mask image, and the target part mask image to obtain region fusion features corresponding to at least one anomaly region of the target object includes:
[0187] Using the second feature fusion layer, the first segmentation result and the first image features, the second segmentation result and the second image features, the fused segmentation result and the fused image features, the reference anomaly mask image and the target part mask image are fused to obtain the region features corresponding to at least one anomaly region of the target object and the part features corresponding to at least one target part.
[0188] An attention mechanism is applied to the region features corresponding to the at least one abnormal region and the location features corresponding to the at least one target location to obtain the region fusion features corresponding to at least one abnormal region of the target object.
[0189] Among them, the regional features corresponding to the abnormal area can be understood as the lesion area features, and the site features corresponding to the target site can be understood as the organ area features of the target site.
[0190] Specifically, in the second feature fusion layer, the first segmentation result, first image features, second segmentation result, second image features, fused segmentation result, fused image features, reference anomaly mask image, and target site mask image can be fused to obtain regional features corresponding to each abnormal region in at least one abnormal region corresponding to the target object, and site features corresponding to each target site in at least one target site. This enables similarity analysis between lesion region features and organ region features. Furthermore, attention mechanisms are applied to the regional features corresponding to each abnormal region and the site features corresponding to each target site to obtain regional fusion features corresponding to at least one abnormal region, thereby enriching the features and representation of each abnormal region.
[0191] In practical applications, extranodal involvement is characterized by assessing the similarity between lesion region features and organ region features. Subsequently, these organ region features and lesion region features are weighted and aggregated to enrich the features and representation of each lesion region.
[0192] In specific implementation, when processing the regional features (i.e., lesion region features) corresponding to at least one abnormal region and the site features (i.e., organ region features) corresponding to at least one target site using the attention mechanism, the lesion region features are transformed by the query weight matrix to obtain the query vector Q. The organ region features are transformed by the key weight matrix and the value weight matrix to obtain the key vector K and the value vector V, respectively. Matrix multiplication is performed on the query vector Q and the key vector K to obtain the attention score matrix. The attention score matrix is then processed by the softmax activation function to obtain the normalized attention weight matrix. This attention weight matrix is used to represent the degree of attention of the lesion region features to different organ region features. Matrix multiplication is performed on the normalized attention weight matrix and the value vector V to obtain the fused feature vector. This fused feature vector combines the information of the lesion region features and the relevant organ region features, and can better reflect the contextual relationship of the lesion region in a specific anatomical structure. This fused feature vector is the obtained regional fusion feature corresponding to at least one abnormal region.
[0193] In summary, to address the lack of background information on lesion areas, extranodal involvement is characterized by the similarity between lesion area features and organ area features. This fully considers the distribution of lesions and their anatomical correlation with surrounding organs. Furthermore, by characterizing the interaction between lesion areas and anatomical structures, the anatomical information of each lesion area is enhanced, thereby improving the performance of lesion areas in their spatial and functional environments and contributing to more accurate prognostic prediction.
[0194] Further, the step of using the score prediction layer to predict the score of the region fusion features corresponding to the at least one abnormal region to obtain the abnormal score corresponding to the target object includes:
[0195] The feature weights of the region fusion features corresponding to the at least one abnormal region are determined using the score prediction layer.
[0196] The feature weights of the region fusion features corresponding to the at least one abnormal region are calculated to obtain the abnormal score corresponding to the target object.
[0197] Specifically, the score prediction layer can be used to assign feature weights to the regional fusion features corresponding to each abnormal region. Based on the feature weights of the regional fusion features corresponding to each abnormal region, the regional fusion features corresponding to each abnormal region are weighted and summed to obtain the weighted regional fusion features. The weighted regional fusion features are used as patient-level features, and the abnormal scores corresponding to the target object are determined based on the weighted regional fusion features.
[0198] In practical applications, DLBCL faces unique challenges due to the heterogeneity of its anatomical distribution and the diversity of the number of lesions. Approaches using a fixed-size full-image input paradigm or a naive average pooling method applying all lesion features cannot adaptively handle lesion regions with different anatomical locations and pathological significance. To overcome these limitations, an attention-based multi-lesion aggregation module (i.e., a score prediction layer) can be used. This module assigns coefficients (i.e., feature weights) to each lesion region using a gated attention mechanism, and then weights these coefficients to form patient-level features for prognostic prediction. Specifically, the region fusion feature is an anatomically enhanced lesion region feature. After inputting the region fusion feature corresponding to at least one abnormal region into the score prediction layer, the hyperbolic tangent function is used to perform a nonlinear transformation on the input region fusion feature. The transformed region fusion feature is then processed through a softmax activation function to obtain a normalized attention score vector. This attention score vector represents the importance or weight of each region fusion feature. The higher the score, the more important the abnormal region corresponding to that region fusion feature is in the current task. Based on the attention score vector corresponding to each region fusion feature, the region fusion features of each abnormal region are weighted and aggregated to obtain the final aggregated feature vector. This aggregated feature vector contains information about all lesion regions, and this information is weighted and fused according to their importance in the current task. This helps the model better capture and utilize the interrelationships between multiple lesion regions, improving the accuracy of subsequent diagnosis. Furthermore, an abnormality score can be output based on this aggregated feature vector.
[0199] In summary, given the highly variable nature of lymphoma lesions, this problem is characterized as a multi-lesion aggregation problem. A gated attention mechanism is used to assign coefficients to each lesion region, which are then weighted to form patient-level features for prognostic prediction. This attention-based multi-lesion aggregation addresses the issue of uneven lesion distribution, making the learned attention scores interpretable and demonstrating how each lesion location contributes to prognostic prediction.
[0200] Furthermore, before inputting the first image and the second image into the second image processing model, the process further includes:
[0201] Determine the first sample image, the second sample image, and the label anomaly score corresponding to the target object;
[0202] The first sample image and the second sample image are input into a second image processing model to obtain a predicted anomaly score corresponding to the target object. The predicted anomaly score is determined based on the region fusion features corresponding to at least one anomaly region of the target object. The region fusion features corresponding to at least one anomaly region are determined based on the first segmentation result and first image features corresponding to the first sample image, the second segmentation result and second image features corresponding to the second sample image, the fused image features and fused segmentation result between the first sample image and the second sample image, and the target part mask image corresponding to the target object. The target part mask image is obtained by masking the first sample image.
[0203] The second image processing model is trained based on the predicted anomaly score and the labeled anomaly score until a fully trained second image processing model is obtained.
[0204] Among them, the first sample image, the second sample image, and the label anomaly score can be used as training pairs for model training. The first sample image can be understood as a sample CT image, and the second sample image can be understood as a sample PET image.
[0205] Specifically, we can determine the sample CT image, sample PET image, and labeled anomaly score corresponding to the target object. The sample CT image and sample PET image are then input into a second image processing model to obtain the predicted anomaly score output by the target object. Based on this predicted anomaly score and the labeled anomaly score, the second image processing model is trained until a fully trained second image processing model is obtained. Then, in subsequent prognostic assessments, anomaly score predictions can be performed based on this trained first image processing model.
[0206] It is understandable that the process of obtaining the predicted anomaly score corresponding to the target object by processing the first and second sample images based on the second image processing model is similar to the process of obtaining the anomaly score corresponding to the target object by processing the first and second images based on the second image processing model in the above application process. The embodiments in this specification will not be described again here.
[0207] In practice, the model loss value can be calculated based on the predicted anomaly score and the labeled anomaly score. The second image processing model can be trained based on the model loss value until a second image processing model that meets the training stopping condition is obtained as the second image processing model that has been trained. The training stopping condition can be that the model loss value reaches a preset loss value threshold and / or the number of training times reaches a preset number threshold.
[0208] In summary, after determining the first and second images of the target object, these images are input into a first image processing model. The model then predicts abnormal information about the target object based on these images, achieving intelligent recognition of both images without requiring manual identification of abnormal regions by clinicians. Furthermore, the first image processing model can perform feature processing and segmentation on the first and second images separately, and then fuse these processes to obtain a first segmentation result and first image features corresponding to the first image, a second segmentation result and second image features corresponding to the second image, and a fused segmentation result and fused image features between the first and second images. Based on preset image metrics, the second image is segmented to obtain the image segmentation result corresponding to the preset metrics, thus obtaining an abnormality mask image of the target object. This further ensures accurate identification of abnormal regions of the target object, reduces the workload of clinicians, and improves the accuracy and efficiency of abnormal information recognition.
[0209] The following is in conjunction with the appendix Figure 3 Taking the application of the image processing method provided in this specification to the training of the first image processing model as an example, the image processing method will be further explained. Figure 3 A flowchart illustrating the training process of a first image processing model in an image processing method provided in one embodiment of this specification is shown.
[0210] Specifically, after determining the sample CT image and sample PET image corresponding to the target object, the sample CT image and sample PET image are input into the first feature extraction layer in the first image processing model. The CT encoder, CT decoder, PET encoder, PET decoder and PET / CT decoder contained in the first feature extraction layer are used to extract features, and the first image features and first segmentation result corresponding to the sample CT image, the second image features and second segmentation result corresponding to the sample PET image, and the fused image features and fused segmentation result between the sample CT image and the sample PET image are obtained.
[0211] Sample PET images are input into the first segmentation layer. In this layer, the sample PET images are processed using an empirical thresholding method and an SUV thresholding method to obtain an SUV threshold map corresponding to the sample PET image. Convolutional blocks are then used to perform feature processing on the SUV threshold map, yielding the image features corresponding to this SUV threshold map as the image segmentation result. Sample CT images are input into a deep learning model designed for medical image segmentation. A full-body mask image of the target object is segmented from the sample CT image. Based on the reference region corresponding to the target object, a reference region mask image is determined from the full-body mask image.
[0212] The first image features, the first segmentation result, the second image features, the second segmentation result, the fused image features, the fused segmentation result, and the image segmentation result are input into the first feature fusion layer to obtain the predicted anomaly mask image output by the first feature fusion layer. In the first feature fusion layer, the first image features and the first segmentation result, the second image features and the second segmentation result, the fused image features and the fused segmentation result, and the image segmentation result can be input into the first feature fusion layer. In the first feature fusion layer, logarithmic probability maps (i.e., logits maps) for both CT and PET modalities are obtained through logarithmic unit extraction blocks, and probability maps (i.e., probability maps) for both CT and PET modalities are obtained through the aforementioned probability extraction blocks. The logarithmic probability maps and probability maps are fused and an activation function is used to obtain the anomaly mask image corresponding to the target object. From the predicted anomaly mask image, determine the mask image containing anomaly mask information and the mask image containing location mask information. Based on the mask image containing anomaly mask information and the label anomaly mask image, calculate the first model loss value. Based on the mask image containing location mask information and the reference location mask image, calculate the second model loss value. Train the first image processing model based on the first model loss value and the second model loss value until a first image processing model that meets the training stopping condition is obtained, which is the first image processing model that has been trained.
[0213] The following is in conjunction with the appendix Figure 4Taking the application of the image processing method provided in this specification to the training of a second image processing model as an example, the image processing method will be further explained. Figure 4 This document illustrates a flowchart of the training process of a second image processing model in an image processing method according to an embodiment of this specification. Specifically, after determining the sample CT image and sample PET image corresponding to the target object, the sample CT image and sample PET image are input into the second image processing model. The model sequentially passes through a first feature extraction layer and a first feature fusion layer to obtain the first image features and first segmentation result corresponding to the sample CT image, the second image features and second segmentation result corresponding to the sample PET image, the fused image features and fused segmentation result between the sample CT image and sample PET image, and a reference anomaly mask image corresponding to the target object. The sample CT image is then input into a deep learning model designed based on medical image segmentation to segment a full-body mask image of the target object from the sample CT image. This full-body mask image is used as the target part mask image of the target object. In the second feature fusion layer, the first segmentation result, first image features, second segmentation result, second image features, fused segmentation result, fused image features, reference anomaly mask image, and target site mask image are fused to obtain regional features corresponding to each abnormal region in at least one abnormal region corresponding to the target object, and site features corresponding to each target site in at least one target site. This enables similarity analysis between lesion region features and organ region features. An attention mechanism is then applied to the regional features corresponding to each abnormal region and the site features corresponding to each target site to obtain regional fusion features corresponding to at least one abnormal region. In the score prediction layer, feature weights are assigned to the regional fusion features corresponding to each abnormal region. Based on these feature weights, the regional fusion features corresponding to each abnormal region are weighted and summed to obtain weighted regional fusion features. These weighted regional fusion features serve as patient-level features, and the predicted anomaly score corresponding to the target object is determined based on these weighted regional fusion features. The second image processing model is trained based on the predicted anomaly score and the labeled anomaly score until a fully trained second image processing model is obtained.
[0214] In specific implementation, in the second feature fusion layer, when processing the regional features (i.e., lesion region features) corresponding to at least one abnormal region and the site features (i.e., organ region features) corresponding to at least one target site, the lesion region features are transformed by the query weight matrix to obtain the query vector Q. The organ region features are transformed by the key weight matrix and the value weight matrix to obtain the key vector K and the value vector V, respectively. The query vector Q and the key vector K are multiplied by matrix to obtain the attention score matrix. The attention score matrix is processed by the softmax activation function to obtain the normalized attention weight matrix. The attention weight matrix is used to represent the degree of attention of the lesion region features to different organ region features. The normalized attention weight matrix is multiplied by the value vector V to obtain the fused feature vector. The fused feature vector is the regional fusion feature corresponding to at least one abnormal region.
[0215] In the score prediction layer, the hyperbolic tangent function is used to perform a nonlinear transformation on the input region fusion features. The transformed region fusion features are then processed by the softmax activation function to obtain a normalized attention score vector. This attention score vector represents the importance or weight of each region fusion feature. The higher the score, the more important the abnormal region corresponding to that region fusion feature is in the current task. Based on the attention score vector corresponding to each region fusion feature, the region fusion features of each abnormal region are weighted and aggregated to obtain the final aggregated feature vector. Furthermore, the abnormal score can be output based on this aggregated feature vector.
[0216] Corresponding to the above method embodiments, this specification also provides embodiments of an image processing apparatus. Figure 5 A schematic diagram of the structure of an image processing apparatus provided in one embodiment of this specification is shown. Figure 5 As shown, the device includes:
[0217] The determination module 502 is configured to determine the first image and the second image corresponding to the target object;
[0218] Input module 504 is configured to input the first image and the second image into a first image processing model to obtain an anomaly mask image corresponding to the target object;
[0219] The abnormal mask image is determined based on the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, the fusion segmentation result and fusion image features between the first image and the second image, and the image segmentation result corresponding to the preset image index. The image segmentation result is obtained by segmenting the second image based on the preset image index.
[0220] In one optional embodiment, the first image processing model includes a first feature extraction layer, a first segmentation layer, and a first feature fusion layer;
[0221] The input module 504 is further configured as follows:
[0222] The first image and the second image are input into the first image processing model. The first feature extraction layer is used to extract and fuse features of the first image and the second image to obtain the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, and the fused segmentation result and fused image features between the first image and the second image.
[0223] Using the first segmentation layer, the second image is segmented based on the preset image index to obtain the image segmentation result;
[0224] Using the first feature fusion layer, the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, the fused segmentation result and fused image features between the first image and the second image, and the image segmentation result are fused to obtain the abnormal mask image corresponding to the target object.
[0225] In one optional embodiment, the first feature extraction layer includes a first encoder, a second encoder, a first decoder, a second decoder, and a fusion decoder;
[0226] The input module 504 is further configured as follows:
[0227] The first image is input into the first encoder to obtain the first encoded feature corresponding to the first image;
[0228] The second image is input into the second encoder to obtain the second encoded feature corresponding to the second image;
[0229] The first encoded feature is input into the first decoder to obtain the first segmentation result and the first image feature corresponding to the first image;
[0230] The second encoded feature is input into the second decoder to obtain the second segmentation result and the second image feature corresponding to the second image;
[0231] The first encoded feature and the second encoded feature are input into the fusion decoder to obtain the fusion segmentation result between the first image and the second image, as well as the fusion image features.
[0232] In an optional embodiment, the input module 504 is further configured to:
[0233] Using the first feature fusion layer, the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, the fused segmentation result and fused image features between the first image and the second image, and the image segmentation result are fused to obtain the abnormal segmentation result corresponding to the target object and the probability corresponding to the abnormal segmentation result.
[0234] Based on the anomaly segmentation results and the probabilities corresponding to the anomaly segmentation results, the anomaly mask image corresponding to the target object is determined.
[0235] In an optional embodiment, the device further includes a training module configured to:
[0236] Determine the first sample image, the second sample image, and the label anomaly mask image corresponding to the target object;
[0237] The first sample image and the second sample image are input into a first image processing model to obtain a predicted anomaly mask image corresponding to the target object. The predicted anomaly mask image is determined based on a first segmentation result and first image features corresponding to the first sample image, a second segmentation result and second image features corresponding to the second sample image, fused image features and fused segmentation results between the first sample image and the second sample image, and image segmentation results corresponding to preset image indicators. The image segmentation results are obtained by segmenting the second sample image based on the preset image indicators.
[0238] The first image processing model is trained based on the predicted anomaly mask image, the labeled anomaly mask image, and the reference part mask image corresponding to the target object until a trained first image processing model is obtained. The reference part mask image is obtained by masking the sample first image.
[0239] In an optional embodiment, the training module is further configured to:
[0240] The first model loss value is calculated based on the anomaly mask information contained in the labeled anomaly mask image and the predicted anomaly mask image.
[0241] The second model loss value is calculated based on the part mask information contained in the reference part mask image and the predicted anomaly mask image.
[0242] The first image processing model is trained based on the first model loss value and the second model loss value until a fully trained first image processing model is obtained.
[0243] In an optional embodiment, the training module is further configured to:
[0244] Based on the reference region corresponding to the target object, the first image of the sample is segmented to obtain a reference region mask image corresponding to the reference region.
[0245] In an optional embodiment, the input module 504 is further configured to:
[0246] The first image and the second image are input into the second image processing model to obtain the anomaly score corresponding to the target object;
[0247] The anomaly score is determined based on the region fusion features corresponding to at least one anomaly region of the target object. The region fusion features corresponding to at least one anomaly region are determined based on the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, the fusion segmentation result and fusion image features between the first image and the second image, and the target part mask image corresponding to the target object. The target part mask image is obtained by masking the first image.
[0248] In one optional embodiment, the second image processing model includes a first feature extraction layer, a first feature fusion layer, a second segmentation layer, a second feature fusion layer, and a score prediction layer;
[0249] The input module 504 is further configured as follows:
[0250] The first image and the second image are input into the second image processing model. The first feature extraction layer is used to extract and fuse features of the first image and the second image to obtain the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, and the fused segmentation result and fused image features between the first image and the second image.
[0251] Using the first feature fusion layer, the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, and the fused segmentation result and fused image features between the first image and the second image are fused to obtain the reference anomaly mask image corresponding to the target object.
[0252] Using the second segmentation layer, the first image is segmented according to the target part corresponding to the target object to obtain the target part mask image corresponding to the target part;
[0253] Using the second feature fusion layer, feature fusion processing is performed on the first segmentation result and the first image features, the second segmentation result and the second image features, the fused segmentation result and the fused image features, the reference anomaly mask image and the target part mask image to obtain the region fusion features corresponding to at least one anomaly region of the target object;
[0254] Using the score prediction layer, score prediction is performed on the region fusion features corresponding to the at least one abnormal region to obtain the abnormal score corresponding to the target object.
[0255] In an optional embodiment, the input module 504 is further configured to:
[0256] Using the second feature fusion layer, the first segmentation result and the first image features, the second segmentation result and the second image features, the fused segmentation result and the fused image features, the reference anomaly mask image and the target part mask image are fused to obtain the region features corresponding to at least one anomaly region of the target object and the part features corresponding to at least one target part.
[0257] An attention mechanism is applied to the region features corresponding to the at least one abnormal region and the location features corresponding to the at least one target location to obtain the region fusion features corresponding to at least one abnormal region of the target object.
[0258] In an optional embodiment, the input module 504 is further configured to:
[0259] The feature weights of the region fusion features corresponding to the at least one abnormal region are determined using the score prediction layer.
[0260] The feature weights of the region fusion features corresponding to the at least one abnormal region are calculated to obtain the abnormal score corresponding to the target object.
[0261] In an optional embodiment, the training module is further configured to:
[0262] Determine the first sample image, the second sample image, and the label anomaly score corresponding to the target object;
[0263] The first sample image and the second sample image are input into a second image processing model to obtain a predicted anomaly score corresponding to the target object. The predicted anomaly score is determined based on the region fusion features corresponding to at least one anomaly region of the target object. The region fusion features corresponding to at least one anomaly region are determined based on the first segmentation result and first image features corresponding to the first sample image, the second segmentation result and second image features corresponding to the second sample image, the fused image features and fused segmentation result between the first sample image and the second sample image, and the target part mask image corresponding to the target object. The target part mask image is obtained by masking the first sample image.
[0264] The second image processing model is trained based on the predicted anomaly score and the labeled anomaly score until a fully trained second image processing model is obtained.
[0265] In summary, after determining the first and second images of the target object, these images are input into a first image processing model. The model then predicts abnormal information about the target object based on these images, achieving intelligent recognition of both images without requiring manual identification of abnormal regions by clinicians. Furthermore, the first image processing model can perform feature processing and segmentation on the first and second images separately, and then fuse these processes to obtain a first segmentation result and first image features corresponding to the first image, a second segmentation result and second image features corresponding to the second image, and a fused segmentation result and fused image features between the first and second images. Based on preset image metrics, the second image is segmented to obtain the image segmentation result corresponding to the preset metrics, thus obtaining an abnormality mask image of the target object. This further ensures accurate identification of abnormal regions of the target object, reduces the workload of clinicians, and improves the accuracy and efficiency of abnormal information recognition.
[0266] The above is an illustrative scheme of an image processing apparatus according to this embodiment. It should be noted that the technical solution of this image processing apparatus and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the image processing apparatus, please refer to the description of the technical solution of the image processing method described above.
[0267] Corresponding to the above method embodiments, this specification also provides another image processing method, see [link to documentation]. Figure 6 , Figure 6 A flowchart of another image processing method according to an embodiment of this specification is shown, which specifically includes the following steps.
[0268] Step 602: Determine the first and second images corresponding to the target object;
[0269] Step 604: Input the first image and the second image into the second image processing model to obtain the anomaly score corresponding to the target object;
[0270] The anomaly score is determined based on the region fusion features corresponding to at least one anomaly region of the target object. The region fusion features corresponding to at least one anomaly region are determined based on the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, the fusion segmentation result and fusion image features between the first image and the second image, and the target part mask image corresponding to the target object. The target part mask image is obtained by masking the first image.
[0271] It is understood that the second image processing model in this image processing method is the second image processing model in the above-mentioned image processing method. The process of processing the first image and the second image using the second image processing model and the training process of the second image processing model are similar to those described above. The embodiments in this specification will not be repeated here.
[0272] Corresponding to the above-described method embodiments, this specification also provides another image processing apparatus. Figure 7 A schematic diagram of another image processing apparatus provided in one embodiment of this specification is shown. Figure 7 As shown, the device includes:
[0273] The determination module 702 is configured to determine the first image and the second image corresponding to the target object;
[0274] Input module 704 is configured to input the first image and the second image into a second image processing model to obtain the anomaly score corresponding to the target object;
[0275] The anomaly score is determined based on the region fusion features corresponding to at least one anomaly region of the target object. The region fusion features corresponding to at least one anomaly region are determined based on the first segmentation result and first image features corresponding to the first image, the second segmentation result and second image features corresponding to the second image, the fusion segmentation result and fusion image features between the first image and the second image, and the target part mask image corresponding to the target object. The target part mask image is obtained by masking the first image.
[0276] The above is an illustrative scheme of an image processing apparatus according to this embodiment. It should be noted that the technical solution of this image processing apparatus and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the image processing apparatus, please refer to the description of the technical solution of the image processing method described above.
[0277] Corresponding to the above method embodiments, this specification also provides an image processing model training method, see [link to documentation]. Figure 8 , Figure 8 A flowchart of an image processing model training method according to an embodiment of this specification is shown, which specifically includes the following steps.
[0278] Step 802: Determine the first sample image, the second sample image, and the label anomaly mask image corresponding to the target object;
[0279] Step 804: Input the first sample image and the second sample image into the first image processing model to obtain the predicted anomaly mask image corresponding to the target object. The predicted anomaly mask image is determined based on the first segmentation result and first image features corresponding to the first sample image, the second segmentation result and second image features corresponding to the second sample image, the fused image features and fused segmentation result between the first sample image and the second sample image, and the image segmentation result corresponding to the preset image index. The image segmentation result is obtained by segmenting the second sample image based on the preset image index.
[0280] Step 806: Train the first image processing model based on the predicted anomaly mask image, the label anomaly mask image, and the reference part mask image corresponding to the target object until a trained first image processing model is obtained, wherein the reference part mask image is obtained by masking the sample first image.
[0281] It is understood that the training method of this image processing model is similar to the training process of the aforementioned first image processing model, and the embodiments in this specification do not limit it.
[0282] Corresponding to the above method embodiments, this specification also provides an image processing model training device. Figure 9 A schematic diagram of an image processing model training apparatus according to one embodiment of this specification is shown. Figure 9 As shown, the device includes:
[0283] The determination module 902 is configured to determine the first sample image, the second sample image, and the label anomaly mask image corresponding to the target object;
[0284] Input module 904 is configured to input the first sample image and the second sample image into a first image processing model to obtain a predicted anomaly mask image corresponding to the target object. The predicted anomaly mask image is determined based on a first segmentation result and first image features corresponding to the first sample image, a second segmentation result and second image features corresponding to the second sample image, fused image features and fused segmentation results between the first sample image and the second sample image, and image segmentation results corresponding to preset image indicators. The image segmentation results are obtained by segmenting the second sample image based on the preset image indicators.
[0285] The training module 906 is configured to train the first image processing model based on the predicted anomaly mask image, the labeled anomaly mask image, and the reference part mask image corresponding to the target object, until a trained first image processing model is obtained, wherein the reference part mask image is obtained by masking the sample first image.
[0286] The above is a schematic scheme of an image processing model training device according to this embodiment. It should be noted that the technical solution of this image processing model training device and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the image processing model training device, please refer to the description of the technical solution of the image processing method described above.
[0287] Corresponding to the above-described method embodiments, this specification also provides another image processing model training method, see [link to documentation]. Figure 10 , Figure 10 A flowchart of another image processing model training method according to an embodiment of this specification is shown, which specifically includes the following steps.
[0288] Step 1002: Determine the first sample image, the second sample image, and the label anomaly score corresponding to the target object;
[0289] Step 1004: Input the first sample image and the second sample image into the second image processing model to obtain the predicted anomaly score corresponding to the target object. The predicted anomaly score is determined based on the region fusion features corresponding to at least one anomaly region of the target object. The region fusion features corresponding to at least one anomaly region are determined based on the first segmentation result and first image features corresponding to the first sample image, the second segmentation result and second image features corresponding to the second sample image, the fused image features and fused segmentation result between the first sample image and the second sample image, and the target part mask image corresponding to the target object. The target part mask image is obtained by masking the first sample image.
[0290] Step 1006: Train the second image processing model based on the predicted anomaly score and the labeled anomaly score until a trained second image processing model is obtained.
[0291] It is understood that the training method of this image processing model is similar to the training process of the aforementioned second image processing model, and the embodiments in this specification do not limit it.
[0292] Corresponding to the above method embodiments, this specification also provides another image processing model training device. Figure 11 A schematic diagram of another image processing model training apparatus provided in one embodiment of this specification is shown. Figure 11 As shown, the device includes:
[0293] The determination module 1102 is configured to determine the first sample image, the second sample image, and the label anomaly score corresponding to the target object;
[0294] Input module 1104 is configured to input the first sample image and the second sample image into a second image processing model to obtain a predicted anomaly score corresponding to the target object. The predicted anomaly score is determined based on the region fusion features corresponding to at least one anomaly region of the target object. The region fusion features corresponding to at least one anomaly region are determined based on the first segmentation result and first image features corresponding to the first sample image, the second segmentation result and second image features corresponding to the second sample image, the fused image features and fused segmentation result between the first sample image and the second sample image, and the target part mask image corresponding to the target object. The target part mask image is obtained by masking the first sample image.
[0295] The training module 1106 is configured to train the second image processing model based on the predicted anomaly score and the labeled anomaly score until a trained second image processing model is obtained.
[0296] The above is a schematic scheme of an image processing model training device according to this embodiment. It should be noted that the technical solution of this image processing model training device and the technical solution of the image processing method described above belong to the same concept. For details not described in detail in the technical solution of the image processing model training device, please refer to the description of the technical solution of the image processing method described above.
[0297] Corresponding to the above-described method embodiments, this specification also provides a CT image processing method, see [link to documentation]. Figure 12 , Figure 12 A flowchart of a CT image processing method according to an embodiment of this specification is shown, which specifically includes the following steps.
[0298] Step 1202: Receive a CT image processing task, wherein the CT image processing task carries a CT image and a PET image corresponding to the target object, and the CT image processing task is used to detect abnormal information of the target object;
[0299] Step 1204: Input the CT image and the PET image into the first image processing model to obtain the anomaly mask image for the target object output by the first image processing model, wherein the first image processing model is trained according to the above-described image processing model training method; and / or
[0300] The CT image and the PET image are input into the second image processing model to obtain the anomaly score for the target object output by the second image processing model, wherein the second image processing model is trained according to the above-described image processing model training method.
[0301] The above is an illustrative scheme of a CT image processing method according to this embodiment. It should be noted that the technical solution of this CT image processing method belongs to the same concept as the technical solution of the image processing method described above. For details not described in detail in the technical solution of the CT image processing method, please refer to the description of the technical solution of the image processing method described above.
[0302] Corresponding to the above method embodiments, this specification also provides an image processing method applied to cloud-side devices, see [link to documentation]. Figure 13 , Figure 13 A flowchart of another image processing method according to an embodiment of this specification is shown, which specifically includes the following steps.
[0303] Step 1302: Receive an image processing task sent by the receiving end device, wherein the image processing task carries a first image and a second image corresponding to the target object, and the image processing task is used to detect abnormal information of the target object;
[0304] Step 1304: Input the first image and the second image into the first image processing model to obtain the anomaly mask image for the target object output by the first image processing model, wherein the first image processing model is trained according to the above-described image processing model training method; and / or
[0305] The first image and the second image are input into the second image processing model to obtain the anomaly score for the target object output by the second image processing model, wherein the second image processing model is trained according to the above image processing model training method;
[0306] Step 1306: Send the anomaly mask image and / or anomaly score to the end-side device.
[0307] The above is an illustrative scheme of an image processing method according to this embodiment. It should be noted that the technical solution of this image processing method belongs to the same concept as the technical solution of the image processing method described above. For details not described in detail in the technical solution of the image processing method, please refer to the description of the technical solution of the image processing method described above.
[0308] Corresponding to the above-described method embodiments, this specification also provides a computer-aided diagnostic method for tumors, see [link to documentation]. Figure 14 , Figure 14 A flowchart of a computer-aided diagnostic method for tumors according to an embodiment of this specification is shown, which specifically includes the following steps.
[0309] Step 1402: Receive a tumor screening task, wherein the tumor screening task carries CT images and PET images corresponding to the target object, and the CT image processing task is used to detect tumor information of the target object;
[0310] Step 1404: Input the CT image and the PET image into the first image processing model to obtain an anomaly mask image for the target object output by the first image processing model, wherein the first image processing model is trained according to the above-described image processing model training method; and / or
[0311] The CT image and the PET image are input into the second image processing model to obtain the anomaly score for the target object output by the second image processing model, wherein the second image processing model is trained according to the above-described image processing model training method.
[0312] The above is an illustrative scheme of a computer-aided diagnosis method for tumors according to this embodiment. It should be noted that the technical solution of this computer-aided diagnosis method for tumors belongs to the same concept as the technical solution of the image processing method described above. Details not described in detail in the technical solution of the computer-aided diagnosis method for tumors can be found in the description of the technical solution of the image processing method described above.
[0313] Corresponding to the above-described method embodiments, this specification also provides a computer-aided diagnostic system for tumors. Figure 15 A schematic diagram of the structure of a computer-aided diagnostic system for tumors according to one embodiment of this specification is shown. Figure 15 As shown, the system includes a client 1502 and a server 1504, wherein,
[0314] The client 1502 is used to send a CT image processing task to the server 1504, wherein the CT image processing task carries a CT image and a PET image corresponding to the target object, and the CT image processing task is used to detect abnormal information of the target object;
[0315] The server 1504 is used to input the CT image and the PET image into a first image processing model to obtain an anomaly mask image for the target object output by the first image processing model, wherein the first image processing model is trained according to the above-described image processing model training method; and / or
[0316] The CT image and the PET image are input into the second image processing model to obtain the anomaly score for the target object output by the second image processing model, wherein the second image processing model is trained according to the above image processing model training method;
[0317] The server 1504 is also used to send the anomaly mask image and / or the anomaly score to the client 1502.
[0318] The above is an illustrative scheme of a computer-aided diagnosis system for tumors according to this embodiment. It should be noted that the technical solution of this computer-aided diagnosis system for tumors and the technical solution of the image processing method described above belong to the same concept. Details not described in detail in the technical solution of the computer-aided diagnosis system for tumors can be found in the description of the technical solution of the image processing method described above.
[0319] Figure 16 A structural block diagram of a computing device 1600 according to one embodiment of this specification is shown. The components of the computing device 1600 include, but are not limited to, a memory 1610 and a processor 1620. The processor 1620 is connected to the memory 1610 via a bus 1630, and a database 1650 is used to store data.
[0320] The computing device 1600 also includes an access device 1640, which enables the computing device 1600 to communicate via one or more networks 1660. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 1640 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0321] In one embodiment of this application, the aforementioned components of the computing device 1600 and Figure 16 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 16 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this application. Those skilled in the art can add or replace other components as needed.
[0322] The computing device 1600 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 1600 can also be a mobile or stationary server.
[0323] The processor 1620 is used to execute the following computer program / instructions, which, when executed by the processor, implement the steps of the above method.
[0324] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computing device embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0325] An embodiment of this specification also provides a computer-readable storage medium storing a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0326] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the computer-readable storage medium embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.
[0327] An embodiment of this specification also provides a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of the above-described method.
[0328] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solution of the above method belong to the same concept, and all details not described in detail in the technical solution of the computer program product can be referred to the description of the technical solution of the above method.
[0329] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0330] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0331] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0332] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0333] The preferred embodiments disclosed above are merely illustrative of this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described herein. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.
Claims
1. An image processing method comprising: determining a first image and a second image corresponding to a target object; inputting the first image and the second image into a first image processing model to obtain an abnormal mask image corresponding to the target object; wherein the abnormal mask image is determined according to a first segmentation result corresponding to the first image and a first image feature, a second segmentation result corresponding to the second image and a second image feature, a fusion segmentation result between the first image and the second image and a fusion image feature, and an image segmentation result corresponding to a preset image index, the image segmentation result being obtained by segmenting the second image based on the preset image index.
2. The method of claim 1, wherein the first image processing model comprises a first feature extraction layer, a first segmentation layer, and a first feature fusion layer; the inputting the first image and the second image into the first image processing model to obtain the abnormal mask image corresponding to the target object comprises: inputting the first image and the second image into the first image processing model, and using the first feature extraction layer to extract and fuse features of the first image and the second image to obtain the first segmentation result corresponding to the first image and the first image feature, the second segmentation result corresponding to the second image and the second image feature, the fusion segmentation result between the first image and the second image and the fusion image feature; using the first segmentation layer to segment the second image based on the preset image index to obtain the image segmentation result; using the first feature fusion layer to fuse the first segmentation result corresponding to the first image and the first image feature, the second segmentation result corresponding to the second image and the second image feature, the fusion segmentation result between the first image and the second image and the fusion image feature, and the image segmentation result to obtain the abnormal mask image corresponding to the target object.
3. The method of claim 2, wherein the first feature extraction layer comprises a first encoder, a second encoder, a first decoder, a second decoder, and a fusion decoder; the using the first feature extraction layer to extract and fuse features of the first image and the second image to obtain the first segmentation result corresponding to the first image and the first image feature, the second segmentation result corresponding to the second image and the second image feature, the fusion segmentation result between the first image and the second image and the fusion image feature comprises: inputting the first image into the first encoder to obtain first encoded features corresponding to the first image; inputting the second image into the second encoder to obtain second encoded features corresponding to the second image; inputting the first encoded features into the first decoder to obtain the first segmentation result corresponding to the first image and the first image feature; inputting the second encoded features into the second decoder to obtain the second segmentation result corresponding to the second image and the second image feature; inputting the first encoding feature and the second encoding feature into the fusion decoder to obtain a fusion segmentation result between the first image and the second image and a fusion image feature.
4. The method of claim 2, wherein the fusion processing of the first image corresponding first segmentation result and first image feature, the second image corresponding second segmentation result and second image feature, the fusion segmentation result between the first image and the second image and the fusion image feature, and the image segmentation result by the first feature fusion layer to obtain the target object corresponding abnormal mask image comprises: fusion processing of the first image corresponding first segmentation result and first image feature, the second image corresponding second segmentation result and second image feature, the fusion segmentation result between the first image and the second image and the fusion image feature, and the image segmentation result by the first feature fusion layer to obtain the target object corresponding abnormal segmentation result and the probability corresponding to the abnormal segmentation result; determining the target object corresponding abnormal mask image according to the abnormal segmentation result and the probability corresponding to the abnormal segmentation result.
5. The method of any one of claims 1-4, wherein before the first image and the second image are input into the first image processing model, the method further comprises: determining a sample first image, a sample second image and a label abnormal mask image corresponding to a target object; inputting the sample first image and the sample second image into the first image processing model to obtain a predicted abnormal mask image corresponding to the target object, wherein the predicted abnormal mask image is determined according to a first segmentation result and a first image feature corresponding to the sample first image, a second segmentation result and a second image feature corresponding to the sample second image, a fusion image feature and a fusion segmentation result between the sample first image and the sample second image, and an image segmentation result corresponding to a preset image index, the image segmentation result being obtained by segmenting the sample second image based on the preset image index; training the first image processing model according to the predicted abnormal mask image, the label abnormal mask image and a reference part mask image corresponding to the target object until a trained first image processing model is obtained, wherein the reference part mask image is obtained by masking the sample first image.
6. The method of claim 5, wherein the training of the first image processing model according to the predicted abnormal mask image, the label abnormal mask image and the reference part mask image corresponding to the target object until the trained first image processing model is obtained comprises: calculating a first model loss value according to abnormal mask information contained in the label abnormal mask image and the predicted abnormal mask image; calculating a second model loss value according to part mask information contained in the reference part mask image and the predicted abnormal mask image; and training the first image processing model according to the first model loss value and the second model loss value until the trained first image processing model is obtained. According to the first model loss value and the second model loss value, the first image processing model is trained until a trained first image processing model is obtained.
7. The method of claim 5, before the training of the first image processing model according to the predicted anomaly mask image, the label anomaly mask image and the reference part mask image corresponding to the target object until the trained first image processing model is obtained, further comprising: segmenting the sample first image according to the reference part corresponding to the target object to obtain a reference part mask image corresponding to the reference part.
8. The method of claim 1, after the determination of the first image and the second image corresponding to the target object, further comprising: inputting the first image and the second image into a second image processing model to obtain an anomaly score corresponding to the target object; wherein the anomaly score is determined according to a region fusion feature corresponding to at least one abnormal region of the target object, and the region fusion feature corresponding to the at least one abnormal region is determined according to a first segmentation result and a first image feature corresponding to the first image, a second segmentation result and a second image feature corresponding to the second image, a fusion segmentation result and a fusion image feature between the first image and the second image, and a target part mask image corresponding to the target object, the target part mask image being obtained by masking the first image.
9. The method of claim 8, the second image processing model comprising a first feature extraction layer, a first feature fusion layer, a second segmentation layer, a second feature fusion layer and a score prediction layer; the inputting of the first image and the second image into the second image processing model to obtain the anomaly score corresponding to the target object comprises: inputting the first image and the second image into the second image processing model, and using the first feature extraction layer to extract and fuse features of the first image and the second image to obtain the first segmentation result and the first image feature corresponding to the first image, the second segmentation result and the second image feature corresponding to the second image, and the fusion segmentation result and the fusion image feature between the first image and the second image; using the first feature fusion layer to fuse the first segmentation result and the first image feature corresponding to the first image, the second segmentation result and the second image feature corresponding to the second image, and the fusion segmentation result and the fusion image feature between the first image and the second image to obtain a reference anomaly mask image corresponding to the target object; using the second segmentation layer to segment the first image according to the target part corresponding to the target object to obtain a target part mask image corresponding to the target part; using the second feature fusion layer to fuse the reference anomaly mask image corresponding to the target object and the target part mask image corresponding to the target part to obtain the anomaly score corresponding to the target object; and using the score prediction layer to predict the anomaly score corresponding to the target object. The second feature fusion layer is used for performing feature fusion processing on the first segmentation result and the first image feature, the second segmentation result and the second image feature, the fusion segmentation result and the fusion image feature, the reference abnormality mask image and the target part mask image, to obtain region fusion features corresponding to at least one abnormal region of the target object; The score prediction layer is used for performing score prediction on the region fusion features corresponding to the at least one abnormal region, to obtain an abnormality score corresponding to the target object.
10. The method of claim 9, wherein the step of using the second feature fusion layer to perform feature fusion processing on the first segmentation result and the first image feature, the second segmentation result and the second image feature, the fusion segmentation result and the fusion image feature, the reference abnormality mask image and the target part mask image, to obtain region fusion features corresponding to at least one abnormal region of the target object comprises: The second feature fusion layer is used for performing fusion processing on the first segmentation result and the first image feature, the second segmentation result and the second image feature, the fusion segmentation result and the fusion image feature, the reference abnormality mask image and the target part mask image, to obtain region features corresponding to the at least one abnormal region of the target object and part features corresponding to at least one target part; The attention mechanism is used for performing attention mechanism processing on the region features corresponding to the at least one abnormal region and the part features corresponding to the at least one target part, to obtain the region fusion features corresponding to the at least one abnormal region of the target object.
11. The method of claim 9, wherein the step of using the score prediction layer to perform score prediction on the region fusion features corresponding to the at least one abnormal region, to obtain an abnormality score corresponding to the target object comprises: The score prediction layer is used for determining feature weights of the region fusion features corresponding to the at least one abnormal region; The feature weights of the region fusion features corresponding to the at least one abnormal region are calculated, to obtain the abnormality score corresponding to the target object.
12. The method of claim 8, wherein before the first image and the second image are input into the second image processing model, the method further comprises: determining sample first images, sample second images and label abnormality scores corresponding to a target object; inputting the sample first image and the sample second image into a second image processing model to obtain a predicted anomaly score corresponding to the target object, wherein The predicted abnormality score is determined according to region fusion features corresponding to at least one abnormal region of the target object, the region fusion features corresponding to the at least one abnormal region of the target object being determined according to a first segmentation result and a first image feature corresponding to a sample first image, a second segmentation result and a second image feature corresponding to a sample second image, a fusion image feature and a fusion segmentation result between the sample first image and the sample second image, and a target part mask image corresponding to the target object, the target part mask image being obtained by performing mask processing on the sample first image; The second image processing model is trained according to the predicted abnormality score and the label abnormality score, until a trained second image processing model is obtained.
13. An image processing method, comprising: determining a first image and a second image corresponding to a target object; inputting the first image and the second image into a second image processing model to obtain an abnormal score corresponding to the target object; wherein the abnormal score is determined according to a region fusion feature corresponding to at least one abnormal region of the target object, and the region fusion feature corresponding to the at least one abnormal region is determined according to a first segmentation result and a first image feature corresponding to the first image, a second segmentation result and a second image feature corresponding to the second image, a fusion segmentation result and a fusion image feature between the first image and the second image, and a target part mask image corresponding to the target object, the target part mask image being obtained by performing mask processing on the first image.
14. An image processing model training method, comprising: determining a sample first image, a sample second image, and a label abnormal mask image corresponding to a target object; inputting the sample first image and the sample second image into a first image processing model to obtain a predicted abnormal mask image corresponding to the target object, wherein the predicted abnormal mask image is determined according to a first segmentation result and a first image feature corresponding to the sample first image, a second segmentation result and a second image feature corresponding to the sample second image, a fusion image feature and a fusion segmentation result between the sample first image and the sample second image, and an image segmentation result corresponding to a preset image index, the image segmentation result being obtained by performing segmentation on the sample second image based on the preset image index; training the first image processing model according to the predicted abnormal mask image, the label abnormal mask image, and a reference part mask image corresponding to the target object until a trained first image processing model is obtained, wherein the reference part mask image is obtained by performing mask processing on the sample first image.
15. An image processing model training method, comprising: determining a sample first image, a sample second image, and a label abnormal score corresponding to a target object; inputting the sample first image and the sample second image into a second image processing model to obtain a predicted abnormal score corresponding to the target object, wherein the predicted abnormal score is determined according to a region fusion feature corresponding to at least one abnormal region of the target object, and the region fusion feature corresponding to the at least one abnormal region is determined according to a first segmentation result and a first image feature corresponding to the sample first image, a second segmentation result and a second image feature corresponding to the sample second image, a fusion image feature and a fusion segmentation result between the sample first image and the sample second image, and a target part mask image corresponding to the target object, the target part mask image being obtained by performing mask processing on the sample first image; training the second image processing model according to the predicted abnormal score and the label abnormal score until a trained second image processing model is obtained.
16. A CT image processing method, comprising: receive a CT image processing task, wherein the CT image processing task carries a CT image and a PET image corresponding to a target object, and the CT image processing task is used to detect abnormal information of the target object; input the CT image and the PET image into a first image processing model to obtain an abnormal mask image for the target object output by the first image processing model, wherein the first image processing model is trained according to the method in claim 14; and / or input the CT image and the PET image into a second image processing model to obtain an abnormal score for the target object output by the second image processing model, wherein the second image processing model is trained according to the method in claim 15.
17. An image processing method applied to a cloud-side device, comprising: receiving an image processing task sent by an end-side device, wherein the image processing task carries a first image and a second image corresponding to a target object, and the image processing task is used to detect abnormal information of the target object; inputting the first image and the second image into a first image processing model to obtain an abnormal mask image for the target object output by the first image processing model, wherein the first image processing model is trained according to the method in claim 14; and / or inputting the first image and the second image into a second image processing model to obtain an abnormal score for the target object output by the second image processing model, wherein the second image processing model is trained according to the method in claim 15; sending the abnormal mask image and / or the abnormal score to the end-side device.
18. A computer-aided diagnosis method of a tumor, comprising: receiving a tumor screening task, wherein the tumor screening task carries a CT image and a PET image corresponding to a target object, and the CT image processing task is used to detect tumor information of the target object; inputting the CT image and the PET image into a first image processing model to obtain an abnormal mask image for the target object output by the first image processing model, wherein the first image processing model is trained according to the method in claim 14; and / or inputting the CT image and the PET image into a second image processing model to obtain an abnormal score for the target object output by the second image processing model, wherein the second image processing model is trained according to the method in claim 15.
19. A computer-aided diagnosis system of a tumor, comprising a client and a server, wherein the client is configured to send a CT image processing task to the server, wherein the CT image processing task carries a CT image and a PET image corresponding to a target object, and the CT image processing task is used to detect abnormal information of the target object; The server is configured to input the CT image and the PET image into a first image processing model to obtain an abnormality mask image of the target object output by the first image processing model, wherein the first image processing model is trained according to the method of claim 14; and / or input the CT image and the PET image into a second image processing model to obtain an abnormality score of the target object output by the second image processing model, wherein the second image processing model is trained according to the method of claim 15. The server is further configured to send the abnormality mask image and / or the abnormality score to the client.
20. A computing device, comprising: a memory and a processor; the memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions, so as to implement the steps of the method of any one of claims 1 to 18.
21. A computer readable storage medium, which stores computer programs / instructions, and the computer programs / instructions are executed by a processor to implement the steps of the method of any one of claims 1 to 18.
22. A computer program product, comprising computer programs / instructions, and the computer programs / instructions are executed by a processor to implement the steps of the method of any one of claims 1 to 18.