Training method of image processing model and image processing method
By training the image processing model, using sample enhancement of the contrast and mutual learning between the image and the target image, the problem of low recognition accuracy and diagnosis rate of plain-scanned CT images in vascular embolization examination is solved, and higher recognition accuracy and diagnosis rate are achieved.
Patent Information
- Application Number
- CN202510309700.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-14
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art uses plain-scanning CT images to perform vascular embolization examination, and the recognition accuracy and diagnosis rate are low, which cannot effectively solve the needs of early diagnosis and treatment of vascular embolism.
Through the training method of the image processing model, multiple training sample pairs are obtained, including sample enhancement images and sample target images, and the initial image processing model is used to calculate the prediction annotation information and encoded feature information, adjust the model parameters until the training stop condition is reached, and the target image processing model is obtained.
The processing performance and recognition performance of the target image processing model for the target image is improved, the recognition accuracy rate when recognizing based on the target image is improved, and the diagnosis rate of vascular embolism is further improved.
Smart Images

Figure CN120163804A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and particularly to a method for training an image processing model and an image processing method. Background Art
[0002] Vascular embolism is a major factor affecting human health. Vascular embolism refers to the formation of a thrombus or other substances in a person's blood vessels, blocking the blood vessels and resulting in blood flow obstruction. Vascular embolism can include arterial embolism and venous thromboembolism, etc. In order to achieve early diagnosis and treatment of vascular embolism, professional doctors are usually required to identify vascular embolism in medical images.
[0003] In practical applications, limited by the experience of doctors, image recognition and analysis are usually carried out with the help of medical images. Vascular embolism has non-specific symptoms and can be detected by enhanced CT. However, enhanced CT exposes patients to contrast agents, which may cause allergic reactions or organ failure. In addition, due to technical and equipment limitations, enhanced CT examinations cannot be used for a long time in all regions. Therefore, plain CT can be used for vascular embolism examinations, which is more convenient and practical. However, since the sensitivity of plain CT for vascular embolism examinations is low, the recognition accuracy when identifying plain CT images is low, further resulting in a low diagnosis rate. Therefore, an effective technical solution is urgently needed to solve the above problems. Summary of the Invention
[0004] In view of this, the embodiments of this specification provide a method for training an image processing model. One or more embodiments of this specification simultaneously relate to a training device for an image processing model, an image processing method, an image processing device, a CT image processing method, a computer-aided diagnosis method for vascular embolism, a computer-aided diagnosis method for tumors, a computer-aided diagnosis system for tumors, a computing device, a computer-readable storage medium, and a computer program product to solve the technical defects existing in the prior art.
[0005] According to the first aspect of the embodiments of this specification, a method for training an image processing model is provided, including:
[0006] Obtain a plurality of training sample pairs, where each training sample pair includes a sample enhanced image and a sample target image. The sample enhanced image includes first sample annotation information for a target detection region, and the sample target image includes second sample annotation information for the target detection region;
[0007] Input the sample enhanced image and the sample target image into the initial image processing model to obtain the first predicted annotation information, the second predicted annotation information, the first encoded feature information, and the second encoded feature information output by the initial image processing model. Among them, the initial image processing model includes an enhanced image processing sub-model and a target image processing sub-model. The enhanced image processing sub-model outputs the first predicted annotation information and the first encoded feature information based on the sample enhanced image, and the target image processing sub-model outputs the second predicted annotation information and the second encoded feature information based on the sample target image;
[0008] Calculate the model loss value according to the first sample annotation information, the second sample annotation information, the first predicted annotation information, the second predicted annotation information, the first encoded feature information, and the second encoded feature information;
[0009] Adjust the model parameters of the initial image processing model according to the model loss value, and continue to train the initial image processing model until the model training stop condition is reached to obtain the target image processing model.
[0010] According to the second aspect of the embodiments of the present specification, there is provided a training device for an image processing model, including:
[0011] An acquisition module configured to acquire a plurality of training sample pairs. Among them, a training sample pair includes a sample enhanced image and a sample target image. The sample enhanced image includes first sample annotation information for a target detection area, and the sample target image includes second sample annotation information for the target detection area;
[0012] An input module configured to input the sample enhanced image and the sample target image into the initial image processing model to obtain the first predicted annotation information, the second predicted annotation information, the first encoded feature information, and the second encoded feature information output by the initial image processing model. Among them, the initial image processing model includes an enhanced image processing sub-model and a target image processing sub-model. The enhanced image processing sub-model outputs the first predicted annotation information and the first encoded feature information based on the sample enhanced image, and the target image processing sub-model outputs the second predicted annotation information and the second encoded feature information based on the sample target image;
[0013] A calculation module configured to calculate the model loss value according to the first sample annotation information, the second sample annotation information, the first predicted annotation information, the second predicted annotation information, the first encoded feature information, and the second encoded feature information;
[0014] A training module, configured to adjust model parameters of the initial image processing model according to the model loss value, and continue to train the initial image processing model until a model training stop condition is reached, to obtain a target image processing model.
[0015] According to a third aspect of the embodiments of the present specification, there is provided an image processing method, including:
[0016] Receiving an image processing task, where the image processing task carries a plurality of target images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormality in the target detection area;
[0017] Inputting the plurality of target images into a target image processing model to obtain a detection result of the target image processing model for the target detection area, where the target image processing model is trained by the image processing model training method provided in the embodiments of the present specification.
[0018] According to a fourth aspect of the embodiments of the present specification, there is provided an image processing apparatus, including:
[0019] A receiving module, configured to receive an image processing task, where the image processing task carries a plurality of target images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormality in the target detection area;
[0020] A prediction module, configured to input the plurality of target images into a target image processing model to obtain a detection result of the target image processing model for the target detection area, where the target image processing model is trained by the image processing model training method provided in the embodiments of the present specification.
[0021] According to a fifth aspect of the embodiments of the present specification, there is provided a CT image processing method, including:
[0022] Receiving a CT image processing task, where the CT image processing task carries a plurality of plain scan CT images corresponding to a target detection area, and the CT image processing task is used to detect whether there is an abnormality in the target detection area;
[0023] Inputting the plurality of plain scan CT images into a CT image processing model to obtain a detection result of the CT image processing model for the target detection area, where the CT image processing model is trained by the image processing model training method provided in the embodiments of the present specification.
[0024] According to a sixth aspect of the embodiments of the present specification, there is provided an image processing model training method, applied to a cloud-side device, including:
[0025] Obtain a plurality of training sample pairs, where each training sample pair includes a sample enhanced image and a sample target image. The sample enhanced image includes first sample annotation information for a target detection region, and the sample target image includes second sample annotation information for the target detection region;
[0026] Input the sample enhanced image and the sample target image into an initial image processing model to obtain first predicted annotation information, second predicted annotation information, first encoded feature information, and second encoded feature information output by the initial image processing model. The initial image processing model includes an enhanced image processing sub-model and a target image processing sub-model. The enhanced image processing sub-model outputs the first predicted annotation information and the first encoded feature information based on the sample enhanced image, and the target image processing sub-model outputs the second predicted annotation information and the second encoded feature information based on the sample target image;
[0027] Calculate a model loss value according to the first sample annotation information, the second sample annotation information, the first predicted annotation information, the second predicted annotation information, the first encoded feature information, and the second encoded feature information;
[0028] Adjust the model parameters of the initial image processing model according to the model loss value, and continue to train the initial image processing model until a model training stop condition is reached to obtain a target image processing model, and send the model parameters of the target image processing model to the edge device.
[0029] According to the seventh aspect of the embodiments of the present specification, there is provided an image processing method applied to a cloud device, including:
[0030] Receive an image processing task sent by an edge device, where the image processing task carries a plurality of target images corresponding to a target detection region, and the image processing task is used to detect whether there is an abnormality in the target detection region;
[0031] Input the plurality of target images into the target image processing model to obtain a detection result for the target detection region output by the target image processing model, where the target image processing model is trained by the image processing model training method provided by the embodiments of the present specification;
[0032] Send the detection result to the edge device.
[0033] According to the eighth aspect of the embodiments of the present specification, there is provided a computer-aided diagnosis method for vascular embolism, including:
[0034] Receive a vascular embolism detection task, where the vascular embolism detection task carries multiple non-contrast CT images corresponding to a vascular region, and the vascular embolism detection task is used to detect whether there is an embolism in the vascular region;
[0035] Input the multiple non-contrast CT images into a target image processing model to obtain a detection result output by the target image processing model regarding whether there is a vascular embolism in the vascular region, where the target image processing model is trained by the image processing model training method provided in the embodiments of this specification.
[0036] According to the ninth aspect of the embodiments of this specification, a computer-aided diagnosis method for tumors is provided, including:
[0037] Receive a tumor screening task, where the tumor screening task carries multiple non-contrast CT images corresponding to a target detection region, and the tumor screening task is used to detect whether there is a tumor in the target detection region;
[0038] Input the multiple non-contrast CT images into a target image processing model to obtain a detection result output by the target image processing model regarding whether there is a tumor in the target detection region, where the target image processing model is trained by the image processing model training method provided in the embodiments of this specification.
[0039] According to the tenth aspect of the embodiments of this specification, a computer-aided diagnosis system for tumors is provided, including a client and a server, where
[0040] The client is used to send a CT image processing task to the server, where the CT image processing task carries multiple non-contrast CT images corresponding to a target detection region, and the CT image processing task is used to detect whether there is an abnormality in the target detection region;
[0041] The server is used to input the multiple non-contrast CT images into a CT image processing model to obtain a detection result output by the CT image processing model regarding whether there is a tumor in the target detection region, and send the detection result to the client, where the CT image processing model is trained by the image processing model training method provided in the embodiments of this specification.
[0042] According to the eleventh aspect of the embodiments of this specification, a computing device is provided, including:
[0043] A memory and a processor;
[0044] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above method are implemented.
[0045] According to the twelfth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above method are implemented.
[0046] According to the thirteenth aspect of the embodiments of the present specification, a computer program product is provided, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above method are implemented.
[0047] An embodiment of the present specification provides a method for training an image processing model, including: obtaining a plurality of training sample pairs, where each training sample pair includes a sample enhanced image and a sample target image. The sample enhanced image includes first sample annotation information for a target detection region, and the sample target image includes second sample annotation information for the target detection region; inputting the sample enhanced image and the sample target image into an initial image processing model to obtain first predicted annotation information, second predicted annotation information, first encoded feature information, and second encoded feature information output by the initial image processing model. The initial image processing model includes an enhanced image processing sub-model and a target image processing sub-model. The enhanced image processing sub-model outputs the first predicted annotation information and the first encoded feature information based on the sample enhanced image, and the target image processing sub-model outputs the second predicted annotation information and the second encoded feature information based on the sample target image; calculating a model loss value according to the first sample annotation information, the second sample annotation information, the first predicted annotation information, the second predicted annotation information, the first encoded feature information, and the second encoded feature information; adjusting the model parameters of the initial image processing model according to the model loss value, and continuing to train the initial image processing model until a model training stop condition is reached to obtain a target image processing model.
[0048] The above method provides a framework for mutual learning between two sub-models, namely an enhanced image processing sub-model and a target image processing sub-model, unifies the classification tasks and segmentation tasks of two types of sample images, namely sample enhanced images and sample target images, processes the two types of sample images based on the two sub-models respectively, obtains the predicted annotation information and encoded feature information corresponding to the two types of sample images respectively, realizes the contrastive mutual learning between the sample enhanced images and the sample target images, trains a target image processing model based on the predicted annotation information and the encoded feature information, realizes the knowledge transfer from the enhanced image processing sub-model for processing sample enhanced images to the target image processing sub-model for processing sample target images, improves the processing performance and recognition performance of the target image processing model for target images, enhances the recognition accuracy in subsequent recognition based on target images, and further enhances the diagnosis rate. Description of the Drawings
[0049] Figure 1 is a flowchart of a method for training an image processing model provided by an embodiment of this specification;
[0050] Figure 2 is a schematic diagram of the model structure of an image processing model provided by an embodiment of this specification;
[0051] Figure 3 is a schematic diagram of the structure of a method for training an image processing model provided by an embodiment of this specification;
[0052] Figure 4 is a schematic diagram of the structure of a device for training an image processing model provided by an embodiment of this specification;
[0053] Figure 5 is a flowchart of an image processing method provided by an embodiment of this specification;
[0054] Figure 6 is a schematic diagram of the structure of an image processing device provided by an embodiment of this specification;
[0055] Figure 7 is a flowchart of a CT image processing method provided by an embodiment of this specification;
[0056] Figure 8 is a schematic diagram of a method for training an image processing model applied to a cloud-side device provided by an embodiment of this specification;
[0057] Figure 9 is a schematic flowchart of an image processing method applied to a cloud-side device provided by an embodiment of this specification;
[0058] Figure 10It is a schematic flowchart of a computer-aided diagnosis method for vascular embolization provided by an embodiment of this specification;
[0059] Figure 11 It is a schematic flowchart of a computer-aided diagnosis method for a tumor provided by an embodiment of this specification;
[0060] Figure 12 It is a schematic diagram of a computer-aided diagnosis system for a tumor provided by an embodiment of this specification;
[0061] Figure 13 It is a structural block diagram of a computing device provided by an embodiment of this specification. Detailed implementation manners
[0062] In the following description, many specific details are set forth in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar generalizations without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0063] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any or all possible combinations of one or more of the associated listed items.
[0064] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0065] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to select to authorize or refuse.
[0066] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, usually including hundreds of millions, tens of billions, hundreds of billions, trillions or even more than one quadrillion model parameters. A large model can also be called a Foundation Model. Through pre-training of the large model with a large amount of unlabeled corpus, a pre-trained model with more than one billion parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large language models (LLMs), multi-modal pre-training models, etc.
[0067] When a large model is actually applied, only a small number of samples are needed to fine-tune the pre-trained model for application to different tasks. Large models can be widely applied in the fields of natural language processing (NLP), computer vision, etc. Specifically, they can be applied to tasks in the field of computer vision such as visual question answering (VQA), image captioning (IC), image generation, etc., as well as tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, machine translation, etc. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0068] First, the noun terms involved in one or more embodiments of this specification are explained.
[0069] CT (Computed Tomography): Computed Tomography. It uses a precisely collimated X-ray beam and a highly sensitive detector to perform one cross-sectional scan after another around a certain part of the human body. It has the characteristics of fast scanning time and clear images and can be used for the examination of various diseases.
[0070] Plain CT: Also known as ordinary scanning, it refers to a scan without intravenous injection of iodine-containing contrast agent.
[0071] Contrast-enhanced CT: It refers to a method of performing a scan after injecting a contrast agent into the blood vessel. The purpose is to increase the density difference between the diseased tissue and the normal tissue to show the lesions that were not shown or were not clearly shown on plain CT. By observing whether there is enhancement and the type of enhancement, it helps in the qualitative diagnosis of the lesions.
[0072] CAD (computer aided diagnosis): Computer-aided diagnosis refers to the use of imaging, medical image processing technologies, and other possible physiological and biochemical means, combined with computer analysis and calculation, to assist in detecting lesions and improving the accuracy of diagnosis.
[0073] CTPA: Computed Tomography Pulmonary Angiography, a medical imaging technique specifically used to evaluate the pulmonary artery system, mainly for diagnosing pulmonary embolism. By injecting a contrast agent and performing a CT scan at specific time points, it can clearly show the blood flow in the pulmonary artery and identify whether there are thrombi or other abnormalities.
[0074] With the continuous development of computer technology, various learning models have gradually been applied to the prediction of various application scenarios. Correspondingly, deep learning models have also achieved remarkable success in the task of computer-aided diagnosis (CAD) of medical images. Case analysis and classification from medical images is an important topic in computer-aided diagnosis.
[0075] Pulmonary embolism is a serious life-threatening disease, second only to myocardial infarction and sudden cardiac death. Early diagnosis and treatment are crucial. Since pulmonary embolism has non-specific symptoms, the current detection mainly relies on enhanced CT examination of the pulmonary artery. However, enhanced CT exposes patients to contrast agents, which may cause allergic reactions or other discomfort symptoms. In addition, in some areas, due to technical and equipment limitations, enhanced CT cannot be popularized. In contrast, plain CT examination is more convenient and practical.
[0076] However, in plain CT images, there is almost no difference in contrast between the embolism area and the surrounding pulmonary vessels. Doctors are prone to misjudgment when using plain CT images to diagnose pulmonary embolism, which may delay treatment. Therefore, researching methods for automatic identification of pulmonary embolism based on plain CT and applying them to corresponding auxiliary diagnosis products can improve the diagnosis efficiency to a certain extent and provide a treatment window period for patients.
[0077] However, automatic identification of vascular embolism based on plain CT faces two challenges. The first is that the Hu values of the embolism area and the surrounding vessels are similar, with low contrast, making it very difficult to accurately segment the embolism area. The second is that the morphological characteristics of embolisms are diverse and their onset sites are not fixed, making it difficult to detect vascular embolisms using morphological features and location information. Therefore, an effective technical solution is urgently needed to solve the above problems.
[0078] In this specification, a method for training an image processing model is provided. This specification also relates to an apparatus for training an image processing model, an image processing method, an image processing apparatus, a CT image processing method, a computer-aided diagnosis method for vascular embolism, a computer-aided diagnosis method for tumors, a computer-aided diagnosis system for tumors, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail one by one in the following embodiments.
[0079] See Figure 1 , Figure 1 FIG. shows a flowchart of a method for training an image processing model according to an embodiment of this specification, which specifically includes the following steps.
[0080] Step 102: Obtain a plurality of training sample pairs, where each training sample pair includes a sample enhanced image and a sample target image. The sample enhanced image includes first sample annotation information for the target detection region, and the sample target image includes second sample annotation information for the target detection region.
[0081] Specifically, the method for training an image processing model provided in the embodiments of this specification can be applied to an image processing model in the field of assisted medical treatment.
[0082] In the field of assisted medical treatment, the target detection region can be understood as an organ part that needs to be detected, such as a lung region, a stomach region, etc. The sample enhanced image can be understood as a CT image obtained by performing a medical scan after injecting a contrast agent into the target detection region. For example, it can be a contrast-enhanced CT image. When the target detection region is the lung region, the sample enhanced image can be a CTPA image for the lung region. The sample target image can be understood as a CT image obtained by performing a medical scan without injecting a contrast agent into the target detection region. For example, it can be a plain CT image. The sample annotation information can be understood as abnormal annotation information for the target detection region, and the sample annotation information can include abnormal position information marked in the target detection region and / or annotation information indicating whether there is an abnormality in the target detection region, etc. Then, the first sample annotation information can be understood as abnormal annotation information marked for the target detection region in the sample enhanced image, and the second sample annotation information can be understood as abnormal annotation information marked for the target detection region in the sample target image.
[0083] Specifically, the training method of the image processing model provided in the embodiments of this specification can be supervised training, which includes multiple training sample pairs. In a training sample pair, there are multiple sample enhanced images and sample target images for the target detection region. Among them, the sample enhanced image specifically refers to an image with enhanced contrast (i.e., injected with a contrast agent) for the target detection region, and the sample target image specifically refers to a normal image (i.e., not injected with a contrast agent) for the target detection region. The first sample annotation information for the target detection region is included in the sample enhanced image, and the second sample annotation information for the target detection region is included in the sample target image.
[0084] Furthermore, in practical applications, the corresponding situation of the target detection region may be different at different time points. To ensure the consistency of the data between the sample enhanced image and the sample target image, in the method provided in the embodiments of this specification, the sample enhanced image and the sample target image are images of the same target detection region at the same time point. For example, taking the CT image of a certain patient as an example for explanation, if the target detection region is the lungs of this patient, then the training sample pair is the enhanced CT image and the plain CT image taken when this patient is examined at a certain time point. Both the enhanced CT image and the plain CT image are images of the lung region, and the two images need to be the images of this patient during the same examination. This is to prevent the situation where the abnormal information in the enhanced CT image and the plain CT image does not match after the patient has received treatment.
[0085] Specifically, when implementing, the obtaining of multiple training sample pairs includes:
[0086] Determine the sample detection object, and obtain the sample enhanced image and the sample target image corresponding to the sample detection object, where the sample enhanced image and the sample target image corresponding to the sample detection object correspond to the same target detection region of the sample detection object;
[0087] Generate the first sample annotation information for the target detection region on the sample enhanced image;
[0088] Based on the first sample annotation information, annotate the second sample annotation information on the sample target image.
[0089] Among them, the sample detection object specifically refers to the object corresponding to the target detection region. For example, when the target detection region is the stomach, the sample detection object is the organism corresponding to the target detection region, such as a human, a horse, a cow, etc. After determining the sample detection object, further obtain the sample enhanced image and the sample target image corresponding to the sample detection object. The sample enhanced image and the sample target image are images of the same target detection region.
[0090] In practical applications, the abnormal information in the sample enhanced image is more obvious than that in the sample target image. Therefore, the first sample annotation information can be generated for the target detection area in the sample enhanced image. Specifically, it can be manually annotated by technicians on the sample enhanced image, or the sample enhanced image can be input into the information annotation model, and after the annotation model recognizes the sample enhanced image, the first sample annotation information is generated.
[0091] Specifically, after determining the first sample annotation information, the sample target image can be annotated according to the first sample annotation information. The abnormal information for the abnormal area in the sample enhanced image is more obvious and can be directly annotated. However, the abnormal information for the abnormal area in the sample target image is not obvious, and it is impossible to use the annotation model or manual annotation method for annotation. At this time, since the sample enhanced image and the sample target image are taken for the same target detection area at the same time, the first sample annotation information in the sample enhanced image can be used to annotate the sample target image. For example, if the range of the abnormal area annotated in the first sample annotation information is range 1, then the sample target image can be image-annotated according to this range 1, and the range of the abnormal area in the obtained second sample annotation information is range 2, where the areas of range 1 and range 2 are the same.
[0092] In practical applications, the registration method can be used to annotate the second sample annotation information for the sample target image according to the first sample annotation information. Registration refers to the process of matching and superimposing two or more images. The purpose of registration is to find the spatial transformation relationship between two or more images so that they are spatially aligned, so as to perform further image analysis or processing.
[0093] The methods and techniques of registration vary according to different application fields and specific requirements. In medical image processing, registration techniques are used to make the corresponding points on medical images spatially consistent. In the method provided in the embodiments of this specification, the first sample annotation information is annotated in the sample enhanced image, and through the image registration method, the first sample annotation information in the sample enhanced image is migrated to the sample target image, and the second sample annotation information is generated in the sample target image.
[0094] Further, generating the first sample annotation information for the target detection area on the sample enhanced image includes:
[0095] Inputting the sample enhanced image into the information annotation model to obtain the first sample segmentation information and the first sample classification information for the target detection area output by the information annotation model;
[0096] Annotating the second sample annotation information on the sample target image based on the first sample annotation information includes:
[0097] Mark the second sample segmentation information on the sample target image according to the first sample segmentation information, and determine the second sample classification information of the sample target image according to the first sample classification information.
[0098] In a specific embodiment provided in this specification, the sample enhanced image can be input into an information annotation model, and the information annotation model performs information annotation on the sample enhanced image. Specifically, the information annotation model marks the first sample segmentation information in the sample enhanced image and labels the first sample classification information for the sample enhanced image. Among them, the first sample segmentation information specifically refers to the segmentation mask information for the abnormal region in the target detection region, and the first sample classification information specifically refers to the abnormal classification information for the target detection region.
[0099] Correspondingly, after determining the first sample segmentation information and the first sample classification information of the sample enhanced image, the second sample segmentation information can be marked on the sample target image according to the first sample segmentation information, and the second sample classification information of the sample target image can be determined according to the first sample classification information.
[0100] Among them, the first sample segmentation information and the first sample classification information can be understood as the first sample annotation information of the sample enhanced image; the second sample segmentation information and the second sample classification information can be understood as the second sample annotation information of the sample target image. The specific contents of the first sample classification information and the second sample classification information are the same.
[0101] For example, in the field of assisted medical treatment, when the target detection region is the lung region, the sample segmentation information can be understood as the segmentation mask information for pulmonary embolism in the target detection region, and the sample classification information can be understood as the probability information for whether it is a pulmonary embolism case or not in the target detection region. Then, it can be understood that the first sample segmentation information can be understood as the sample segmentation information of the target detection region in the sample enhanced image, the second sample segmentation information can be understood as the sample segmentation information of the target detection region in the sample target image, the first sample classification information can be understood as the sample classification information of the target detection region in the sample enhanced image, and the second sample classification information can be understood as the sample classification information of the target detection region in the sample target image.
[0102] Step 104: Input the sample enhanced image and the sample target image into the initial image processing model to obtain the first predicted annotation information, the second predicted annotation information, the first encoded feature information, and the second encoded feature information output by the initial image processing model. The initial image processing model includes an enhanced image processing sub-model and a target image processing sub-model. The enhanced image processing sub-model outputs the first predicted annotation information and the first encoded feature information based on the sample enhanced image, and the target image processing sub-model outputs the second predicted annotation information and the second encoded feature information based on the sample target image.
[0103] Among them, the initial image processing model can be understood as an image processing model that has not undergone model training. In practical applications, the initial image processing model includes two parts, namely an enhanced image processing sub-model and a target image processing sub-model. The model structures of the two sub-models are the same and are used to process the sample enhanced image and the sample target image respectively. Specifically, the enhanced image processing sub-model is used to process the sample enhanced image, and the target image processing sub-model is used to process the sample target image. The first predicted annotation information can be understood as the annotation information predicted by the initial image processing model for the sample enhanced image, the second predicted annotation information can be understood as the annotation information predicted by the initial image processing model for the sample target image, the first encoded feature information can be understood as the encoded feature information obtained by the initial image processing model for processing the sample enhanced image, and the second encoded feature information can be understood as the encoded feature information obtained by the initial image processing model for processing the sample target image.
[0104] See Figure 2 , Figure 2 which shows the schematic diagram of the model structure of the image processing model provided by an embodiment of this specification. As Figure 2 shown, the initial image processing model includes two sub-models, namely an enhanced image processing sub-model and a target image processing sub-model. Each sub-model includes an encoder, a decoder, and a classifier. The encoder is used to extract the encoded feature information of the image, the decoder is used to generate image segmentation information, and the classifier is used to generate image classification information. The image segmentation information and the image classification information together form the annotation information. Further, the decoder and the classifier in the enhanced image processing sub-model output the first predicted annotation information, and the decoder and the classifier in the target image processing sub-model output the second predicted annotation information.
[0105] Specifically, when implementing, the step of inputting the sample enhanced image and the sample target image into the initial image processing model to obtain the first predicted annotation information, the second predicted annotation information, the first encoded feature information, and the second encoded feature information includes:
[0106] Input the sample enhanced image into the enhanced image processing sub-model to obtain first predicted annotation information and first encoded feature information;
[0107] Input the sample target image into the target image processing sub-model to obtain second predicted annotation information and second encoded feature information.
[0108] Specifically, the sample enhanced image can be input into the enhanced image processing sub-model in the initial image processing model, and the enhanced image processing sub-model is used to process the sample enhanced image to obtain first predicted annotation information and first encoded feature information; the sample target image is input into the target image processing sub-model in the initial image processing model, and the target image processing sub-model is used to process the sample target image to obtain second predicted annotation information and second encoded feature information.
[0109] In summary, by using the enhanced image processing sub-model and the target image processing sub-model to process the sample enhanced image and the sample target image respectively, the contrastive mutual learning between the sample enhanced image and the sample target image is realized, and the phase privilege information unique to the enhancement based on the contrast agent in the sample enhanced image is transferred to the phase of the target image processing sub-model, improving the recognition accuracy of the image processing model in the phase of the sample target image.
[0110] In practical applications, the enhanced image processing sub-model includes a first encoder, a first decoder, and a first classifier;
[0111] The step of inputting the sample enhanced image into the enhanced image processing sub-model to obtain first predicted annotation information and first encoded feature information includes:
[0112] Input the sample enhanced image into the first encoder to obtain at least one enhanced encoded feature information, where at least one enhanced encoded feature information includes classification enhanced encoded feature information;
[0113] Input the classification enhanced encoded feature information into the first classifier to obtain first predicted classification information;
[0114] Input each enhanced encoded feature information into the first decoder to obtain first predicted segmentation information;
[0115] Determine the first encoded feature information from the at least one enhanced encoded feature information.
[0116] The enhanced image processing sub-model may include a first encoder, a first decoder and a first classifier, wherein the first encoder is used to extract enhanced coding feature information of the sample enhanced image, the first decoder is used to predict segmentation information of the sample enhanced image, and the first classifier is used to predict classification information of the sample enhanced image. The first predicted annotation information of the sample enhanced image may include first predicted classification information and first predicted segmentation information.
[0117] Specifically, combined with the above Figure 2 , the sample enhanced image can be input into the first encoder, and after being processed by the first encoder, at least one enhanced coding feature information is obtained. In practical applications, the first encoder includes multiple sequentially connected coding layers, and each coding layer outputs enhanced coding feature information. For example, taking the first encoder having 5 coding layers as an example, the sample enhanced image is input into the first encoder, the first coding layer encodes the sample enhanced image and outputs the first enhanced coding feature information, the first enhanced coding feature information is input into the second coding layer, and after encoding, the second enhanced coding feature information is output, and so on, the fourth enhanced coding feature information output by the fourth coding layer is input into the fifth coding layer, and after encoding, the fifth enhanced coding feature information is obtained. At this point, the first encoder outputs 5 enhanced coding feature information, and these 5 enhanced coding feature information are the at least one enhanced coding feature information obtained above, and the enhanced coding feature information output by the last coding layer is the classification enhanced coding feature information.
[0118] Afterwards, the classified enhanced coding feature information can be input into the first classifier, and the first classifier classifies the information through the category feature information in the classified enhanced coding feature information to obtain the first predicted classification information; and at least one enhanced coding feature information output by the first encoder is input into the first decoder, and the at least one enhanced coding feature information is decoded in the first decoder to obtain the first predicted segmentation information.
[0119] Furthermore, in order to reflect the coding difference between the enhanced image processing submodel and the target image processing submodel during the coding process, the first coding feature information is determined in at least one enhanced coding feature information, and further, the first coding feature information can be determined in at least one enhanced coding feature information according to a preset ratio. For example, the first 70% of the enhanced coding feature information is selected as the first coding feature information, and when there are 5 enhanced coding feature information, the first 3 or 4 enhanced coding feature information are selected as the first coding feature information according to the coding order of the enhanced coding feature information.
[0120] In practical applications, in the first encoder, the feature vector dimensions of the enhanced encoded feature information output by each encoding layer are different. Generally, the feature vector dimension of the enhanced encoded feature information output by the previous encoding layer is greater than that of the enhanced encoded feature information output by the subsequent encoding layer. Continuing with the above example, the feature vector dimension of the first enhanced encoded feature information output by the first encoding layer is greater than that of the second enhanced encoded feature information output by the second encoding layer. Based on this, it is also possible to select, from at least one enhanced encoded feature information, the enhanced encoded feature information whose feature vector dimension is greater than a preset dimension threshold as the first encoded feature information. For example, the enhanced encoded feature information output by the first two encoding layers can be determined as the first encoded feature information, or the enhanced encoded feature information output by the second and third encoding layers can also be determined as the first encoded feature information. The embodiments of this specification do not limit this.
[0121] In summary, by using the enhanced image processing sub-model, the prediction of the classification information and segmentation information of the enhanced sample image and the determination of the encoded feature information are realized, which is convenient for subsequent contrast learning.
[0122] Correspondingly, the target image processing sub-model includes a second encoder, a second decoder, and a second classifier;
[0123] Inputting the sample target image into the target image processing sub-model to obtain second prediction annotation information and second encoded feature information includes:
[0124] Inputting the sample target image into the second encoder to obtain at least one target encoded feature information, where at least one target encoded feature information includes classification target encoded feature information;
[0125] Inputting the classification target encoded feature information into the second classifier to obtain second prediction classification information;
[0126] Inputting each target encoded feature information into the second decoder to obtain second prediction segmentation information;
[0127] Determining second encoded feature information from the at least one target encoded feature information.
[0128] Specifically, combining the above Figure 2 , the target image processing sub-model includes a second encoder, a second decoder, and a second classifier. The model structure of the target image processing sub-model is the same as that of the enhanced image processing sub-model. Regarding the processing method of the sample target image by the target image processing sub-model, reference can be made to the processing method of the sample enhanced image by the enhanced image processing sub-model above, and details will not be repeated here.
[0129] Based on this, in the initial image processing model, the first predicted classification information, the first predicted segmentation information, and the first encoded feature information are obtained in the enhanced image processing sub-model branch; the second predicted classification information, the second predicted segmentation information, and the second encoded feature information are obtained in the target image processing sub-model branch.
[0130] In practical applications, the inputting of the sample enhanced image and the sample target image into the initial image processing model includes:
[0131] Cropping the sample enhanced image and the sample target image based on a preset cropping size;
[0132] Inputting the cropped sample enhanced image and the sample target image into the initial image processing model.
[0133] Specifically, usually the abnormal area in the target detection area is smaller than the target detection area. If the sample enhanced image and the sample target image are input into the initial image processing model for processing, there will be a problem of inaccurate recognition. Based on this, before inputting the sample enhanced image and the sample target image into the initial image processing model, the sample enhanced image and the sample target image are first cropped according to the preset cropping size. The cropped sample enhanced image and the sample target image are input into the initial image processing model for processing. This enables the initial image processing model to better identify the abnormal information in the target detection area.
[0134] In practical applications, the preset cropping size can be 224×224×96. Based on this, the sample enhanced image and the sample target image are respectively randomly cropped into three-dimensional images with an image size of 224×224×96 (i.e., the cropped sample enhanced image and the sample target image), and then the cropped sample enhanced image and the sample target image are input into the initial image processing model to perform the above image processing and prediction in the initial image processing model.
[0135] Step 106: Calculate the model loss value according to the first sample annotation information, the second sample annotation information, the first predicted annotation information, the second predicted annotation information, the first encoded feature information, and the second encoded feature information.
[0136] Among them, the first sample annotation information includes the first sample segmentation information and the first sample classification information, the first predicted annotation information includes the first predicted segmentation information and the first predicted classification information, the second sample annotation information includes the second sample segmentation information and the second sample classification information, and the second predicted annotation information includes the second predicted segmentation information and the second predicted classification information.
[0137] Specifically, calculating the model loss value according to the first sample annotation information, the second sample annotation information, the first predicted annotation information, the second predicted annotation information, the first encoded feature information, and the second encoded feature information includes:
[0138] Calculating a first model loss value according to the first sample annotation information and the first predicted annotation information;
[0139] Calculating a second model loss value according to the second sample annotation information and the second predicted annotation information;
[0140] Calculating an information difference loss value according to the first predicted annotation information and the second predicted annotation information;
[0141] Calculating a feature information comparison loss value according to the first encoded feature information and the second encoded feature information;
[0142] Calculating the model loss value according to the first model loss value, the second model loss value, the information difference loss value, and the feature information comparison loss value.
[0143] Among them, the first model loss value can be understood as the model loss value of the enhanced image processing sub-model itself, the second model loss value can be understood as the model loss value of the target image processing sub-model itself, and the information difference loss value and the feature information comparison loss value can be understood as the model loss value between the enhanced image processing sub-model and the target image processing sub-model.
[0144] Specifically, a first segmentation loss value can be calculated according to the first sample segmentation information and the first predicted segmentation information, a first classification loss value can be calculated according to the first sample classification information and the first predicted classification information, and a first model loss value can be calculated according to the first segmentation loss value and the first classification loss value; a second segmentation loss value can be calculated according to the second sample segmentation information and the second predicted segmentation information, a second classification loss value can be calculated according to the second sample classification information and the second predicted classification information, and a second model loss value can be calculated according to the second segmentation loss value and the second classification loss value; an information difference loss value can be calculated according to the first predicted annotation information and the second predicted annotation information; a feature information comparison loss value can be calculated according to the first encoded feature information and the second encoded feature information, and the model loss value can be calculated according to the first model loss value, the second model loss value, the information difference loss value, and the feature information comparison loss value.
[0145] In the specific embodiments provided in this specification, since the enhanced image processing sub-model processes the sample enhanced image, abnormal information in the target detection area can be identified through the sample enhanced image. However, the target image processing sub-model processes the sample target image and cannot accurately identify the abnormal information in the sample target image. Therefore, in the embodiments provided in this specification, it is desired that the target image processing sub-model can learn the feature information in the recognition process of the enhanced image processing sub-model. Therefore, in the process of the initial image processing model recognizing the sample enhanced image and the sample target image, in addition to generating the first predicted annotation information corresponding to the sample enhanced image and the second predicted annotation information corresponding to the sample target image, it can further generate a classification difference loss value, a segmentation difference loss value, and a feature information comparison loss value.
[0146] Specifically, calculating the feature information comparison loss value according to the first encoded feature information and the second encoded feature information includes:
[0147] Performing normalization processing on the first encoded feature information and the second encoded feature information to obtain first embedding vector information and second embedding vector information;
[0148] Calculating similarity information according to the first embedding vector information and the second embedding vector information;
[0149] Calculating the feature information comparison loss value according to the similarity information.
[0150] Among them, the similarity information can be understood as a similarity matrix.
[0151] Specifically, in order to enhance the feature alignment between the bimodal images (i.e., the sample enhanced image and the sample target image), based on the multi-modal contrast learning method, the embedding vectors of similar sample images in different modalities can be made closer, and at the same time, the separation degree between different sample images can be increased. Specifically, the first encoded feature information can be processed by normalization and adaptive average pooling to obtain the first embedding vector information, the second encoded feature information can be processed by normalization and adaptive average pooling to obtain the second embedding vector information, and the similarity matrix can be calculated according to the first embedding vector information, the second embedding vector information, and the temperature parameter, and the feature information comparison loss value can be calculated according to the similarity matrix, and the first predicted segmentation information and the second predicted segmentation information corresponding to the sample enhanced image and the sample target image.
[0152] For example, for training sample pair A, training sample pair B, and training sample pair C, by comparing the loss value based on the feature information, the embedding vectors between the sample enhanced image and the sample target image in a single training sample pair become closer. For example, the embedding vectors between the sample enhanced image A and the sample target image A in training sample pair A become closer, while the embedding vectors between different training sample pairs are made farther apart, thereby increasing the separation degree between different sample images. For example, the embedding vectors between the sample enhanced image A in training sample pair A and the sample enhanced image B in training sample pair B become farther apart, and the embedding vectors between the sample enhanced image A in training sample pair A and the sample target image B in training sample pair B become farther apart, etc.
[0153] In practical applications, the formula for calculating the feature information comparison loss value is as shown in formula (1) below.
[0154]
[0155] Among them, i and j are the sample enhanced image and the sample target image in the training sample pair, and M ij is the sample mask corresponding to the training sample pair (i.e., the first predicted segmentation information and the second predicted segmentation information). If the sample mask is equal to 1, it indicates that the training sample pair corresponding to the sample mask is a positive sample pair (i.e., there is an anomaly). If the sample mask is equal to 0, it indicates that the training sample pair corresponding to the sample mask is a negative sample pair (i.e., there is no anomaly). is the feature information comparison loss value, the similarity matrix is is the first embedding vector information, is the second embedding vector information, is the temperature parameter, T is the exponent of the temperature parameter, N is the number of samples in a training batch (i.e., the number of training sample pairs), and τ is the temperature parameter.
[0156] In summary, through the multi-modal learning method, a feature-based contrast learning strategy is proposed to promote the embedding of the same sample in different modalities to be closer, while increasing the separation degree of different sample embeddings and improving the effect of cross-modal embeddings.
[0157] In specific implementation, the first predicted annotation information includes first predicted classification information and first predicted segmentation information, and the second predicted annotation information includes second predicted classification information and second predicted segmentation information;
[0158] Calculating the information difference loss value according to the first predicted annotation information and the second predicted annotation information includes:
[0159] Calculating the classification difference loss value according to the first predicted classification information and the second predicted classification information;
[0160] Calculate the segmentation difference loss value according to the first predicted segmentation information and the second predicted segmentation information;
[0161] Determine the information difference loss value according to the classification difference loss value and the segmentation difference loss value.
[0162] Specifically, calculate the classification difference loss value according to the first predicted classification information and the second predicted classification information. In the method provided in the embodiments of this specification, KL divergence is used as a measure of the difference between the two. See the following formula (2).
[0163]
[0164] Among them, L KL (p2|p1) represents the classification difference loss value, p1 represents the first predicted classification information, and p2 represents the second predicted classification information. m represents the classification result, m = 0 indicates no anomaly, and m = 1 indicates an anomaly. x i represents the i-th sample pair.
[0165] Calculate the segmentation difference loss value according to the first predicted segmentation information and the second predicted segmentation information. The segmentation difference loss value is used to improve the ability to distinguish discriminative features between the background and the abnormal region in the segmentation scenario. In the method provided in the embodiments of this specification, an intra-class feature differentiation strategy based on a dense center loss function is provided. In each training iteration, the center is calculated as the centroid feature of the pixels belonging to the corresponding class in the segmentation mask. For each center loss function, see the following formula (3):
[0166]
[0167] Among them, L disc represents the segmentation difference loss value, x k is the feature of the pixel belonging to class k, c k represents the center of the k-th class of the depth feature. This way enables the network to more effectively learn a compact and independent cluster corresponding to each class in the feature space.
[0168] Specifically, when calculating the model loss value according to the first model loss value, the second model loss value, the information difference loss value, and the feature information comparison loss value, the model loss value can be calculated according to the following formula (4).
[0169] L Total = L clas + L seg + λ1L KL + λ2L cfl + λ3L disc (4)
[0170] Among them, LTotal It can be understood as the model loss value, L clas It can be understood as the first classification loss value and the second classification loss value. The first classification loss value and the second classification loss value can be binary cross-entropy losses, L seg It can be understood as the first segmentation loss value and the second segmentation loss value. The first segmentation loss value and the second segmentation loss value can be optimized by focal loss to solve the class imbalance ratio between the abnormal region and the background. λ1, λ2, and λ3 are the weights of each loss value. In an embodiment of this specification, the weights can all be set to 0.4 to make the ranges of each loss value equivalent.
[0171] In practical applications, data augmentation can include random flipping, rotation, spatial padding, and random cropping to a unified size with a probability of 30%. In the inference stage, sliding window inference can be used. The window size and the preset cropping size can be the same, both being 224×224×96, and the window sliding overlap ratio is 50%. The central patch is cropped to the same size as the input of the classifier. The Adam optimizer and the Cosine Annealing learning rate scheduler are paired for use. The initial learning rate is 0.001, and the learning rate is adjusted strategically during the training process. The minimum learning rate is set to 0.0001.
[0172] In summary, by the differences between the first predicted classification information and the second predicted classification, the differences between the first predicted segmentation information and the second predicted segmentation information, and the differences between the first encoded feature information and the second encoded feature information, the differences between the enhanced image processing sub-model and the target image processing sub-model in the image recognition process are determined. Thus, it is convenient to better train the initial image processing model, enabling the target image processing model to better recognize the target image.
[0173] Step 108: Adjust the model parameters of the initial image processing model according to the model loss value, and continue to train the initial image processing model until the model training stop condition is reached to obtain the target image processing model.
[0174] Specifically, according to the model loss value, the model parameters of the enhanced image processing sub-model and the target image processing sub-model in the initial image processing model can be adjusted. After that, the above steps can be repeated to continue training the initial image processing model until the model training stop condition is reached to obtain the target image processing model.
[0175] In practical applications, the model training stop conditions of the initial image processing model include:
[0176] The total loss value is less than the preset threshold, and / or the number of training rounds reaches the preset number of training rounds.
[0177] Specifically, during the training process of the initial image processing model, the training stop condition of the model can be set to that the model loss value is less than a preset threshold. In practical applications, when the model loss value is less than the preset threshold, there is no need to adjust the model parameters of the initial image processing model. Further,
[0178] Furthermore, the training stop condition of the initial image processing model can be further set to that the number of training rounds reaches a preset number of training rounds. For example, if the preset number of training rounds is 10 rounds, then when the number of training rounds of the model reaches 10 rounds, the model training stop condition is reached.
[0179] In practical applications, adjusting the model parameters of the initial image processing model according to the model loss value and continuing to train the initial image processing model until the model training stop condition is reached to obtain the target image processing model includes:
[0180] Based on the gradient descent algorithm, adjusting the model parameters of the initial image processing model according to the model loss value and continuing to train the initial image processing model until the model training stop condition is reached to obtain the target image processing model.
[0181] In practical applications, the gradient descent algorithm can be understood as a multi-task collaborative optimization method based on Jacobian descent. Among them, Jacobian-based gradient descent is a technique for optimizing multivariate functions. The Jacobian matrix is a matrix composed of the first-order partial derivatives of a multivariate function. In order to optimize the multi-task objective, the gradient conflict between different tasks can be processed based on a conflict-free gradient projection aggregator. Different tasks can be understood as the classification task, segmentation task, and the optimization tasks corresponding to the above formulas (1), (2), and (3) proposed in the embodiments of this specification. The gradient is updated as shown in the following formula (5) during the model training process.
[0182]
[0183] Among them, η represents the learning rate, represents based on the conflict-free gradient projection aggregator, represents the Jacobian gradient descent, x t represents the parameters of the model after gradient update, x t-1 represents the parameters of the model before gradient update.
[0184] Specifically, the conflict-free gradient projection aggregator can project the gradient of each task onto the dual cone of the Jacobian row and take the average to ensure conflict-free updates. It improves the coordination and efficiency of the optimization process.
[0185] Further, adjusting the model parameters of the initial image processing model according to the model loss value, and continuing to train the initial image processing model until a model training stop condition is reached to obtain a target image processing model, includes:
[0186] Adjusting the model parameters of the initial image processing model according to the model loss value, and continuing to train the initial image processing model until a model training stop condition is reached to obtain a reference image processing model;
[0187] Generating a target image processing model based on the target image processing sub-model in the reference image processing model.
[0188] At this time, the reference image processing model is not the final image processing model to be trained. In the method provided in this specification, ultimately, it is desired that the obtained image processing model can process the target image. The model structure of the reference image processing model is the same as that of the initial image processing model, and it also includes an enhanced image processing sub-model and a target image processing sub-model. In order to enable the final image processing model to only process the target image, the target image processing model can be obtained according to the target image processing sub-model in the reference image processing model. That is, the enhanced image processing sub-model in the reference image processing model is deleted, and the target image processing sub-model in the reference image processing model is retained to obtain the final target image processing model.
[0189] In summary, the above method provides a framework for mutual learning between two sub-models, namely an enhanced image processing sub-model and a target image processing sub-model, unifies the classification tasks and segmentation tasks of two types of sample images, namely sample enhanced images and sample target images. Based on the two sub-models, the two types of sample images are processed respectively to obtain the predicted annotation information and encoded feature information corresponding to the two types of sample images respectively, realizes the contrastive mutual learning between the sample enhanced image and the sample target image, and trains the target image processing model based on the predicted annotation information and the encoded feature information, realizes the knowledge transfer from the enhanced image processing sub-model for processing the sample enhanced image to the target image processing sub-model for processing the sample target image, improves the processing performance and recognition performance of the target image processing model for the target image, enhances the recognition accuracy rate when recognizing based on the target image subsequently, and further enhances the diagnosis rate.
[0190] The following combines the attached Figure 3 , taking the application of the training method of the image processing model provided in this specification in pulmonary embolism as an example, to further illustrate the training method of the image processing model. Among them, Figure 3 shows a schematic structural diagram of a training method of an image processing model provided by an embodiment of this specification.
[0191] In this embodiment, the model training of the image detection model corresponding to pulmonary artery embolism will be explained as an example. First, a plurality of pairs of training samples that meet the image quality requirements are collected from a hospital. The pairs of training samples include pulmonary artery enhanced CT images and plain scan CT images. The enhanced CT images are input into the image labeling model for recognition to obtain a voxel-level pulmonary artery embolism segmentation mask and classification label. Based on the pulmonary artery embolism segmentation mask and classification label of the enhanced CT images, the plain scan CT images are labeled to obtain the voxel-level pulmonary artery embolism segmentation mask and classification label on the plain scan CT images.
[0192] Voxel level refers to the level of detail or fineness of the basic unit considered or operated on in three-dimensional data processing, analysis, or imaging. A voxel is short for volume element and is the smallest unit in three-dimensional space segmentation. Solids containing voxels can be represented by volume rendering or by extracting polygon isosurfaces of a given threshold contour. Voxels are used in fields such as three-dimensional imaging, scientific data, and medical imaging. Conceptually, it is similar to the smallest unit pixel in two-dimensional space.
[0193] Use the lungmask python library to extract the maximum circumscribed three-dimensional matrix containing the lung region in the enhanced CT images and plain scan CT images. The pairs of training samples are divided into a training set and a test set in a ratio of 8:2, ensuring that the ratio of positive and negative cases in the training set and the test set remains the same. The lungmask python is a python library mainly used in the field of medical image processing, especially for the automatic segmentation of the lung region in chest CT images.
[0194] Use an encoder based on xLSTM as the backbone network and a decoder with U-net as the segmentation framework to build a pulmonary artery embolism CT image classification and segmentation framework for multi-task learning based on mutual learning, combined with the above Figure 3 , including an enhanced CT path network (i.e., an enhanced image processing sub-model) and a plain scan CT path network (i.e., a target image processing sub-model). In the method provided in the embodiments of this specification, the proposed cross-period mutual learning framework aims to simultaneously train two convolutional neural networks with different tasks (classification and segmentation). The enhanced CT path network is responsible for classifying and segmenting the enhanced CT images, and the plain scan CT path network is responsible for classifying and segmenting the plain scan CT images.
[0195] xLSTM is an extension or improvement of the traditional LSTM (Long Short-Term Memory) network. LSTM is a special type of recurrent neural network (RNN) designed to address the vanishing gradient and exploding gradient problems encountered by traditional RNNs when dealing with long sequences, thus better capturing long-term dependencies. Among them, the xLSTM architecture can be understood as an xLSTM architecture constructed by integrating sLSTM and mLSTM into a residual block. Here, sLSTM can refer to the introduction of exponential gating and a new storage mixing technique based on LSTM, allowing LSTM to revise its storage decisions. mLSTM can refer to the expansion of the memory unit of LSTM from a scalar to a matrix, improving the storage capacity, and introducing a covariance update rule, enabling mLSTM to be fully parallelized.
[0196] In the framework provided in this specification, whether it is the enhanced CT pathway network or the non-enhanced CT pathway network, the 3D version of xLSTM is used as the backbone network to obtain features at different scales, and the decoder of the 3D version of U-net is selected as the decoder for the segmentation task. A classifier is added after the backbone network for the classification task. The architecture of the classifier can include an adaptive three-dimensional average pooling layer, a fully connected layer of size 64×32, a relu activation layer, and a fully connected layer of size 32×2. The architectures of the enhanced CT pathway network and the non-enhanced CT pathway network are the same, and there is no distinction between the enhanced CT pathway network and the non-enhanced CT pathway network in the subsequent processing of the pathway network.
[0197] In the application, a 224×224×96-dimensional segmented three-dimensional image is randomly cropped from the input 3D CT image (which can be an enhanced CT image or a non-enhanced CT image). Selecting the size of 224×224×96 can ensure that each 3D image includes the mask area of pulmonary embolism, enabling the model to identify the pulmonary embolism area, and at the same time reducing the volume of the image input into the model and the data processing volume of the image processing model.
[0198] The segmented three-dimensional image is input into the encoder of the pathway network (the enhanced CT image is input into the first encoder, and the non-enhanced CT image is input into the second encoder) to obtain at least one encoded feature information E = {E1, E2, E3, E4, E5, E6} corresponding to the segmented three-dimensional image. E1 - E6 are the encoded feature information output by each encoding layer in the encoder (the enhanced CT image is the enhanced encoded feature information, and the non-enhanced CT image is the target encoded feature information). Among them, E6 is the encoded feature information output by the last encoding layer. E6 can also be understood as the classification encoded feature information.
[0199] Input E6 into the classifier for processing. The logits obtained from the classifier represent the probability that the segmented three-dimensional image is an embolism case or a normal case, that is, the predicted classification information. Align the logits obtained from the enhanced CT pathway network and the plain CT pathway network respectively for the prediction distribution, use the KL divergence as a measure of the difference between the logits, and calculate the classification difference loss value through the above formula (2).
[0200] Input each encoded feature information E into the decoder for decoding processing to obtain the segmented output mask It means the probability that each pixel point in the image is a pulmonary embolism. In order to improve the ability of the segmentation network to distinguish the discriminative features between the background and pulmonary embolism, in the method provided in this specification, an intra-class feature differentiation strategy based on the dense center loss function is proposed. In each training iteration, the center is calculated as the centroid feature of the pixels belonging to the corresponding class in the segmentation mask, and the segmentation difference loss value is calculated through the above formula (3).
[0201] In addition, in order to enhance the feature alignment between the bimodal images, based on the multi-modal contrast learning method, similar samples are encouraged to have closer embedding vectors in different modalities, while increasing the separation between different samples. Specifically, normalized embedding vectors can be obtained from at least one encoded feature information output by the encoder. For example, E1 and E2 are selected from at least one encoded feature information E = {E1, E2, E3, E4, E5, E6} for calculation. In the method provided in this specification, during the process of calculating the feature information contrast loss value, the contrast loss is calculated by calculating the similarity matrix, and the positive sample pairs are selectively included through the positive sample mask M, and the feature information contrast loss value is calculated through the above formula (1).
[0202] On this basis, each predicted segmentation information and predicted classification information will be compared with the annotation information on the sample image to calculate the model loss value of the enhanced CT pathway network and the model loss value of the plain CT pathway network.
[0203] Through the above formula (4), adjust the model parameters of the enhanced CT pathway network and the plain CT pathway network, and continue training until the model training stop condition is reached to obtain the reference image processing model corresponding to the initial image processing model.
[0204] Through the above formula (5), perform model gradient update based on the gradient descent algorithm during the model training process to avoid gradient conflict.
[0205] On the basis of obtaining a reference image processing model, the enhanced CT path network in the reference image processing model is removed, and only the plain CT path network is retained. A target image processing model for assisting in the diagnosis and treatment of pulmonary artery embolism through plain CT images is obtained.
[0206] In the method provided in this specification, after obtaining the target image processing model, for the binary classification task, the model performance of the target image processing model can be evaluated using the area under the receiver operating characteristic curve, sensitivity, and specificity. For the segmentation task, the dice coefficient is used to evaluate the model performance. Three doctors with cardiopulmonary imaging experience are invited to compare with the target image processing model. The 289 plain CT images from the test set are given to the three doctors respectively. The doctors are asked to judge whether there is pulmonary artery embolism based on the CT images without obtaining any patient information, and it is informed that the incidence of pulmonary embolism cases in the dataset may be higher than the typical prevalence in routine screening, but the specific case type distribution is not disclosed to the doctors. These 289 plain CT images are respectively input into the image processing model for pulmonary artery embolism annotation recognition. During the test, it is required that the accuracy of the target image processing model is greater than 0.8. If this requirement is met, the training and optimization of the target image processing model are terminated. If not, the model parameters of the target image processing model are adjusted according to the above steps until this requirement is met.
[0207] After testing, the recognition results of the target image processing model for the plain CT in the test set are higher than the conclusions given by the three doctors. Thus, the training of the target image processing model with the ability to recognize plain CT images and obtain the diagnosis and treatment results of pulmonary artery embolism in plain CT images is successful.
[0208] The embodiment of this specification provides a framework for mutual learning between two sub-models, namely an enhanced image processing sub-model and a target image processing sub-model, which unifies the classification task and segmentation task of two types of sample images, namely sample enhanced images and sample target images. Based on the two sub-models, the two types of sample images are processed respectively to obtain the predicted annotation information and encoded feature information corresponding to the two types of sample images respectively, realizing the contrastive mutual learning between the sample enhanced images and the sample target images. The target image processing model is trained based on the predicted annotation information and encoded feature information, realizing the knowledge transfer from the enhanced image processing sub-model's processing of sample enhanced images to the target image processing sub-model's processing of sample target images, improving the processing performance and recognition performance of the target image processing model for target images, enhancing the recognition accuracy when recognizing based on target images subsequently, and further enhancing the diagnosis rate.
[0209] Corresponding to the above method embodiment, this specification also provides an embodiment of a training device for an image processing model. Figure 4The figure shows a schematic structural diagram of a training device for an image processing model provided by an embodiment of this specification. As Figure 4 shown, the device includes:
[0210] An acquisition module 402, configured to acquire a plurality of training sample pairs, where a training sample pair includes a sample enhanced image and a sample target image, the sample enhanced image includes first sample annotation information for a target detection region, and the sample target image includes second sample annotation information for the target detection region;
[0211] An input module 404, configured to input the sample enhanced image and the sample target image into an initial image processing model, and obtain first predicted annotation information, second predicted annotation information, first encoded feature information, and second encoded feature information output by the initial image processing model, where the initial image processing model includes an enhanced image processing sub-model and a target image processing sub-model, the enhanced image processing sub-model outputs the first predicted annotation information and the first encoded feature information based on the sample enhanced image, and the target image processing sub-model outputs the second predicted annotation information and the second encoded feature information based on the sample target image;
[0212] A calculation module 406, configured to calculate a model loss value according to the first sample annotation information, the second sample annotation information, the first predicted annotation information, the second predicted annotation information, the first encoded feature information, and the second encoded feature information;
[0213] A training module 408, configured to adjust model parameters of the initial image processing model according to the model loss value, and continue to train the initial image processing model until a model training stop condition is reached, to obtain a target image processing model.
[0214] In an optional embodiment, the calculation module 406 is further configured to:
[0215] Calculate a first model loss value according to the first sample annotation information and the first predicted annotation information;
[0216] Calculate a second model loss value according to the second sample annotation information and the second predicted annotation information;
[0217] Calculate an information difference loss value according to the first predicted annotation information and the second predicted annotation information;
[0218] Calculate a feature information comparison loss value according to the first encoded feature information and the second encoded feature information;
[0219] Calculate the model loss value according to the first model loss value, the second model loss value, the information difference loss value, and the feature information comparison loss value.
[0220] In an optional embodiment, the calculation module 406 is further configured to:
[0221] Normalize the first encoded feature information and the second encoded feature information to obtain first embedding vector information and second embedding vector information;
[0222] Calculate similarity information according to the first embedding vector information and the second embedding vector information;
[0223] Calculate the feature information comparison loss value according to the similarity information.
[0224] In an optional embodiment, the first predicted annotation information includes first predicted classification information and first predicted segmentation information, and the second predicted annotation information includes second predicted classification information and second predicted segmentation information;
[0225] The calculation module 406 is further configured to:
[0226] Calculate a classification difference loss value according to the first predicted classification information and the second predicted classification information;
[0227] Calculate a segmentation difference loss value according to the first predicted segmentation information and the second predicted segmentation information;
[0228] Determine the information difference loss value according to the classification difference loss value and the segmentation difference loss value.
[0229] In an optional embodiment, the training module 408 is further configured to:
[0230] Based on the gradient descent algorithm, adjust the model parameters of the initial image processing model according to the model loss value, and continue to train the initial image processing model until the model training stop condition is reached to obtain a target image processing model.
[0231] In an optional embodiment, the input module 404 is further configured to:
[0232] Input the sample enhanced image into the enhanced image processing sub-model to obtain first predicted annotation information and first encoded feature information;
[0233] Input the sample target image into the target image processing sub-model to obtain second predicted annotation information and second encoded feature information.
[0234] In an alternative embodiment, the enhanced image processing sub-model includes a first encoder, a first decoder, and a first classifier;
[0235] The input module 404 is further configured to:
[0236] Input the sample enhanced image into the first encoder to obtain at least one enhanced encoded feature information, where the at least one enhanced encoded feature information includes classification enhanced encoded feature information;
[0237] Input the classification enhanced encoded feature information into the first classifier to obtain first predicted classification information;
[0238] Input each enhanced encoded feature information into the first decoder to obtain first predicted segmentation information;
[0239] Determine first encoded feature information from the at least one enhanced encoded feature information.
[0240] In an alternative embodiment, the target image processing sub-model includes a second encoder, a second decoder, and a second classifier;
[0241] The input module 404 is further configured to:
[0242] Input the sample target image into the second encoder to obtain at least one target encoded feature information, where the at least one target encoded feature information includes classification target encoded feature information;
[0243] Input the classification target encoded feature information into the second classifier to obtain second predicted classification information;
[0244] Input each target encoded feature information into the second decoder to obtain second predicted segmentation information;
[0245] Determine second encoded feature information from the at least one target encoded feature information.
[0246] In an alternative embodiment, the acquisition module 402 is further configured to:
[0247] Determine a sample detection object, and acquire a sample enhanced image and a sample target image corresponding to the sample detection object, where the sample enhanced image and the sample target image corresponding to the sample detection object correspond to the same target detection area of the sample detection object;
[0248] Generate first sample annotation information for the target detection area on the sample enhanced image;
[0249] Based on the first sample annotation information, annotate second sample annotation information on the sample target image.
[0250] In an alternative embodiment, the obtaining module 402 is further configured to:
[0251] Input the sample enhanced image into an information annotation model to obtain first sample segmentation information and first sample classification information for the target detection region output by the information annotation model;
[0252] Annotating second sample annotation information on the sample target image based on the first sample annotation information includes:
[0253] Annotate second sample segmentation information on the sample target image according to the first sample segmentation information, and determine the second sample classification information of the sample target image according to the first sample classification information.
[0254] In an alternative embodiment, the input module 404 is further configured to:
[0255] Crop the sample enhanced image and the sample target image based on a preset cropping size;
[0256] Input the cropped sample enhanced image and sample target image into an initial image processing model.
[0257] In an alternative embodiment, the training module 408 is further configured to:
[0258] Adjust the model parameters of the initial image processing model according to the model loss value, and continue to train the initial image processing model until a model training stop condition is reached to obtain a reference image processing model;
[0259] Generate a target image processing model based on the target image processing sub-model in the reference image processing model.
[0260] The above device provides a framework for mutual learning between two sub-models, namely an enhanced image processing sub-model and a target image processing sub-model, unifying the classification tasks and segmentation tasks of two types of sample images, namely sample enhanced images and sample target images. Based on the two sub-models, the two types of sample images are processed respectively to obtain the predicted annotation information and encoded feature information corresponding to the two types of sample images, realizing the contrastive mutual learning between the sample enhanced image and the sample target image. The target image processing model is trained based on the predicted annotation information and encoded feature information, realizing the knowledge transfer from the enhanced image processing sub-model for processing the sample enhanced image to the target image processing sub-model for processing the sample target image, improving the processing performance and recognition performance of the target image processing model for target images, enhancing the recognition accuracy in subsequent recognition based on target images, and further enhancing the diagnostic rate.
[0261] The above is a schematic solution of a training device for an image processing model according to this embodiment. It should be noted that the technical solution of the training device for the image processing model and the technical solution of the above-mentioned image processing model training method belong to the same concept. For the details not described in detail in the technical solution of the training device for the image processing model, reference can be made to the description of the technical solution of the above-mentioned image processing model training method.
[0262] See Figure 5 , Figure 5 shows a flowchart of an image processing method provided according to an embodiment of this specification. As Figure 5 shown, it specifically includes the following steps.
[0263] Step 502: Receive an image processing task, where the image processing task carries multiple target images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormality in the target detection area.
[0264] Among them, the image processing task can be understood as a task for detecting whether there is an abnormality in the target detection area, and the image processing task carries multiple target images corresponding to the target detection area. Further, the target detection area can be understood as a partition for predicting whether an abnormality occurs. For example, the target detection area can be any organ, such as the liver, spleen, lung, stomach, etc. By predicting whether the target detection area is abnormal, the state of the object to be detected can be further assisted to be judged according to the prediction result, so as to help determine the state of the object to be detected. Detecting whether there is an abnormality in the target detection area can be understood as detecting whether there is a tumor in an organ, whether there is an embolism in a blood vessel, etc.
[0265] It should be noted that in one or more embodiments of this specification, the image processing task can be used to recognize plain CT images and judge whether there is an abnormality in the target detection area in the plain CT image according to the image features. For example, in the diagnosis scenario of pulmonary artery embolism, the plain CT image of the lung can be segmented and classified to predict whether there is a pulmonary artery embolism, so as to assist the doctor to judge whether there is an abnormality in the target detection area, and then facilitate subsequent treatment.
[0266] In a specific embodiment provided in this specification, taking the detection of whether there is an embolism in the pulmonary artery as an example for explanation. Obtain the plain CT image of the lung, that is, the multiple target images obtained are the plain CT images including the lung area, and the multiple plain CT images form a 3D image of the lung. By performing image detection on the multiple plain CT images, detect whether there is an embolism in the pulmonary artery.
[0267] Step 504: Input the multiple target images into a target image processing model to obtain a detection result for the target detection region output by the target image processing model, where the target image processing model is trained by the image processing model training method provided in the embodiments of this specification.
[0268] In practical applications, after receiving an image processing task, multiple target images carried by the image processing task are obtained from the task and input into a target image processing model for processing, and a detection result corresponding to the target detection region output by the target image processing model can be obtained. The detection result specifically includes information such as whether an abnormality appears in the target detection region and the probability of an abnormality occurring.
[0269] It should be noted that the target image processing model in the image processing method of the embodiments of this specification is obtained by training with the image processing model training method in the above embodiments. This target image processing model can not only detect target images including all target detection regions but also detect target images including partial target detection regions.
[0270] In an optional embodiment, the receiving the image processing task includes:
[0271] Receiving an image processing request sent by a user, where the image processing request carries an image processing task;
[0272] Correspondingly, the method further includes:
[0273] Sending the detection result to the user.
[0274] Specifically, the image processing task can be triggered locally on the terminal, that is, multiple target images are saved locally on the terminal, and the image processing task is triggered through the terminal to perform subsequent processes of image processing. In addition, the image processing task can also be sent by the user, that is, the terminal receives an image processing request sent by the user, the image processing request carries an image processing task, and after the image processing is completed and the detection result is obtained through the image processing method mentioned in the above embodiments, the detection result is returned to the user so that the user can perform corresponding subsequent processing according to the detection result.
[0275] Through the image processing method provided in the embodiments of this specification, a scanning and screening scheme for target images corresponding to target detection regions is provided, enabling better identification of abnormal information in the target detection region during the detection of target images, achieving high sensitivity for target image processing, and improving the accuracy of detection results.
[0276] Corresponding to the above method embodiments, this specification also provides embodiments of an image processing apparatus. Figure 6The figure shows a schematic structural diagram of an image processing device provided by an embodiment of this specification. As Figure 6 shown, the device includes:
[0277] A receiving module 602, configured to receive an image processing task, where the image processing task carries a plurality of target images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormality in the target detection area;
[0278] A prediction module 604, configured to input the plurality of target images into a target image processing model, and obtain a detection result for the target detection area output by the target image processing model, where the target image processing model is trained by the image processing model training method provided by the embodiment of this specification.
[0279] The above is a schematic solution of an image processing device in this embodiment. It should be noted that the technical solution of this image processing device and the technical solution of the above image processing method belong to the same concept. For the details not described in detail in the technical solution of the image processing device, reference can be made to the description of the technical solution of the above image processing method.
[0280] See Figure 7 , Figure 7 The figure shows a flowchart of a CT image processing method provided by an embodiment of this specification, which specifically includes the following steps:
[0281] Step 702: Receive a CT image processing task, where the CT image processing task carries a plurality of plain scan CT images corresponding to a target detection area, and the CT image processing task is used to detect whether there is an abnormality in the target detection area;
[0282] Step 704: Input the plurality of plain scan CT images into a CT image processing model, and obtain a detection result for the target detection area output by the CT image processing model, where the CT image processing model is trained by the image processing model training method provided by the embodiment of this specification.
[0283] It should be noted that the implementation manners of steps 702 to 704 are the same as those of steps 502 to 504 above, and will not be elaborated in this embodiment of this specification.
[0284] Specifically, in the method provided in this embodiment, taking the CT image processing task of detecting pulmonary embolism as an example for further explanation, based on this, the CT image processing task includes plain scan CT images of the lung region. The plain scan CT images are input into the CT image processing model for recognition. The CT image processing model can identify whether there is an embolism in the pulmonary artery in the plain scan CT images, solving the problem that the current algorithms cannot detect pulmonary embolism based on plain scan CT images and reducing the threshold for pulmonary embolism detection.
[0285] Figure 8 It is a schematic diagram of a method for training an image processing model applied to a cloud-side device provided in an embodiment of this specification. As Figure 8 shown, the method includes:
[0286] Step 802: Obtain a plurality of training sample pairs, where a training sample pair includes a sample enhanced image and a sample target image. The sample enhanced image includes first sample annotation information for the target detection region, and the sample target image includes second sample annotation information for the target detection region;
[0287] Step 804: Input the sample enhanced image and the sample target image into the initial image processing model to obtain first predicted annotation information, second predicted annotation information, first encoded feature information, and second encoded feature information output by the initial image processing model. The initial image processing model includes an enhanced image processing sub-model and a target image processing sub-model. The enhanced image processing sub-model outputs the first predicted annotation information and the first encoded feature information based on the sample enhanced image, and the target image processing sub-model outputs the second predicted annotation information and the second encoded feature information based on the sample target image;
[0288] Step 806: Calculate a model loss value according to the first sample annotation information, the second sample annotation information, the first predicted annotation information, the second predicted annotation information, the first encoded feature information, and the second encoded feature information;
[0289] Step 808: Adjust the model parameters of the initial image processing model according to the model loss value, and continue to train the initial image processing model until the model training stop condition is reached to obtain a target image processing model, and send the model parameters of the target image processing model to the edge-side device.
[0290] It should be noted that the implementation manners of steps 802 to 808 are the same as those of steps 102 to 108 above, and will not be elaborated in the embodiments of this specification.
[0291] In practical applications, since training a model requires a large amount of data and high computing resources, edge devices may not have the corresponding processing capabilities. Therefore, the model training process can be implemented on cloud devices. After obtaining the model parameters of the target image processing model, the cloud devices can also send the model parameters to the edge devices. The edge devices can construct the target image processing model locally according to the model parameters of the target image processing model and further perform image processing using the target image processing model.
[0292] Figure 9 FIG. 4 is a schematic flowchart of an image processing method applied to a cloud device according to an embodiment of the present specification, which specifically includes:
[0293] Step 902: Receive an image processing task sent by an edge device, where the image processing task carries a plurality of target images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormality in the target detection area;
[0294] Step 904: Input the plurality of target images into a target image processing model, and obtain a detection result of the target image processing model for the target detection area, where the target image processing model is trained by the image processing model training method provided in the embodiments of the present specification;
[0295] Step 906: Send the detection result to the edge device.
[0296] Figure 10 FIG. 10 is a schematic flowchart of a computer-aided diagnosis method for vascular embolism according to an embodiment of the present specification, which specifically includes:
[0297] Step 1002: Receive a vascular embolism detection task, where the vascular embolism detection task carries a plurality of plain CT images corresponding to a vascular area, and the vascular embolism detection task is used to detect whether there is an embolism in the vascular area;
[0298] Step 1004: Input the plurality of plain CT images into a target image processing model, and obtain a detection result of whether there is a vascular embolism in the vascular area output by the target image processing model, where the target image processing model is trained by the image processing model training method provided in the embodiments of the present specification.
[0299] In the present embodiment, the target image processing model is trained to detect the classification result of vascular embolism.
[0300] The computer-aided diagnosis method for vascular embolism provided in this embodiment can be applied to scenarios for screening whether there is vascular embolism in arterial vessels and venous vessels. It provides guiding suggestions for doctors and helps improve the diagnostic accuracy of doctors.
[0301] Figure 11 It is a schematic flowchart of a computer-aided diagnosis method for tumors provided in an embodiment of this specification, specifically including:
[0302] Step 1102: Receive a tumor screening task, where the tumor screening task carries multiple non-contrast CT images corresponding to the target detection area, and the tumor screening task is used to detect whether there is a tumor in the target detection area;
[0303] Step 1104: Input the multiple non-contrast CT images into a target image processing model, and obtain a detection result of whether there is a tumor in the target detection area output by the target image processing model, where the target image processing model is trained by the image processing model training method provided in the embodiments of this specification.
[0304] The computer-aided diagnosis method for tumors provided in this embodiment can be applied to tumor screenings such as pancreatic cancer, esophageal cancer, liver cancer, lung cancer, breast cancer, colorectal cancer, gastric cancer, lymphoma, etc., and scenarios for screening whether there is a tumor in various organs of the whole body such as the pancreas, esophagus, liver, lung, breast, intestine, stomach, and systemic lymph nodes. It provides guiding suggestions for doctors and helps improve the diagnostic accuracy of doctors.
[0305] Figure 12 It is a schematic diagram of a computer-aided diagnosis system for tumors provided in an embodiment of this specification. The system includes a client 1202 and a server 1204, where
[0306] The client 1202 is used to send a CT image processing task to the server 1204, where the CT image processing task carries multiple non-contrast CT images corresponding to the target detection area, and the CT image processing task is used to detect whether there is an abnormality in the target detection area;
[0307] The server 1204 is used to input the multiple non-contrast CT images into a CT image processing model, obtain a detection result of whether there is a tumor in the target detection area output by the CT image processing model, and send the detection result to the client 1202, where the CT image processing model is trained by the image processing model training method provided in the embodiments of this specification.
[0308] Figure 13FIG. 0 shows a structural block diagram of a computing device 1300 provided according to an embodiment of this specification. The components of the computing device 1300 include, but are not limited to, a memory 1310 and a processor 1320. The processor 1320 is connected to the memory 1310 via a bus 1330, and a database 1350 is used to store data.
[0309] The computing device 1300 further includes an access device 1340, which enables the computing device 1300 to communicate via one or more networks 1360. Examples of these networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 1340 may include one or more of any type of wired or wireless network interfaces (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0310] In an embodiment of the present application, the above components of the computing device 1300 and Figure 13 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 13 the shown structural block diagram of the computing device is for illustrative purposes only and is not a limitation on the scope of the present application. Those skilled in the art can add or replace other components as needed.
[0311] The computing device 1300 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook, etc.), a mobile phone (e.g., a smart phone), a wearable computing device (e.g., a smart watch, smart glasses, etc.) or other types of mobile devices, or a stationary computing device such as a desktop computer or a personal computer (PC). The computing device 1300 can also be a mobile or stationary server.
[0312] Wherein, the processor 1320 is configured to execute the following computer program / instructions, and when the computer program / instructions are executed by the processor, the steps of the above method are implemented.
[0313] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the computing device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the description of the method embodiment.
[0314] An embodiment of this specification also provides a computer-readable storage medium, which stores computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the above method are implemented.
[0315] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the computer-readable storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the relevant parts, reference can be made to the description of the method embodiment.
[0316] An embodiment of this specification also provides a computer program product, including computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the above method are implemented.
[0317] The above is a schematic solution of a computer program product of this embodiment. It should be noted that the technical solution of the computer program product and the technical solution of the above method belong to the same concept. For the details not described in detail in the technical solution of the computer program product, reference can be made to the description of the technical solution of the above method.
[0318] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0319] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0320] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described order of actions, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0321] In the above embodiments, the descriptions of the various embodiments each have their own emphasis. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0322] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not elaborate on all the details and do not limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A training method for an image processing model, comprising: Acquire a plurality of training sample pairs, wherein the training sample pairs include a sample enhanced image and a sample target image, the sample enhanced image includes first sample annotation information for a target detection area, and the sample target image includes second sample annotation information for the target detection area; Inputting the sample enhanced image and the sample target image into an initial image processing model, obtaining first predicted labeling information, second predicted labeling information, first coded feature information, and second coded feature information output by the initial image processing model, wherein the initial image processing model includes an enhanced image processing sub-model and a target image processing sub-model, the enhanced image processing sub-model outputs the first predicted labeling information and the first coded feature information based on the sample enhanced image, and the target image processing sub-model outputs the second predicted labeling information and the second coded feature information based on the sample target image; Calculate a model loss value according to the first sample labeling information, the second sample labeling information, the first prediction labeling information, the second prediction labeling information, the first encoding feature information, and the second encoding feature information; The model parameters of the initial image processing model are adjusted according to the model loss value, and the initial image processing model is continuously trained until the model training stop condition is reached to obtain the target image processing model.
2. The method according to claim 1, wherein the step of calculating the model loss value according to the first sample labeling information, the second sample labeling information, the first prediction labeling information, the second prediction labeling information, the first encoding feature information, and the second encoding feature information comprises: Calculating a first model loss value according to the first sample labeling information and the first prediction labeling information; Calculating a second model loss value according to the second sample labeling information and the second prediction labeling information; Calculating an information difference loss value according to the first predicted labeling information and the second predicted labeling information; Calculating a feature information contrast loss value according to the first encoding feature information and the second encoding feature information; The model loss value is calculated according to the first model loss value, the second model loss value, the information difference loss value and the feature information comparison loss value.
3. The method according to claim 2, wherein the step of calculating the feature information contrast loss value according to the first encoding feature information and the second encoding feature information comprises: Normalizing the first encoding feature information and the second encoding feature information to obtain first embedded vector information and second embedded vector information; Calculating similarity information according to the first embedding vector information and the second embedding vector information; The feature information comparison loss value is calculated according to the similarity information.
4. The method of claim 2, wherein the first prediction labeling information includes first prediction classification information and first prediction segmentation information, and the second prediction labeling information includes second prediction classification information and second prediction segmentation information; The calculating the information difference loss value according to the first predicted labeling information and the second predicted labeling information includes: Calculating a classification difference loss value according to the first predicted classification information and the second predicted classification information; Calculating a segmentation difference loss value according to the first predicted segmentation information and the second predicted segmentation information; The information difference loss value is determined according to the classification difference loss value and the segmentation difference loss value.
5. The method according to claim 1, wherein adjusting the model parameters of the initial image processing model according to the model loss value, and continuing to train the initial image processing model until a model training stop condition is reached to obtain a target image processing model, comprises: Based on the gradient descent algorithm, the model parameters of the initial image processing model are adjusted according to the model loss value, and the initial image processing model is continuously trained until the model training stop condition is reached to obtain the target image processing model.
6. The method of claim 1, wherein the sample enhanced image and the sample target image are input into an initial image processing model to obtain first predicted labeling information, second predicted labeling information, first encoded feature information, and second encoded feature information output by the initial image processing model, comprising: Inputting the sample enhanced image into the enhanced image processing sub-model to obtain first predicted labeling information and first encoded feature information; The sample target image is input into the target image processing sub-model to obtain second predicted labeling information and second encoded feature information.
7. The method of claim 6, wherein the enhanced image processing sub-model comprises a first encoder, a first decoder, and a first classifier; The step of inputting the sample enhanced image into the enhanced image processing sub-model to obtain first prediction labeling information and first encoding feature information includes: Inputting the sample enhanced image into the first encoder to obtain at least one enhanced encoding feature information, wherein the at least one enhanced encoding feature information includes classification enhanced encoding feature information; Inputting the classification enhancement coding feature information into the first classifier to obtain first predicted classification information; Inputting each enhanced coding feature information into the first decoder to obtain first prediction segmentation information; A first encoding feature information is determined in the at least one enhanced encoding feature information.
8. An image processing method, comprising: receiving an image processing task, wherein the image processing task carries a plurality of target images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormality in the target detection area; The multiple target images are input into a target image processing model to obtain a detection result for the target detection area output by the target image processing model, wherein the target image processing model is trained by the training method of the image processing model described in any one of claims 1-7.
9. A CT image processing method, comprising: receiving a CT image processing task, wherein the CT image processing task carries a plurality of plain scan CT images corresponding to a target detection area, and the CT image processing task is used to detect whether there is an abnormality in the target detection area; The multiple plain scan CT images are input into a CT image processing model to obtain a detection result for the target detection area output by the CT image processing model, wherein the CT image processing model is trained by the training method of the image processing model described in any one of claims 1-7.
10. A training method for an image processing model, applied to a cloud-side device, comprising: Acquire a plurality of training sample pairs, wherein the training sample pairs include a sample enhanced image and a sample target image, the sample enhanced image includes first sample annotation information for a target detection area, and the sample target image includes second sample annotation information for the target detection area; Inputting the sample enhanced image and the sample target image into an initial image processing model, obtaining first predicted labeling information, second predicted labeling information, first coded feature information, and second coded feature information output by the initial image processing model, wherein the initial image processing model includes an enhanced image processing sub-model and a target image processing sub-model, the enhanced image processing sub-model outputs the first predicted labeling information and the first coded feature information based on the sample enhanced image, and the target image processing sub-model outputs the second predicted labeling information and the second coded feature information based on the sample target image; Calculate a model loss value according to the first sample labeling information, the second sample labeling information, the first prediction labeling information, the second prediction labeling information, the first encoding feature information, and the second encoding feature information; Adjust the model parameters of the initial image processing model according to the model loss value, continue to train the initial image processing model until the model training stop condition is reached, obtain the target image processing model, and send the model parameters of the target image processing model to the end-side device.
11. An image processing method, applied to a cloud-side device, comprising: An image processing task sent by a receiving device, wherein the image processing task carries a plurality of target images corresponding to a target detection area, and the image processing task is used to detect whether there is an abnormality in the target detection area; Inputting the multiple target images into a target image processing model to obtain a detection result output by the target image processing model for the target detection area, wherein the target image processing model is trained by the training method of the image processing model according to any one of claims 1 to 7; Send the detection result to the terminal side device.
12. A computer-aided diagnosis method for vascular embolism, comprising: receiving a vascular embolism detection task, wherein the vascular embolism detection task carries a plurality of plain scan CT images corresponding to a vascular region, and the vascular embolism detection task is used to detect whether there is embolism in the vascular region; The multiple plain scan CT images are input into a target image processing model to obtain a detection result output by the target image processing model regarding whether vascular embolism exists in the vascular region, wherein the target image processing model is trained using the training method of the image processing model described in any one of claims 1 to 7.
13. A computer-aided diagnosis method for tumors, comprising: receiving a tumor screening task, wherein the tumor screening task carries a plurality of plain scan CT images corresponding to a target detection area, and the tumor screening task is used to detect whether a tumor exists in the target detection area; The multiple plain scan CT images are input into a target image processing model to obtain a detection result output by the target image processing model regarding whether a tumor exists in the target detection area, wherein the target image processing model is trained using the training method of the image processing model described in any one of claims 1 to 7.
14. A computer-aided diagnosis system for tumors, comprising a client and a server, wherein: The client is used to send a CT image processing task to the server, wherein the CT image processing task carries a plurality of plain scan CT images corresponding to a target detection area, and the CT image processing task is used to detect whether there is an abnormality in the target detection area; The server is used to input the multiple plain scan CT images into a CT image processing model, obtain the detection result output by the CT image processing model regarding whether a tumor exists in the target detection area, and send the detection result to the client, wherein the CT image processing model is trained by the training method of the image processing model described in any one of claims 1-7.
15. A computing device comprising: Memory and processor; The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer program / instructions are executed by the processor, the steps of the method described in any one of claims 1 to 13 are implemented.
16. A computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction, when executed by a processor, implements the steps of the method according to any one of claims 1 to 13.
17. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 13.