Training method and device, image processing method, storage medium and electronic equipment
By distilling knowledge and pseudo-label training on medical image processing models, combined with case text, the problem of low accuracy in existing medical image processing methods is solved, and more accurate lesion detection and recognition is achieved.
Patent Information
- Application Number
- CN202311805604.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2025-06-27
AI Technical Summary
The existing medical imaging processing methods have the problem of low accuracy in processing results, especially in lesion detection and identification, and it is difficult to obtain accurate identification and matching.
A training method is adopted to distil the image detection module knowledge by inputting medical images into a contrasting language image pretraining network as a teacher model, obtaining pseudo-labels of the lesion area, and training the medical image processing model in combination with case text.
It improves the accuracy of medical imaging processing results, can detect and identify lesions more accurately, and accurately match the case text, reduces redundant information and improves the readability of processing results.
Smart Images

Figure CN120219273A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the technical field of medical image processing, and relates to a training method, in particular to a training method and device, an image processing method, a storage medium, and an electronic device. Background Art
[0002] With the rapid development of medical imaging, medical images have become increasingly important in the modern medical field. Currently, widely used medical images include X-ray images, CT (Computed Tomography) images, MRI (Magnetic Resonance Imaging) images, etc. However, the existing methods for processing medical images generally have the problem of low accuracy of processing results. Summary of the Invention
[0003] Embodiments of this application provide a training method and device, an image processing method, a storage medium, and an electronic device for improving the accuracy of medical image processing results.
[0004] In a first aspect, embodiments of this application provide a training method for a medical image processing model. The medical image processing model includes an image detection module, a text encoder, and a fusion module. Among them, the fusion module is used to fuse the outputs of the image detection module and the text encoder to generate a medical image processing result. The training method includes: using a medical image as an input, and performing knowledge distillation on the image detection module with the image encoder in a contrastive language-image pre-training network as a teacher model; using the image detection module to process the medical image to obtain its lesion area; using the contrastive language-image pre-training network to match the lesion area and medical entity vocabulary to obtain a pseudo-label corresponding to the lesion area; and training the medical image processing model according to the medical image, the lesion area, and the pseudo-label.
[0005] In an implementation manner of the first aspect, the image detection module includes a region proposal network and a detection head. Using a medical image as an input and performing knowledge distillation on the image detection module with the image encoder in a contrastive language-image pre-training network as a teacher model includes: the teacher model performing object detection on the medical image through a two-stage object detection model to generate corresponding detection regions, and inputting the detection regions into the image encoder to obtain image embeddings; the image detection module processing the medical image through the region proposal network and the detection head to obtain region embeddings; obtaining a distillation loss according to the image embeddings and the region embeddings; and training the region proposal network and the detection head using the distillation loss.
[0006] In one implementation of the first aspect, the training method further includes: using MetaMap to convert the pseudo-labels corresponding to the lesion regions into a unified standardized medical language.
[0007] In one implementation of the first aspect, the image detection module processes the medical image to obtain image features, the text encoder processes the case text to obtain text features, and the fusion module fuses the image features and the text features to obtain the medical image processing result.
[0008] In one implementation of the first aspect, the fusion module includes multiple fusion layers, and the fusion layer uses the multi-head attention mechanism to fuse the image features and the text features to obtain the medical image processing result.
[0009] In one implementation of the first aspect, the loss function used in training the medical image processing model is a semantic similarity loss function, and the semantic similarity loss function is determined according to the cross-entropy loss from the image to the label and the cross-entropy loss from the label to the image.
[0010] In a second aspect, an embodiment of the present application provides a medical image processing method, and the medical image processing method includes: obtaining a pair of a medical image to be processed and case text; using a medical image processing model to process the pair of the medical image to be processed and the case text, where the medical image processing model is trained by using the training method described in any one of the first aspects of the embodiments of the present application.
[0011] In a third aspect, an embodiment of the present application provides a training device for a medical image processing model. The medical image processing model includes an image detection module, a text encoder, and a fusion module, where the fusion module is used to fuse the outputs of the image detection module and the text encoder to generate a medical image processing result. The training device includes: a knowledge distillation module, which is used to take a medical image as an input and perform knowledge distillation on the image detection module by using the image encoder in the contrastive language-image pre-training network as a teacher model; a lesion region acquisition module, which is used to use the image detection module to process the medical image to obtain its lesion region; a pseudo-label acquisition module, which is used to use the contrastive language-image pre-training network to match the lesion region and medical entity vocabulary to obtain the pseudo-labels corresponding to the lesion region; and a model training module, which is used to train the medical image processing model according to the medical image, the lesion region, and the pseudo-labels.
[0012] Fourthly, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described in any item of the first aspect of the embodiments of the present application is implemented.
[0013] Fifthly, an embodiment of the present application provides an electronic device, which includes: a memory storing a computer program; a processor communicatively connected to the memory, and when the computer program is called, the method described in any item of the first aspect of the embodiments of the present application is executed.
[0014] An embodiment of the present application provides a training method for training a medical image processing model. After training, the medical image processing model can process medical images and case texts to obtain a processing result. Among them, the case text is, for example, the relevant text describing the image in the case. It can be seen that the medical image processing model can process by combining the case text and the medical image, and thus can obtain a more accurate processing result.
[0015] In addition, in some embodiments, the trained medical image processing model can be used for lesion detection and recognition. At this time, the medical image processing model can detect and recognize the key lesions proposed by the doctor according to the case text, so as to avoid identifying redundant information, making the detection and recognition results have higher readability and being convenient for non-professionals to understand. Description of the Drawings
[0016] Figure 1 It shows a schematic diagram of an application scenario of an embodiment of the present application.
[0017] Figure 2A It shows a schematic diagram of the structure of the medical image processing model in an embodiment of the present application.
[0018] Figure 2B It shows a flowchart of the training method provided by an embodiment of the present application.
[0019] Figure 3A It shows a schematic diagram of the structure of the image detection module in an embodiment of the present application.
[0020] Figure 3B It shows a flowchart of knowledge distillation in an embodiment of the present application.
[0021] Figure 4 It shows a flowchart of the medical image processing method provided by the present application.
[0022] Figure 5 It shows a schematic diagram of the structure of the training device provided by an embodiment of the present application.
[0023] Description of Reference Numerals
[0024] 1 Electronic device
[0025] 11 Memory
[0026] 12 Display
[0027] 13 General - purpose processor
[0028] 131 Central processing unit
[0029] 132 Neural network processor
[0030] 2 Medical image processing model
[0031] 21 Image detection module
[0032] 211 Region proposal network
[0033] 212 Detection head
[0034] 22 Text encoder
[0035] 23 Fusion module
[0036] 5 Training device
[0037] 51 Knowledge organization module
[0038] 52 Lesion area acquisition module
[0039] 53 Pseudo - label acquisition module
[0040] 54 Model training module
[0041] Steps S21 - S24
[0042] Steps S31 - S34
[0043] Steps S41 - S42 Specific implementation manners
[0044] The following uses specific specific examples to illustrate the implementation manners of the present application. Those skilled in the art can easily understand other advantages and effects of the present application from the content disclosed in this specification. The present application can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0045] It should be noted that the illustrations provided in the following embodiments only schematically illustrate the basic concept of the present application. Therefore, only the components related to the present application are shown in the illustrations, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.
[0046] With the rapid development of medical imaging, medical images have become increasingly important in the modern medical field. Currently, widely used medical images include X-ray images, CT images, MRI, etc. However, the examination results of medical images usually only have text descriptions and do not mark the lesions in the images, which increases the difficulty for non-professionals to understand medical images.
[0047] In addition, in some work, the marking results of lesions are often required. Due to the lack of lesion markings in medical images, manual methods are generally used in related technologies to mark the lesions in medical images, which will lead to an increase in time and cost. To improve efficiency and reduce costs, in some technical solutions, electronic devices such as computers are used to detect and identify lesions. However, the inventors found that these technical solutions only use medical images for detection and identification, and do not use text information such as case texts. This results in these technical solutions being unable to mark different sizes and types of lesions well, and only obtaining the most basic classification, and there is a large amount of irrelevant information in the detection and identification results. In addition, due to the lack of use of text information such as case texts, the lesions identified in these technical solutions cannot be accurately matched with the case content.
[0048] At least for the above problems, the embodiments of the present application provide a training method, which is used to train an image processing model. After training, the image processing model can combine text and medical images for processing to obtain more accurate processing results.
[0049] The training method provided by the embodiments of the present application can be applied, for example, to Figure 1 the electronic device 1 shown in the figure. As Figure 1 shown, the electronic device 1 includes a memory 11, a display 12, and at least one general-purpose processor 13.
[0050] The memory 11 may include volatile memory, such as random access memory (RAM), cache. The memory 11 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD). The memory 11 can be used to store program instructions for the processor to call and execute at least some of the steps of the training method provided in the embodiments of the present application.
[0051] The display 12 may specifically include a display screen (display panel). In some implementations, the display panel may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. In addition, the display 12 may also be a touch panel (touch screen, touch display screen), and the touch panel may include a display screen and a touch-sensitive surface. When the touch-sensitive surface detects a touch operation on or near it, it is transmitted to the general-purpose processor 13 to determine the type of touch event, and then the general-purpose processor 13 provides a corresponding visual output on the display device according to the type of touch event.
[0052] The general-purpose processor 13 can be any type of device capable of processing electronic instructions. In the embodiments of the present application, the electronic device 1 may include one or more general-purpose processors 13, such as one or both of a central processing unit (CPU) 131 and a neural-network processing unit (NPU) 132. In addition, it may also include one or more of a graphics processing unit (GPU), a microprocessor, a microcontroller, a main processor, a controller, and an application specific integrated circuit (ASIC), etc. The general-purpose processor 13 is configured to execute various types of digital storage instructions, such as software or firmware programs stored in the memory 11, which enables the electronic device 1 to provide various services. For example, the processor 11 can execute programs or process data to execute at least some of the steps of the training method provided in the embodiments of the present application.
[0053] The function of the central processing unit 131 is mainly to parse computer instructions and process data in computer software, so as to realize the overall control of the electronic device 1 and control and allocate all hardware resources of the electronic device 1 (such as storage resources, communication resources, I / O interfaces, etc.).
[0054] The neural network processor is a general term for a new type of processor based on neural network algorithms and acceleration. It is specifically designed for artificial intelligence and is used to accelerate the operation of neural networks to solve the problem of low efficiency of traditional chips in neural network operations.
[0055] It should be noted that the name of the neural network processor does not limit this application. For example, in other application scenarios, the neural network processor can also be deformed and replaced with other processors with similar functions, such as Tensor Processing Unit (TPU), Deep learning Processing Unit (DPU), etc.
[0056] Figure 2A It shows a schematic structural diagram of the medical image processing model 2 in the embodiment of this application. As Figure 2A shown, the medical image processing model 2 inputs case texts and medical images and outputs the processing results of the medical images. The processing results include but are not limited to lesion detection and recognition results. For example, in some other embodiments, the medical image processing model can also be used to perform processing such as segmentation on medical images. The medical image processing model 2 includes an image detection module 21, a text encoder 22, and a fusion module 23. The image detection module 21 is used to process the input medical images, the text encoder 22 is used to process the input case texts, and the fusion module is used to fuse the processing results of the image detection module 21 and the text encoder 22 and output the final processing results.
[0057] Figure 2B It shows a flowchart of the training method provided by the embodiment of this application. This training method can be used to Figure 2A train the medical image processing model 21 shown. As Figure 2B shown, the training method includes the following steps S21 to S24.
[0058] S21. Take the medical image as input, and use the image encoder in the Contrastive Language-Image Pretraining network as the teacher model to perform knowledge distillation on the image detection module. CLIP can process images and text simultaneously and can understand the semantic relationship between images and text. CLIP includes an image encoder, which is used to convert the input image into semantic embeddings. Knowledge distillation is a model compression technique that trains by transferring knowledge from a large and complex model (the teacher model, i.e., the image encoder in CLIP) to a small and simplified model (the student model, i.e., the image detection module). The complex knowledge of the teacher model is converted into a more concise and lightweight form to help the student model learn and generalize more efficiently. This helps to use a more lightweight model in an environment with limited resources while maintaining the model performance.
[0059] S22. Use the image detection module after knowledge distillation to process the medical image to obtain its lesion area.
[0060] S23. Use the Contrastive Language-Image Pretraining network to match the lesion area with medical entity words to obtain the pseudo-label corresponding to the lesion area. Among them, the medical entity words are, for example, relevant words in the case text, such as cyst, hypertension, etc.
[0061] S24. Train the medical image processing model according to the medical image, the lesion area, and the pseudo-label. After training, the medical image processing model can be used to infer new medical image and case text pairs.
[0062] In some implementation manners, the medical image is, for example, an X-ray image, but the present application is not limited thereto.
[0063] Figure 3A It shows a schematic diagram of the architecture of the image detection module 21 in the embodiment of the present application. As Figure 3A shown, the image detection module 21 includes a Region Proposal Network (RPN) 211 and a detection head 212. The Region Proposal Network 211 is, for example, a Swin-L RPN, but the present application is not limited thereto. Figure 3B It shows a flowchart of taking the medical image as input and using the image encoder in the Contrastive Language-Image Pretraining network as the teacher model to perform knowledge distillation on the image detection module in the embodiment of the present application. As Figure 3B shown, the above knowledge distillation process includes the following steps S31 to S34.
[0064] In S31, the teacher model performs object detection on medical images through a two-stage object detection model to generate corresponding detection regions, and inputs the detection regions into an image encoder to obtain image embeddings. Among them, the two-stage object detection model is, for example, Mask R-CNN.
[0065] In S32, the image detection module processes medical images through a region proposal network and a detection head to obtain region embeddings. The region embeddings are, for example, vectors of suspicious lesions in medical images.
[0066] In S33, a distillation loss is obtained based on the image embeddings and the region embeddings.
[0067] In some implementation manners, the distillation loss can be represented by the following formula (1):
[0068]
[0069] Among them, E r represents the region embeddings, E i represents the image embeddings, and m represents the number of image embeddings.
[0070] In S34, the region proposal network and the detection head are trained using the distillation loss.
[0071] Through the above process, knowledge distillation from the teacher model to the image detection module can be achieved. During the knowledge distillation process, the weights of the teacher model are not updated and learned. Through knowledge distillation, the image detection module can obtain the strong robustness of CLIP and similar feature expressions to match the corresponding text feature expressions.
[0072] In some implementation manners, the training method may further include: extracting medical entity information from case texts and converting it into a unified medical language (University of Missouri–St. Louis, UMSL) to obtain medical entity vocabulary. Inputting the medical entity vocabulary and medical images into the trained image detection module for inference and matching can obtain the matching relationship between the lesion regions and the medical entity vocabulary, thereby obtaining the pseudo-labels corresponding to the lesion regions.
[0073] Exemplarily, MetaMap can be used to convert medical entity information into medical entity vocabulary.
[0074] In some implementation manners, the training method may further include: using MetaMap to convert the pseudo-labels corresponding to the lesion regions into UMSL.
[0075] In some implementations, the image detection module processes medical images to obtain image features, and the text encoder processes case texts to obtain text features. The fusion module fuses the image features and the text features to obtain the medical image processing result.
[0076] In some implementations, the fusion module includes multiple fusion layers. The fusion layers can use the Cross-Modal Multi-Head Attention Mechanism (X-MHA) to fuse the image features and the text features, as shown in the following formulas 2 to 4:
[0077]
[0078] Among them, L is the number of fusion layers in the fusion module, O 0 and P 0 respectively represent the feature information output by the image detection module and the text encoder. X-MHA represents the multi-head attention mechanism, whose input is O i and P i , and the output is the image features integrated with text features and the text features integrated with image features
[0079] Cross-modal information can be obtained through the multi-head attention mechanism and this information is fused into a single modality for updating. Specifically, as shown in formula 3, the fused feature and its original feature O i are input into the corresponding network for training to obtain the output result O i+1 of the next layer. As shown in formula 4, the fused feature and the original feature P i are input into the corresponding network for training to obtain the output result P i+1 of the next layer.
[0080]
[0081]
[0082] Among them, D represents the fusion operation, O represents the output of a certain layer of the image detection module, and P represents the output of a certain layer of the text encoder.
[0083] In the multi-head attention mechanism, each head calculates a context vector of one modality by paying attention to the other modality, and the specific formulas are as shown in the following formulas 5 to 7:
[0084]
[0085] P (v) = PW (v,L) , O t2i = softmax(Atnn)P (v) W (out,I) , Equation 6;
[0086] O (v) = OW (v,I) , P t2i = softmax(Atnn T )O (v) W (out,L) , Equation 7;
[0087] Among them, {W (symbol,L) , W (symbol,I) : symbol ∈ {q, v, out}} are all trainable parameters, representing the query (q), value (v), and output (out) linear layers in the multi-head attention mechanism respectively. softmax represents the softmax function, and d represents the size of O (q) .
[0088] In the embodiments of the present application, through the cross-modal multi-head attention mechanism, information interaction between text features and image features can be achieved, thereby improving the performance of the medical image processing model.
[0089] In some implementation manners, the loss function used in training the medical image processing model is a semantic similarity loss function, and the semantic similarity loss function is determined according to the cross-entropy loss from the image to the label and the cross-entropy loss from the label to the image.
[0090] Specifically, in the embodiments of the present application, the following Equation 8 can be used to calculate the medical semantic similarity s between the text label and the image:
[0091]
[0092] Among them, I img and I txt represent the medical image and the text label respectively. Based on the above Equation 8, for each medical image i, the medical semantic similarity s of the corresponding text label j can be obtained ij . In the medical image processing model, calculating a soft target between the medical image i and the text label j can be obtained by normalizing j through softmax, as shown in the following Equation 9:
[0093]
[0094] Among them, N represents the number of text labels included in the medical image in one training.
[0095] Based on the above formula (8), for each text label j, the medical semantic similarity s of the corresponding medical image i can be obtained. ji In the medical image processing model, a soft target between the text label j and the medical image i can be obtained by normalizing i through softmax, as shown in the following formula (10):
[0096]
[0097] In the medical image processing model, logits can be obtained from the cosine similarity between the text embedding and the image embedding, which is the same as the alignment loss of the model result. Among them, Logits represents the output of the model before passing through activation functions such as softmax or sigmoid. The alignment loss can be calculated using the following formula (11):
[0098]
[0099] Among them, v represents an image vector in the image embedding, and t represents a text vector in the text embedding. and respectively represent the normalized v p and t p That is, after normalizing the p-th image vector v p the i-th normalization result is obtained. After normalizing the p-th text vector t p the j-th normalization result is obtained. The similarity of the result of the medical image processing model can be obtained through softmax, as shown in the following formula (12):
[0100]
[0101] Among them, τ is a set value, and its value is, for example, 0.07.
[0102] In some implementation manners, the cross-entropy calculation method can be used to calculate the semantic similarity between the medical image and the text label, as shown in the following formulas (13) and (14):
[0103]
[0104]
[0105] According to the above formulas (13) and (14), the semantic similarity loss function L adopted by the medical image processing model during training can be obtained:
[0106]
[0107] The above semantic similarity loss function L integrates the alignment loss of the CLIP model itself and the semantic similarity loss related to medical knowledge.
[0108] According to the above description, it can be known that the training method provided by the embodiments of the present application can use the image encoder in CLIP as a guide, adopt the knowledge distillation method to train the image detection part of the medical image processing model, and can use MetaMap to extract case texts and use CLIP to generate pseudo-labels for medical images. The medical image processing model adopts the overall structure of self-supervised and case text description to detect the suspicious lesion area of the medical image, and uses the semantic similarity loss function in training. In this way, the trained medical image processing model can have higher accuracy.
[0109] The embodiments of the present application also provide a medical image processing method. Figure 4 Shown is a flowchart of the medical image processing method provided by the embodiments of the present application. As Figure 4 shown, the medical image processing method includes the following steps S41 and S42.
[0110] S41, obtaining a pair of medical image and case text to be processed. Specifically, when generating a medical image, relevant medical staff will give the case text corresponding to the medical image, including the description of the lesion state. The medical image and its corresponding case text can be used as a pair of medical image and case text.
[0111] S42, using the medical image processing model to process the pair of medical image and case text to be processed. Among them, the medical image processing model is trained by the training method provided by the embodiments of the present application.
[0112] The protection scope of the training method and the medical image processing method provided by the embodiments of the present application is not limited to the execution order of the steps listed in this embodiment. Any solution realized by adding or subtracting steps of the prior art and replacing steps according to the principle of the present application is included in the protection scope of the present application.
[0113] The embodiments of the present application also provide a training device for a medical image processing model. Among them, the medical image processing model includes an image detection module, a text encoder, and a fusion module, and the fusion module is used to fuse the outputs of the image detection module and the text encoder to generate a medical image processing result. Figure 5 Shown is a schematic structural diagram of the training device 5 provided by the embodiments of the present application. As Figure 5As shown, the training device 5 includes a knowledge distillation module 51, a lesion area acquisition module 52, a pseudo-label acquisition module 53, and a model training module 54. The knowledge distillation module 51 is used to take a medical image as input and perform knowledge distillation on the image detection module using the image encoder in the contrastive language-image pre-training network as the teacher model. The lesion area acquisition module 52 is used to process the medical image using the image detection module to obtain its lesion area. The pseudo-label acquisition module 53 is used to match the lesion area with medical entity vocabulary using the contrastive language-image pre-training network to obtain the pseudo-label corresponding to the lesion area. The model training module 54 is used to train the medical image processing model based on the medical image, the lesion area, and the pseudo-label.
[0114] It should be noted that each of the above modules included in the training device 5 corresponds one-to-one to Figure 2B steps S21 to S24 in, and details are not described herein.
[0115] In several embodiments provided in the present application, it should be understood that the disclosed system, device, or method can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules / units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or modules or units can be in electrical, mechanical, or other forms.
[0116] The modules / units described as separate components may or may not be physically separated. The components shown as modules / units may or may not be physical modules, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules / units can be selected according to actual needs to achieve the purpose of the embodiments of the present application. For example, in each embodiment of the present application, the various functional modules / units can be integrated in a processing module, or each module / unit can exist physically alone, or two or more modules / units can be integrated in one module / unit.
[0117] Those of ordinary skill in the art should further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.
[0118] The embodiments of this application also provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the training method and / or medical image processing method provided by the embodiments of this application. Those of ordinary skill in the art can understand that all or part of the steps in the methods of the above embodiments can be completed by instructing a processor through a program. The described program can be stored in a computer-readable storage medium, and the storage medium is a non-transitory medium, such as random access memory, read-only memory, flash memory, hard disk, solid-state drive, magnetic tape, floppy disk, optical disc, and any combination thereof. The above storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a digital video disc (DVD)), or a semiconductor medium (such as a solid-state disk (SSD)).
[0119] The embodiments of this application can also provide an electronic device, which includes a memory and a processor. The memory is used to store a computer program. In some possible implementation manners, the memory may include various media that can store program codes, such as ROM, RAM, magnetic disk, USB flash drive, memory card, or optical disc. The processor is connected to the memory and is used to execute the computer program stored in the memory, so that the electronic device executes the training method and / or medical image processing method provided by the embodiments of this application.
[0120] In some implementation manners, the electronic device provided by the embodiments of this application may further include a display. The display YYY30 is communicatively connected to the memory and the processor and is used to display the relevant graphical user interface (GUI) of the training method and / or medical image processing method.
[0121] The descriptions of the processes or structures corresponding to the above-mentioned respective drawings each have their own focuses. For parts not detailed in a certain process or structure, reference may be made to the relevant descriptions of other processes or structures.
[0122] The above embodiments are only illustrative of the principles and effects of the present application and are not intended to limit the present application. Any person familiar with this technology can modify or change the above embodiments without departing from the spirit and scope of the present application. Therefore, all equivalent modifications or changes made by those with ordinary knowledge in the technical field without departing from the spirit and technical idea disclosed by the present application should still be covered by the claims of the present application.
Claims
1. A training method for a medical image processing model, characterized in that, The medical image processing model includes an image detection module, a text encoder, and a fusion module. Among them, the fusion module is used to fuse the outputs of the image detection module and the text encoder to generate a medical image processing result. The training method includes: Taking a medical image as input, and performing knowledge distillation on the image detection module using the image encoder in the contrastive language-image pre-training network as the teacher model; Using the image detection module to process the medical image to obtain its lesion area; Using the contrastive language-image pre-training network to match the lesion area and medical entity vocabulary to obtain the pseudo-label corresponding to the lesion area; Training the medical image processing model according to the medical image, the lesion area, and the pseudo-label.
2. The training method according to claim 1, wherein The image detection module includes a region proposal network and a detection head. Taking a medical image as input and performing knowledge distillation on the image detection module using the image encoder in the contrastive language-image pre-training network as the teacher model includes: The teacher model performs object detection on the medical image through a two-stage object detection model to generate corresponding detection regions, and inputs the detection regions into the image encoder to obtain image embeddings; The image detection module processes the medical image through the region proposal network and the detection head to obtain region embeddings; Obtaining a distillation loss according to the image embedding and the region embedding; Training the region proposal network and the detection head using the distillation loss.
3. The training method according to claim 1, characterized in that It also includes: Using MetaMap to convert the pseudo-label corresponding to the lesion area into a unified standardized medical language.
4. The training method according to claim 1, characterized in that, The image detection module processes the medical image to obtain image features, the text encoder processes the case text to obtain text features, and the fusion module fuses the image features and the text features to obtain the medical image processing result.
5. The training method according to claim 4, characterized in that The fusion module includes multiple fusion layers, and the fusion layer uses a multi-head attention mechanism to fuse the image features and the text features to obtain the medical image processing result.
6. The training method according to claim 1, wherein When training the medical image processing model, the loss function used is a semantic similarity loss function, and the semantic similarity loss function is determined according to the cross-entropy loss from the image to the label and the cross-entropy loss from the label to the image.
7. A medical image processing method, characterized in that, The medical image processing method includes: Obtaining a pair of a medical image to be processed and case text; Using the medical image processing model to process the pair of the medical image to be processed and case text, where the medical image processing model is trained using the training method described in any one of claims 1 to 6.
8. A training device for a medical image processing model, characterized in that The medical image processing model includes an image detection module, a text encoder, and a fusion module. Among them, the fusion module is used to fuse the outputs of the image detection module and the text encoder to generate a medical image processing result. The training device includes: A knowledge distillation module, which is used to take a medical image as input and perform knowledge distillation on the image detection module using the image encoder in the contrastive language-image pre-training network as the teacher model; A lesion area acquisition module, configured to process the medical image by using the image detection module to obtain its lesion area; A pseudo-label acquisition module, configured to match the lesion area and medical entity vocabulary by using the contrastive language-image pre-training network to obtain a pseudo-label corresponding to the lesion area; A model training module, configured to train the medical image processing model according to the medical image, the lesion area, and the pseudo-label.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.
10. An electronic device, characterized in that, The electronic device includes: A memory storing a computer program; A processor communicatively connected to the memory, and when calling the computer program, executes the method according to any one of claims 1 to 7.