Medical image processing methods, devices, computer equipment and storage media
A medical image processing model using a saliency prediction sub-model and an attribution prediction sub-model generates a fused image of saliency image and attribution map, and outputs a scoring attribution map. This solves the problem of relying on the experience of medical personnel in existing technologies and improves the efficiency and accuracy of medical image recognition.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PING AN BANK CO LTD
- Filing Date
- 2022-06-15
- Publication Date
- 2026-05-26
AI Technical Summary
In the medical field, existing neural network models rely on the experience of medical personnel, which makes it difficult to guarantee the effectiveness and efficiency of medical image recognition, and may lead to misdiagnosis when medical personnel lack experience.
A medical image processing model employing a saliency prediction sub-model, an attribution prediction sub-model, and a decision sub-model generates a fused image of saliency image and attribution map, outputting a score attribution map to represent health status and score.
It improves the efficiency and effectiveness of medical image recognition, enables a deeper interpretation of target regions in medical images, and enhances the accuracy and efficiency of diagnosis.
Smart Images

Figure CN115205217B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computers, specifically to a medical image processing method, model, device, computer equipment, and storage medium. Background Technology
[0002] With the development of science, neural networks have achieved excellent results in the field of image recognition in recent years. For example, neural networks can be applied to the medical field to recognize medical images.
[0003] However, in the medical field, medical personnel usually make diagnoses based on medical images combined with their own experience. If their experience is insufficient, misdiagnosis may occur. If medical images are identified through neural network models, they still need to rely on the experience of medical personnel in some aspects, making it difficult to guarantee the recognition effect and efficiency of the model. Therefore, improving the efficiency and accuracy of medical images has become an urgent problem to be solved. Summary of the Invention
[0004] This application provides a medical image processing method, apparatus, computer equipment, and storage medium, which can effectively provide the health status represented by the medical image and the corresponding score, thereby improving the recognition effect and efficiency of medical images.
[0005] This application provides a medical image processing method, including:
[0006] An application to a medical image processing model, wherein the medical image processing model includes a saliency prediction sub-model, an attribution prediction sub-model, and a decision sub-model, the method includes:
[0007] The medical images are input into the trained medical image processing model;
[0008] The saliency prediction sub-model in the medical image processing model outputs a saliency image corresponding to the medical image; the saliency image is used to highlight the target objects in the medical image.
[0009] The attribution prediction sub-model in the medical image processing model outputs an attribution map of the medical image; the attribution map is used to represent the probability distribution of the contribution of each pixel of the input image to the prediction process; wherein, the prediction process is the process of obtaining an output image from the input image;
[0010] The saliency image and the attribution map are fused and then input into the decision sub-model to generate a rating attribution map corresponding to the medical image. The rating attribution map has a rating and a category. The rating is used to characterize the severity of the health state represented by the medical image, and the category is used to characterize the name of the stage of the health state.
[0011] This application also provides a medical image processing apparatus, including:
[0012] An apparatus for use in a medical image processing model, the medical image processing model comprising a saliency prediction sub-model, an attribution prediction sub-model, and a decision sub-model, the apparatus comprising:
[0013] An image input module is used to input medical images into the trained medical image processing model;
[0014] A saliency image acquisition module is used to output a saliency image corresponding to the medical image through the saliency prediction sub-model in the medical image processing model; the saliency image is used to highlight the target objects in the medical image;
[0015] The attribution graph acquisition module is used to output an attribution graph of the medical image through the attribution prediction sub-model in the medical image processing model; the attribution graph is used to represent the probability distribution of the contribution of each pixel of the input image to the realization of the prediction process; wherein, the prediction process is the process of obtaining an output image from the input image;
[0016] The rating attribution map generation module is used to fuse the saliency image and the attribution map and input them into the decision sub-model to generate a rating attribution map corresponding to the medical image. The rating attribution map has a rating and a category. The rating is used to characterize the severity of the health state represented by the medical image, and the category is used to characterize the name of the stage of the health state.
[0017] This application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-described medical image processing method.
[0018] This application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described medical image processing method.
[0019] In this embodiment, a medical image can be input into a trained medical image processing model. This model first generates a fused image of the saliency image and attribution map corresponding to the medical image, and then outputs a scoring attribution map corresponding to the medical image based on the fused image, thus obtaining the health status and score represented by the medical image. Through this processing, a deeper interpretation of the representational information of the target region in the medical image can be achieved, thereby improving the efficiency and effectiveness of medical image recognition. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a schematic diagram of a scenario for the medical image processing method provided in the embodiments of this application;
[0022] Figure 2 This is a schematic flowchart of the medical image processing method provided in the embodiments of this application;
[0023] Figure 3 This is a schematic diagram of the medical image processing model provided in the embodiments of this application recognizing medical images;
[0024] Figure 4 This is a schematic diagram of the training medical image processing model provided in the embodiments of this application;
[0025] Figure 5 This is a schematic diagram of the structure of the medical image processing device provided in the embodiments of this application;
[0026] Figure 6 This is a schematic diagram of the computer device provided in the embodiments of this application. Detailed Implementation
[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0028] The medical image processing method provided in this application embodiment can be applied to, for example... Figure 1 In this application environment, a fused image of the saliency image and attribution map corresponding to the medical image is first generated. Then, based on the fused image, a scoring attribution map corresponding to the medical image is output, yielding the health status and score represented by the medical image. Through this processing, a deeper interpretation of the representational information of the target region in the medical image can be achieved, thereby improving the efficiency and effectiveness of medical image recognition. The user end can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server end can be implemented using a standalone server or a server cluster consisting of multiple servers.
[0029] In one embodiment, such as Figure 2As shown, a medical image processing method is provided, which is applied to... Figure 1 Taking the server-side as an example, the specific steps are as follows:
[0030] 201. Input the medical image into the trained medical image processing model.
[0031] Medical images refer to images that interact with the human body through a medium (such as X-rays, electromagnetic fields, and ultrasound) to represent the structure and density of internal tissues and organs. Examples include CT images, MRI images, and ultrasound images. Medical images can include target areas used to characterize the location and extent of lesions. For example, a target area could be a shadowed area on a lung CT image. The larger the area of the shadowed area and the closer the newly added shadowed area is to the current moment, the more severe the patient's lesion is.
[0032] The medical image processing model may include a saliency prediction sub-model, an attribution prediction sub-model, and a decision sub-model. These three sub-models together constitute the medical image processing model, and can be a neural network model, such as a convolutional neural network model. In this embodiment, the attribution prediction sub-model may be an improved version of the CAM (Class Activation Mapping) model, namely the F-CALM (Focus-Class Activation Latent Mapping) model. The difference lies in that the F-CALM model corresponding to the attribution prediction sub-model, based on the CAM model, introduces the EM algorithm during the loss function construction process to provide a deeper interpretation of the contribution distribution of each pixel in the image through maximum likelihood estimation.
[0033] Prior to step 201, the method of this application embodiment may further include:
[0034] Train the initial saliency prediction sub-model and the initial attribution prediction sub-model to obtain the saliency prediction sub-model and the attribution prediction sub-model after training;
[0035] The trained medical image processing model is obtained by utilizing the saliency prediction sub-model, the attribution prediction sub-model, and the decision sub-model.
[0036] like Figure 4 As shown, Figure 4This is a schematic diagram illustrating the training of a medical image processing model provided in an embodiment of this application. This embodiment trains the medical image processing model using reinforcement learning. The decision sub-model can be configured as a reinforcement learning agent to make the final decision on the recognition of medical images. The saliency prediction sub-model and the attribution prediction sub-model can be configured as a reinforcement learning environment for predicting and processing medical images.
[0037] It is understandable that, since the saliency prediction sub-model and the attribution prediction sub-model are environmental factors in reinforcement learning, their intermediate predictions will influence the final decision of the decision sub-model. Specifically, one of the initial saliency prediction sub-model and the initial attribution prediction sub-model can be trained first. After one model is trained, the other model is trained, resulting in fully trained saliency prediction sub-models and attribution prediction sub-models. During training, the output of the reinforcement learning agent under the maximum reward mechanism can be used as a supervision signal; that is, the optimal output of the decision sub-model can be used as the label of the training sample for subsequent training of the saliency prediction sub-model and the attribution prediction sub-model.
[0038] Optionally, the step "training the initial saliency prediction sub-model and the initial attribution prediction sub-model to obtain the trained saliency prediction sub-model and the attribution prediction sub-model" may include:
[0039] Obtain training samples;
[0040] The training samples are input into the initial medical image processing model to obtain the first loss function of the initial attribution prediction model and the second loss function of the initial saliency prediction model.
[0041] Repeat the above steps until the first loss function and the second loss function converge, thus obtaining the trained medical image processing model.
[0042] The training samples may include at least one first medical image with a score and at least one second medical image with a salient region. The first medical image and the second medical image are the same as the original image corresponding to them. The score of the first medical image and the salient region of the second image are determined by the decision sub-model.
[0043] It can be understood that the first medical image and the second medical image can be the same original image with different labels. Based on their respective label properties, the first medical image can be used to train the initial attribution prediction sub-model, and the second medical image can be used to train the initial saliency prediction sub-model. When the first loss function and the second loss function converge during the training process, it indicates that the initial saliency prediction sub-model and the initial attribution prediction sub-model have been trained successfully.
[0044] It should be noted that the labels of the first and second medical images can be generated by the output of the decision sub-model or obtained by manual annotation. This embodiment does not impose any restrictions.
[0045] Optionally, the step "obtaining the first loss function of the initial attribution prediction model and obtaining the second loss function of the initial saliency prediction model" may include:
[0046] The first medical image is input into the initial attribution prediction sub-model to obtain the predicted attribution map of the first medical image;
[0047] Based on the predicted attribution map and the true attribution map of the first medical image, obtain the first loss function of the initial attribution prediction sub-model;
[0048] The second medical image is input into the initial saliency prediction sub-model to obtain the predicted saliency image of the second medical image;
[0049] Based on the predicted saliency image and the true saliency image of the first medical image, the second loss function of the initial saliency prediction sub-model is obtained.
[0050] It is understandable that the first loss function is used to characterize the difference between the predicted attribution map and the true attribution map, and the second loss function is used to characterize the difference between the predicted saliency image and the true saliency image. Based on this, the performance of the model is tested, and the model is iteratively trained until the performance of the two models reaches the optimal level.
[0051] The first loss function may include a region loss function and an area loss function. The region loss function is used to characterize the accuracy of the location of the pixels included in the contribution probability distribution, and the area loss function is used to characterize the accuracy of the area of the pixels included in the contribution probability distribution.
[0052] It can be understood that the attribution map is a heatmap that represents the degree of influence of each pixel in a medical image on the medical image. In other words, the heat on the attribution map can be understood as the degree of contribution of each pixel in the input medical image to the model's prediction process.
[0053] The first loss function can include a region loss function and an area loss function. The region loss function characterizes the accuracy of the location of pixels included in the contribution probability distribution, and the area loss function characterizes the accuracy of the area of pixels included in the contribution probability distribution. It can be understood that the first loss function can be constructed for the latent prediction sub-model from these two aspects.
[0054] Optionally, the step "obtaining the first loss function of the initial attribution prediction sub-model" may include:
[0055] Based on the first predicted value and the EM algorithm, the region loss function is constructed; the first predicted value is the first probability distribution of the location of the pixels contained in the first medical image by the initial attribution prediction model for the known classification.
[0056] The area loss function is constructed based on the second predicted value; the second predicted value is the sum of all probabilities contained in the first probability distribution.
[0057] Obtain the sum of the region loss function and the area loss function, and use the sum as the value of the first loss function.
[0058] Specifically, the latent variables used in the EM algorithm can be determined first. The position Z of each pixel in the input medical image is taken as the dependent variable. The EM algorithm is then used to enable the medical image processing model to learn the conditional probability distribution of Z, i.e., p(zly), under the premise of the known classification result Y. The classification result can be the category of the rating attribution map.
[0059] Specifically, for an image X of size H×W, assume that the number of categories in image X is C, and that Z∈{1,2,…,HW}. Here, H is the height of each pixel in image X, and W is the width of each pixel.
[0060] The image X is passed through two convolutional layers to change the number of channels in the feature map extracted by the classification network to C and 1 respectively. The feature map of size C×H×W is then activated along each channel. For example, activation using the softmax function can yield the following activation values:
[0061] g yz =p(y|x,z)
[0062] Activating a feature map of size 1×H×W, for example, by using L1 normalization, can yield the following activation values:
[0063] h z =p(z|x)
[0064] h after broadcast mechanism processing z With g yz Element-wise multiplication yields the joint probability distribution of category y and location z predicted based on the input medical image x:
[0065] p1 = p(y, z|x)
[0066] Furthermore, based on formula (3), the probability distribution of the correct attribution map output by the initial attribution prediction sub-model can be derived, that is, the first predicted value can be expressed as:
[0067]
[0068] in, The true category corresponding to the input medical image x.
[0069] Specifically, let p θ Let ′(z|x,y) be the probability distribution of each pixel position z given that the classification result of the input medical image x is y. Then, according to the EM algorithm, the region loss function is:
[0070]
[0071] Where θ represents the parameters of the initial attribution prediction sub-model, and θ′ represents the distribution parameters of z.
[0072] By processing equation (5) according to Bayes' theorem, we can obtain:
[0073]
[0074] Where l is the position of a pixel in the position set Z.
[0075] Therefore, the final region loss function is:
[0076]
[0077] It is understandable that the region loss function can be used to characterize the accuracy of the location of pixels included in the probability distribution of contribution, and the area loss function can be used to characterize the accuracy of the area of pixels included in the probability distribution of contribution. Therefore, by adding an area loss function on top of the region loss function, and using the convergence of both loss functions as the convergence condition of the loss function of the initial attribution prediction sub-model, we can not only focus on the accuracy of the location of each pixel in the attribution map, but also further control the accuracy of the area formed by these pixels. This makes the location of each pixel and the area formed by all pixels in the heatmap displayed by the attribution map more accurate, and makes the focus of high-heat regions in the attribution map stronger.
[0078] The area loss function can be constructed based on the second predicted value; the second predicted value can be the sum of all probabilities contained in the first probability distribution, specifically expressed as:
[0079]
[0080] The first loss function can be the sum of the region loss function and the area loss function, specifically expressed as:
[0081] L = L EM +λL area
[0082] Where λ represents a preset hyperparameter used to control the area loss function L. area The proportion in the first loss function L.
[0083] In some embodiments, a second loss function for the initial saliency prediction sub-model can be obtained based on the difference between the predicted saliency image and the true saliency image of the first medical image. For example, as the initial saliency prediction sub-model is iteratively trained, for any stage in the training phase, the similarity distance between the two images can be obtained as the second loss function. When the similarity distance between the two reaches its minimum or zero, it indicates that the second loss function has converged, and the training of the initial saliency prediction sub-model is complete.
[0084] Step 202: Output the saliency image corresponding to the medical image through the saliency prediction sub-model in the medical image processing model.
[0085] The saliency image is used to highlight targets in the medical image. The target can refer to a target identification object; in a medical image, this could be a target area, a lesion area, an unidentified shadow, or other objects used to characterize medical image information. The saliency image effectively displays the location of the target. For example, to identify a target area in a medical image, the saliency area can highlight the target area, while the brightness of non-target areas remains unchanged or is reduced, thus clearly identifying the contour and extent of the target area through the saliency image.
[0086] Step 203: Output the attribution map of the medical image through the attribution prediction sub-model in the medical image processing model.
[0087] The attribution graph is used to represent the probability distribution of the contribution of each pixel of the input image to the prediction process; and the prediction process is the process of obtaining an output image from the input image.
[0088] It is understood that the attribution prediction sub-model in this application embodiment can be an improved F-CALM model based on the CAM model. The difference lies in that the F-CALM model corresponding to the attribution prediction sub-model introduces the EM algorithm in the process of constructing the loss function, based on the CAM model, to provide a deeper interpretation of the contribution distribution of each pixel in the image through maximum likelihood estimation. Therefore, the attribution map has a deeper interpretability than the traditional CAM map. It can not only reflect the contribution distribution of each pixel in the image to the displayed heat of the attribution map through the heat of each pixel in the heat map, but also has stronger focusing, making the heat area of different depth regions in the attribution map more accurate.
[0089] Step 204: After fusing the saliency image and the attribution map, input them into the decision sub-model to generate the scoring attribution map corresponding to the medical image.
[0090] The rating attribution map has a rating and a category. The rating is used to characterize the severity of the health state represented by the medical image, and the category is used to characterize the name of the stage of the health state.
[0091] Here, health status refers to the user's current health condition corresponding to the medical image. Health status can be divided into multiple levels, such as healthy, fair, sub-healthy, and ill. The levels of health status can also be represented by numbers, such as level 1, level 2, level 3, etc., with each level representing a different health state. Optionally, the health status levels and scores can be set in descending order from most severe to least severe.
[0092] Optionally, different types of medical images can correspond to different health states. Taking CT or MRI images as examples, the severity of the health state can be determined by the distribution, area, and generation time of lesions in the medical image. For instance, the more numerous, concentrated, and larger the lesions in a medical image, the worse the user's health state. Furthermore, the closer the lesions were identified to the current time, the worse the symptoms have become and the worse the health state. Optionally, in the scoring attribution map output by the medical image recognition model, the darker areas in the heatmap can focus on the lesions. The depth and area of the heatmap can also reflect the severity of the lesions, allowing the scoring attribution map to determine the classification of the health state based on these features, i.e., to obtain the level of health state, and thus determine the score of the scoring attribution map.
[0093] Taking medical images, which are images of the human body surface, as an example, such as facial images, a medical image processing model can be used to focus on the key pixel locations of the facial image and calculate the probability distribution of the contribution of each pixel to give a specific score. For example, the heat depth in the scoring attribution map can focus on features that lead to poor skin quality, such as wrinkles, dark circles, and large pores. Finally, by combining the heat depth, distribution, and quantity features in the heat map, the health status is classified to determine the corresponding level, and the final score of the scoring attribution map is determined based on the level.
[0094] Prior to step 204, the method of this application embodiment may further include:
[0095] Calculate the first weight and the second weight;
[0096] The step of fusing the saliency image and the attribution map and then inputting them into the decision sub-model includes:
[0097] Based on the first weight and the second weight, the saliency image and the attribution map are fused to obtain a fusion result;
[0098] The fusion result is input into the decision sub-model.
[0099] Optionally, the step "calculating the first weight and the second weight" may further include:
[0100] Obtain a score from preset parameters and a third medical image; the third medical image is any one of the medical images.
[0101] Obtain the minimum value of the product of the preset parameter and the score of the third medical image, and use the minimum value as the first weight for fusing the saliency image;
[0102] Obtain the difference between the preset value and the first weight, and use the difference as the second weight for fusing the attribution graph.
[0103] like Figure 3As shown, before the input image enters the decision sub-model, the saliency image and attribution map need to be fused and pooled before being input into the decision sub-model. To ensure the rationality of the fused image as a fused feature, the weights of the two image fusions need to be pre-set, including a first weight for saliency prediction and a second weight for the attribution map. Optionally, in this embodiment, the weights of top-down attention and bottom-up attention for the input image are controlled by the score of the attribution map. Specifically, the saliency prediction sub-model can output the saliency image through top-down attention, and the attribution prediction sub-model can output the attribution map through a bottom-up attention mechanism.
[0104] The first weight can be set to Where m is a preset hyperparameter, This represents the score output by the decision sub-model under the condition of maximizing reward. It can be understood that since the first and second weights are pre-set, their values can be adjusted and optimized during training to obtain the optimal solutions for the first and second weights.
[0105] It can be understood that for the fused image, the sum of the first and second weights is 1, therefore the second weight is 1-ρ. Thus, the mapping value of the fused image can be expressed as:
[0106] S=(1-ρ)S bu +ρS td
[0107] Here, bu represents front-to-back attention, and td represents top-to-bottom attention.
[0108] As can be seen from the above, in this embodiment of the application, a medical image can be input into a trained medical image processing model. This model first generates a fused image of the saliency image and attribution map corresponding to the medical image, and then outputs a scoring attribution map corresponding to the medical image based on the fused image, thus obtaining the health status and score represented by the medical image. Through the above processing, a deeper interpretation of the representational information of the target region in the medical image can be achieved, thereby improving the recognition efficiency and effectiveness of the medical image.
[0109] To better implement the above methods, this application also provides a medical image processing apparatus, which corresponds one-to-one with the medical image processing methods described above. For example... Figure 5As shown, this medical image processing device is applied to a medical image processing model, which includes a saliency prediction sub-model, an attribution prediction sub-model, and a decision sub-model. The device includes an image input module 301, a saliency image acquisition module 302, an attribution map acquisition module 303, and a score attribution map generation module 304. Detailed descriptions of each functional module are as follows:
[0110] Image input module 301 is used to input medical images into the trained medical image processing model;
[0111] The saliency image acquisition module 302 is used to output a saliency image corresponding to the medical image through the saliency prediction sub-model in the medical image processing model; the saliency image is used to highlight the target objects in the medical image;
[0112] The attribution graph acquisition module 303 is used to output the attribution graph of the medical image through the attribution prediction sub-model in the medical image processing model; the attribution graph is used to represent the probability distribution of the contribution of each pixel of the input image to the realization of the prediction process; wherein, the prediction process is the process of obtaining the output image from the input image;
[0113] The rating attribution map generation module 304 is used to fuse the saliency image and the attribution map and input them into the decision sub-model to generate a rating attribution map corresponding to the medical image; the rating attribution map has a rating and a category, the rating is used to characterize the severity of the health state represented by the medical image, and the category is used to characterize the name of the stage of the health state.
[0114] Specific limitations regarding the medical image processing device can be found in the limitations of the medical image processing method described above, and will not be repeated here. Each module in the aforementioned medical image processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.
[0115] Therefore, in this embodiment, a medical image can be input into a trained medical image processing model. This model first generates a fused image of the saliency image and attribution map corresponding to the medical image, and then outputs a scoring attribution map corresponding to the medical image based on the fused image, thus obtaining the health status and score represented by the medical image. Through this processing, a deeper interpretation of the representational information of the target region in the medical image can be achieved, thereby improving the efficiency and effectiveness of medical image recognition.
[0116] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores data generated or acquired during the medical image processing process, such as medical image processing models. The network interface communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements a medical image processing method.
[0117] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the medical image processing method described in the above embodiments, for example, Figure 2 Steps 201 to 204 are shown. Alternatively, when the computer program is executed by the processor, it implements the functions of each module in the medical image processing device described above, for example... Figure 5 The functions of modules 301 to 304 are shown. To avoid repetition, they will not be described again here.
[0118] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, the computer program implements the medical image processing method described in the above-described method embodiment. For example, Figure 2 Steps 201 to 204 are shown. Alternatively, when the computer program is executed by the processor, it implements the functions of each module in the medical image processing device described above, for example... Figure 4 The functions of modules 301 to 304 are shown. To avoid repetition, they will not be described again here.
[0119] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0120] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0121] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A medical image processing method, characterized in that, An application is made to a medical image processing model, which includes a saliency prediction sub-model, an attribution prediction sub-model, and a decision sub-model. The saliency prediction sub-model and the attribution prediction sub-model are both neural network models, and the decision sub-model is a reinforcement learning-based agent used to make decisions regarding the recognition of medical images. The method includes: The medical images are input into the trained medical image processing model; The saliency prediction sub-model in the medical image processing model outputs a saliency image corresponding to the medical image; the saliency image is used to highlight the target objects in the medical image. The attribution prediction sub-model in the medical image processing model outputs an attribution map of the medical image; the attribution map is used to represent the probability distribution of the contribution of each pixel of the input image to the prediction process; wherein, the prediction process is the process of obtaining an output image from the input image; The saliency image and the attribution map are fused and then input into the decision sub-model to generate a rating attribution map corresponding to the medical image. The rating attribution map has a rating and a category. The rating is used to characterize the severity of the health state represented by the medical image, and the category is used to characterize the name of the stage of the health state. The method further includes, before inputting the medical image into the trained medical image processing model, the following steps: obtaining training samples; the training samples include at least one first medical image with a score, and at least one second medical image with a salient region, wherein the original image corresponding to the first medical image and the second medical image are identical, and the score of the first medical image and the salient region of the second image are determined by the decision sub-model; inputting the training samples into the initial medical image processing model, inputting the first medical image into the initial attribution prediction sub-model, and obtaining a predicted attribution map of the first medical image; and based on the predicted attribution map and the first medical image... The true attribution graph is used to obtain the first loss function of the initial attribution prediction sub-model; the second medical image is input into the initial saliency prediction sub-model to obtain the predicted saliency image of the second medical image; based on the predicted saliency image and the true saliency image of the first medical image, the second loss function of the initial saliency prediction sub-model is obtained; when the first loss function and the second loss function converge, the trained saliency prediction sub-model and the attribution prediction sub-model are obtained; the trained saliency prediction sub-model, the attribution prediction sub-model, and the decision sub-model are used to obtain the trained medical image processing model.
2. The medical image processing method as described in claim 1, characterized in that, The first loss function includes a region loss function and an area loss function. The region loss function is used to characterize the accuracy of the location of pixels included in the contribution probability distribution, and the area loss function is used to characterize the accuracy of the area of pixels included in the contribution probability distribution. Obtaining the first loss function of the initial attribution prediction sub-model includes: Based on the first predicted value and the EM algorithm, the region loss function is constructed; the first predicted value is the first probability distribution of the location of the pixels contained in the first medical image by the initial attribution prediction model for the known classification. The area loss function is constructed based on the second predicted value; the second predicted value is the sum of all probabilities contained in the first probability distribution. Obtain the sum of the region loss function and the area loss function, and use the sum as the value of the first loss function.
3. The medical image processing method as described in claim 1, characterized in that, Before the saliency image and the attribution map are fused and input into the decision sub-model, the following steps are included: Calculate the first weight and the second weight; The step of fusing the saliency image and the attribution map and then inputting them into the decision sub-model includes: Based on the first weight and the second weight, the saliency image and the attribution map are fused to obtain a fusion result; The fusion result is input into the decision sub-model.
4. The medical image processing method as described in claim 3, characterized in that, The calculation of the first weight and the second weight includes: Obtain a score from preset parameters and a third medical image; the third medical image is any one of the medical images. Obtain the minimum value of the product of the preset parameter and the score of the third medical image, and use the minimum value as the first weight for fusing the saliency image; Obtain the difference between the preset value and the first weight, and use the difference as the second weight for fusing the attribution graph.
5. A medical image processing device, characterized in that, An application is made to a medical image processing model, which includes a saliency prediction sub-model, an attribution prediction sub-model, and a decision sub-model. The saliency prediction sub-model and the attribution prediction sub-model are both neural network models, and the decision sub-model is a reinforcement learning-based agent used to make decisions for the recognition of medical images. The device includes: An image input module is used to input medical images into the trained medical image processing model; A saliency image acquisition module is used to output a saliency image corresponding to the medical image through the saliency prediction sub-model in the medical image processing model; the saliency image is used to highlight the target objects in the medical image; The attribution graph acquisition module is used to output an attribution graph of the medical image through the attribution prediction sub-model in the medical image processing model; the attribution graph is used to represent the probability distribution of the contribution of each pixel of the input image to the realization of the prediction process; wherein, the prediction process is the process of obtaining an output image from the input image; The rating attribution map generation module is used to fuse the saliency image and the attribution map and input them into the decision sub-model to generate a rating attribution map corresponding to the medical image; the rating attribution map has a rating and a category, the rating is used to characterize the severity of the health state represented by the medical image, and the category is used to characterize the name of the stage of the health state; The method further includes, before inputting the medical image into the trained medical image processing model, the following steps: obtaining training samples; the training samples include at least one first medical image with a score, and at least one second medical image with a salient region, wherein the original image corresponding to the first medical image and the second medical image are identical, and the score of the first medical image and the salient region of the second image are determined by the decision sub-model; inputting the training samples into the initial medical image processing model, inputting the first medical image into the initial attribution prediction sub-model, and obtaining a predicted attribution map of the first medical image; and based on the predicted attribution map and the first medical image... The true attribution graph is used to obtain the first loss function of the initial attribution prediction sub-model; the second medical image is input into the initial saliency prediction sub-model to obtain the predicted saliency image of the second medical image; based on the predicted saliency image and the true saliency image of the first medical image, the second loss function of the initial saliency prediction sub-model is obtained; when the first loss function and the second loss function converge, the trained saliency prediction sub-model and the attribution prediction sub-model are obtained; the trained saliency prediction sub-model, the attribution prediction sub-model, and the decision sub-model are used to obtain the trained medical image processing model.
6. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the medical image processing method as described in any one of claims 1 to 4.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the medical image processing method as described in any one of claims 1 to 4.