Medical image interpretable classification prediction method, system, equipment and medium
By performing concept label transformation and image preprocessing on medical imaging data, and inputting the classification prediction network built by multimodal learning module, concept optimization module and classification module for training, the limitations of medical imaging concept recognition and interpretation are solved and the accuracy of classification prediction is improved.
Patent Information
- Application Number
- CN202510078518.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-05-16
AI Technical Summary
The prior art has limitations in the concept identification and interpretation of medical images, and it is difficult to effectively identify and label multiple medical concepts that may overlap or be associated, which affects the accuracy of prediction.
A medical image interpretable classification prediction method is proposed. The classification prediction model is obtained by obtaining medical data sets, performing concept label transformation and image preprocessing, and inputting data into a pre-constructed classification prediction network for training. The network includes a multimodal learning module, a concept optimization module and a classification module, which is used to extract and optimize conceptual features and improve the accuracy of classification prediction.
By extracting and optimizing concept features, reducing interfering information, maximizing concept differences, the model's ability to learn concepts and the accuracy of classification prediction results is improved.
Smart Images

Figure CN120014335A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, system, device and medium for interpretable classification prediction of medical images. Background Art
[0002] In related technologies, deep learning models have shown high accuracy in judging medical images, but they still have significant limitations in the recognition and specific interpretation of concepts. In practical applications, it is found that concept recognition often involves complex multi-label classification tasks, and it is difficult to identify and label multiple medical concepts that may overlap or be related, which makes it difficult to truly identify interpretable classification predictions of medical images, affecting the accuracy of the predictions. In summary, the technical problems existing in related technologies need to be improved. Summary of the invention
[0003] The main purpose of the embodiments of the present application is to propose a method, system, device and medium for interpretable classification prediction of medical images, which can improve the accuracy of classification prediction.
[0004] To achieve the above-mentioned purpose, an embodiment of the present application provides a method for interpretable classification prediction of medical images, the method comprising:
[0005] Access to medical datasets;
[0006] Performing concept label conversion and image preprocessing on the medical data set to obtain a training data set;
[0007] Inputting the training data set into a pre-built classification prediction network for training processing to obtain a classification prediction model; the classification prediction network includes a multimodal learning module, a concept optimization module and a classification module;
[0008] The medical image to be processed is input into the classification prediction model for classification prediction processing to obtain a classification prediction result.
[0009] In some embodiments, the step of inputting the training data set into a pre-built classification prediction network for training processing to obtain a classification prediction model comprises the following steps:
[0010] Inputting the training data set into the multimodal learning module for feature extraction processing to obtain an initial concept feature map; the initial concept feature map is used to characterize the response intensity of the concept label at the image space position;
[0011] Inputting the initial concept feature graph into the concept optimization module for whitening and gradient optimization to obtain a target concept feature graph;
[0012] Inputting the target concept feature graph into the classification module for classification processing to obtain a classification result;
[0013] Perform loss calculation processing on the classification result according to the classification prediction loss function to obtain a classification prediction loss value;
[0014] The classification prediction network is subjected to parameter update processing according to the classification prediction loss value to obtain the classification prediction model.
[0015] In some embodiments, the step of inputting the training data set into the multimodal learning module for feature extraction to obtain an initial concept feature map comprises the following steps:
[0016] Inputting the training data set into the multimodal learning module for network training processing to obtain a feature extraction layer;
[0017] Inputting the training data set into the feature extraction layer for feature extraction processing to obtain a high-level feature map;
[0018] The high-level feature map is converted through a convolutional layer to obtain the initial concept feature map.
[0019] In some embodiments, the multimodal learning module includes a backbone network and a text encoder, and the inputting of the training data set into the multimodal learning module for network training processing to obtain a feature extraction layer includes the following steps:
[0020] Inputting the training data set into the backbone network and the text encoder for feature extraction processing to obtain image representation and text representation;
[0021] Performing linear mapping conversion processing and contrast loss calculation processing on the image representation and the text representation to obtain a contrast loss value;
[0022] The parameters of the backbone network are updated according to the contrast loss value to obtain the feature extraction layer.
[0023] In some embodiments, the step of inputting the initial concept feature graph into the concept optimization module for whitening and gradient optimization to obtain a target concept feature graph comprises the following steps:
[0024] Performing iterative whitening processing on the initial concept feature map to obtain a whitening matrix;
[0025] Constructing an initial orthogonal matrix according to the output result of the multimodal learning module;
[0026] Performing iterative optimization processing on the initial orthogonal matrix to obtain an optimized orthogonal matrix;
[0027] The whitening matrix is subjected to feature space orthogonalization processing according to the optimized orthogonal matrix to obtain the target concept feature map.
[0028] In some embodiments, the iterative whitening process is performed on the initial concept feature map to obtain a whitening matrix, comprising the following steps:
[0029] Decentralizing the initial concept feature graph to obtain an initial matrix;
[0030] Performing covariance calculation processing on the initial matrix to obtain a covariance matrix;
[0031] Initializing the initial matrix to obtain a unit matrix;
[0032] Perform Newton iterative calculation processing on the unit matrix according to the covariance matrix to obtain an iterative matrix;
[0033] The whitening matrix is obtained by performing calculation processing according to the iteration matrix and the covariance matrix.
[0034] In some embodiments, the iterative optimization process of the initial orthogonal matrix to obtain an optimized orthogonal matrix comprises the following steps:
[0035] Performing binarization calculation processing on the initial concept feature map to obtain a concept segmentation map;
[0036] Performing maximization regularization processing on the initial orthogonal matrix according to the concept segmentation graph to construct a target optimization function;
[0037] The target optimization function is iteratively calculated according to a conjugate gradient search algorithm to obtain the optimized orthogonal matrix.
[0038] To achieve the above-mentioned purpose, another aspect of the embodiment of the present application provides a medical image interpretable classification prediction system, the system comprising:
[0039] The first module is used to obtain medical data sets;
[0040] The second module is used to perform concept label conversion and image preprocessing on the medical data set to obtain a training data set;
[0041] The third module is used to input the training data set into a pre-built classification prediction network for training processing to obtain a classification prediction model; the classification prediction network includes a multimodal learning module, a concept optimization module and a classification module;
[0042] The fourth module is used to input the medical image to be processed into the classification prediction model for classification prediction processing to obtain the classification prediction result.
[0043] To achieve the above-mentioned purpose, another aspect of an embodiment of the present application provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the above-mentioned method when executing the computer program.
[0044] To achieve the above objective, another aspect of an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described above is implemented.
[0045] The embodiments of the present application include at least the following beneficial effects: The embodiments of the present application provide a method, system, device and medium for interpretable classification prediction of medical images. The scheme obtains a medical data set, performs concept label conversion and image preprocessing on the medical data set to obtain a training data set, and can convert the text data in the medical data set into concept labels for subsequent training and processing of the model, thereby improving the model's ability to learn concepts. In addition, the embodiments of the present application input the training data set into a pre-built classification prediction network for training processing to obtain a classification prediction model, and input the medical image to be processed into the classification prediction model for classification prediction processing to obtain a classification prediction result. The embodiments of the present application can extract the principal components of the extracted features through the concept optimization module in the classification prediction network, reduce interference information, and maximize the concept difference so that the concept features learned by the model are clear and accurate, thereby improving the accuracy of the classification prediction results. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 It is a flowchart of a medical image interpretable classification prediction method provided by an embodiment of the present application;
[0047] Figure 2 It is a structural diagram of a multimodal learning module provided in an embodiment of the present application;
[0048] Figure 3 It is a structural diagram of a classification prediction network provided in an embodiment of the present application;
[0049] Figure 4 It is a structural schematic diagram of a medical image interpretable classification prediction system provided in an embodiment of the present application;
[0050] Figure 5 It is a schematic diagram of the hardware structure of the electronic device provided in the embodiment of the present application. DETAILED DESCRIPTION
[0051] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in conjunction with the accompanying drawings and examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of the present application. They are only examples of systems and methods consistent with some aspects of the embodiments of the present application as detailed in the attached claims.
[0052] It is understood that the terms "first", "second", etc. used in this application can be used to describe various concepts in this article, but unless otherwise specified, these concepts are not limited by these terms. These terms are only used to distinguish one concept from another concept. For example, without departing from the scope of the embodiment of the present application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the words "if" and "if" as used herein can be interpreted as "at the time of" or "when" or "in response to determination".
[0053] The terms "at least one", "multiple", "each", "any", etc. used in this application, at least one includes one, two or more, multiple includes two or more, each refers to each of the corresponding multiple, and any refers to any one of the multiple.
[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0055] Before describing the embodiments of the present application in detail, some nouns and terms involved in the embodiments of the present application are first described. The nouns and terms involved in the embodiments of the present application are subject to the following explanations.
[0056] 1) Convolutional Neural Network: Convolutional Neural Network is a deep feedforward neural network with local connections and weight sharing. It is one of the representative algorithms of deep learning and is good at processing images, especially image recognition and other related machine learning problems. For example, it has a significant improvement effect in various visual tasks such as image classification, target detection, and image segmentation. It is one of the most widely used models. It has the ability to represent learning, can classify input information according to its hierarchical structure in a translation-invariant manner, and can perform supervised learning and unsupervised learning. The convolution kernel parameter sharing in the hidden layer and the sparsity of the inter-layer connection enable the convolutional neural network to learn grid features such as pixels and audio with a small amount of computation, with stable results and no additional feature engineering requirements for data. It is widely used in computer vision, natural language processing and other fields.
[0057] 2) Explainable AI: Explainable AI is a set of processes and methods that allow human users to understand and trust the results and outputs created by machine learning algorithms. Explainable AI is used to explain AI models, their intended impacts, and potential biases. It helps describe the accuracy, fairness, transparency, and results of AI-driven decisions. Explainable AI is critical to help organizations build trust and confidence when they put AI models into production. AI explainability also helps organizations adopt responsible AI development methods. As AI becomes more and more advanced, it has become difficult for humans to understand and trace how algorithms reach their results. The entire computing process has become what is commonly known as an unexplainable "black box." These black box models are created directly from data. Moreover, even the engineers or data scientists who created the algorithms cannot understand or explain what exactly happened inside these algorithms, or how the AI algorithms reached a specific result. Understanding how an AI-enabled system produces a specific output has many benefits. Explainability helps developers ensure that the system operates as expected, and may also be required to meet regulatory standards, or if you want to allow people affected by the decision to question or change the results, then explainability is also very important.
[0058] 3) Multimodal Learning: Multimodal learning is a method in the field of machine learning that involves associating information from multiple modalities (each source or form of information can be defined as a modality, such as images and text), aiming to fuse multiple types of data to help the model obtain more comprehensive knowledge and more accurate results. The key methods of multimodal learning are data alignment, modal fusion, cross-modal attention mechanism, and collaborative learning, which can be applied to many application scenarios such as image generation and text description, video understanding, etc. In multimodal learning, the extraction and integration of information from different modalities can help the model better understand the context or feature relationship when dealing with complex tasks, thereby reducing the limitations of single modality information, improving the accuracy, robustness and breadth of application of the model, and helping to achieve smarter and more efficient AI applications.
[0059] 4) Concept-Based Learning: Concept-Based Learning is a learning method that uses artificially defined concept features as intermediate or final outputs. In this learning method, the model analyzes the input data to abstract some special feature concepts, and obtains the final prediction results based on these concepts. This method is usually used more in the medical field. The results of existing models are becoming more and more complex. It is difficult for people to understand the machine's algorithms, and it is difficult for medical personnel to directly believe the results of the model output. People often need additional, more reliable intermediate or final results to make additional judgments. The concept learning method is a learning method that abstracts these data that have a basis for the final result, uses them as the final result or the main judgment basis for the final result, and presents more detailed prediction basis.
[0060] 5) ZCA whitening (Zero-phase Component Analysis Whitening): ZCA whitening is an algorithm used for image processing and data preprocessing. It achieves whitening by converting image data into its principal component space and then reprojecting these principal components back to the original space. The main purpose of ZCA whitening is to retain the main features of the image while reducing the variance in the data. The specific steps of ZCA whitening are divided into the following parts: 1. Standardize the data, 2. Calculate the covariance matrix, 3. Find the eigenvalues and eigenvectors, 4. Select the principal components, and 5. Reproject the data. Compared with other whitening algorithms (such as PCA whitening and LDA whitening), ZCA whitening performs better in image processing, avoiding the color change of the image caused by PCA whitening, while retaining the main features of the image and reducing the variance in the data, retaining the spatial structure of the original data, and reducing redundant information. It plays a good role in accelerating data processing, improving the convergence speed of machine learning algorithms, and improving the stability of the algorithm.
[0061] 6) Conjugate Gradient Method: The conjugate gradient method is an optimization algorithm between the steepest descent method and the Newton method. For the optimization problem Φ(x) = 1 / 2x T Ax-b T Each search direction of the x-conjugate gradient algorithm is conjugate to each other, and the new direction used in each iteration is a linear combination of the negative residual and the previous search direction, so it requires less storage and is easy to calculate. At the same time, it also has the characteristics of step convergence and high stability. In theory, the optimal solution can be found within n steps. It is one of the most useful methods for solving large linear equations and one of the most effective algorithms for solving large nonlinear optimization.
[0062] In related technologies, traditional black box models, such as deep convolutional neural networks, can perform well in complex tasks, but because their internal working mechanisms are difficult to explain, it is difficult for doctors and medical staff to understand how the model draws conclusions. Although deep learning can surpass human recognition accuracy in some projects, deep neural networks lack interpretability and make judgments in one go, making it difficult for doctors to trust them. The characteristics of black box models also pose more security risks. Some studies have shown that using special methods for training may make the model perform normally most of the time, but when encountering some selected features that are not related to the model results, it will show completely opposite results. The characteristics of black box models make it difficult for regulators and users to find out whether the model has been attacked, so it is unsafe to use black box models directly for diagnosis.
[0063] The current method for interpretable classification prediction of medical images is not perfect when annotating data sets. The ideal concept explanation requires that each specific medical concept be accurately located in a specific area of the original medical image, that is, to provide clear spatial location annotations on the image. However, existing data sets usually only provide simplified text labels and lack accurate annotations of specific parts in the image. To extend the concept annotations to the exact location in the image, a large number of professionals are required to manually annotate, which is not only time-consuming, but also requires rich medical knowledge and image recognition experience, so it is difficult to achieve in practical applications.
[0064] In addition, although deep learning models have shown high accuracy in medical judgment, they still have significant limitations in identifying and explaining specific concepts. Concept recognition often involves complex multi-label classification tasks, which require the model to accurately diagnose the disease while being able to identify and label multiple medical concepts that may overlap or be related. It is still a very challenging engineering problem to allow the model to meet the standard in diagnostic accuracy while also being able to give a clear explanation and label the location of each concept. This not only requires the model to have a high level of understanding and detailed recognition capabilities, but also depends on the comprehensiveness and accuracy of the dataset annotation.
[0065] In view of this, an embodiment of the present application provides a method, system, device and medium for interpretable classification prediction of medical images. The scheme obtains a medical data set, converts the medical data set into concept labels and performs image preprocessing to obtain a training data set, and can convert the text data in the medical data set into concept labels for subsequent training and processing of the model, thereby improving the model's ability to learn concepts. In addition, the embodiment of the present application inputs the training data set into a pre-built classification prediction network for training processing to obtain a classification prediction model, and inputs the medical image to be processed into the classification prediction model for classification prediction processing to obtain a classification prediction result. The embodiment of the present application can extract the principal components of the extracted features through the concept optimization module in the classification prediction network, thereby reducing interference information, and maximizing the concept difference so that the concept features learned by the model are clear and accurate, thereby improving the accuracy of the classification prediction results.
[0066] The medical image interpretable classification prediction method provided in the embodiment of the present application relates to the field of artificial intelligence technology. The medical image interpretable classification prediction method provided in the embodiment of the present application can be applied to a terminal, can also be applied to a server, and can also be software running in a terminal or a server. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, and a car terminal, etc., but is not limited to this; the server side can be configured as an independent physical server, or it can be configured as a server cluster or distributed system composed of multiple physical servers, and can also be configured to provide cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms and other basic cloud computing services. The cloud server, the server can also be a node server in the blockchain network; the software can be an application that implements the medical image interpretable classification prediction method, etc., but is not limited to the above forms.
[0067] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in distributed computing environments, in which tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0068] It should be noted that in each specific implementation of the present application, when it comes to the need to perform relevant processing based on data related to user identity or characteristics such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first, and the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiment of the present application needs to obtain the user's sensitive personal information, the user's separate permission or consent will be obtained through a pop-up window or by jumping to a confirmation page. After clearly obtaining the user's separate permission or consent, the necessary user-related data for the normal operation of the embodiment of the present application will be obtained.
[0069] Figure 1 is an optional flowchart of the medical image interpretable classification prediction method provided in the embodiment of the present application. Figure 1 The method may include but is not limited to steps S101 to S104.
[0070] Step S101, obtaining a medical data set;
[0071] Step S102, performing concept label conversion and image preprocessing on the medical data set to obtain a training data set;
[0072] Step S103, inputting the training data set into a pre-built classification prediction network for training processing to obtain a classification prediction model; the classification prediction network includes a multimodal learning module, a concept optimization module and a classification module;
[0073] Step S104: input the medical image to be processed into the classification prediction model for classification prediction processing to obtain a classification prediction result.
[0074] Steps S101 to S104 shown in the embodiment of the present application, by acquiring a medical data set, the embodiment of the present application can use a medical data set with text concept annotations to train the model, by converting the text data in the medical data set into a concept label, and then preprocessing the image data in the medical data set to obtain a training data set. The classification prediction model is obtained by inputting the training data set into a pre-constructed classification prediction network for training processing, wherein the classification prediction network includes a multimodal learning module, a concept optimization module and a classification module, the multimodal learning module is used to extract data from the multimodal data in the training data, and the multimodal learning module includes a backbone network and a text encoder for multimodal learning of the data. The concept optimization module is used to optimize the concept feature of the feature data extracted by the multimodal learning module, and by maximizing the variance between concepts, the concept features learned by the model are clear and accurate. The classification module is used to perform classification prediction processing according to the concept features, and the classification module includes a global average pooling layer and a fully connected layer. The embodiment of the present application obtains a classification prediction model by training the classification prediction network, so that the medical image to be processed can be input into the classification prediction model for classification prediction to obtain a classification prediction result.
[0075] In step S101 of some embodiments, a medical data set may be obtained from a medical database, or relevant medical data sets may be collected or crawled from a medical system. A medical data set is image data and text data related to medical data. For example, a dermoscopic data set includes dermoscopic images, diagnosis results, some patient information, and some textual concept descriptions.
[0076] In step S102 of some embodiments, it is necessary to perform concept label conversion processing on the text data in the medical data set, and perform image preprocessing on the image data in the medical data set, and use the processed text data and image data as training data sets to train the classification prediction network. In order to obtain trainable information, the text data in the medical data set needs to be converted into digital labels that can be learned by the model. For example, for the vascular structure in the skin, there are tree-like, hairpin-like, comma-like, dot-like, and degenerative structures. According to the characteristics of these blood vessels, whether the vascular structure is regular can be used as a concept for learning. Similarly, conceptual information such as color and area can be converted into concept labels for learning. Since there is a large amount of missing information in the data, it is necessary to use whether a specific concept exists as a concept label.
[0077] Before the image data enters the model, data preprocessing is required. Since deep learning models usually require input data to be in floating point format, the original image data needs to be converted from its initial format, such as integer or different color channel formats, to floating point format to ensure that the data can be correctly read and analyzed by the model. At the same time, in order to ensure data consistency, all images must be resampled to the same size to meet the input size requirements of the model. This step is very critical because the resolution and size differences of the images will affect the training effect and prediction stability of the model. By resizing the images to a consistent size, the model can avoid training bias due to different input image sizes, thereby improving the generalization ability of the model. In addition, in order to further improve the accuracy of the model and reduce overfitting, a series of data enhancement operations need to be performed on the data. For example, multiple variants can be generated by rotation, random cropping, mirror flipping, etc. These operations not only enrich the diversity of the training data, but also make the model more robust when processing images with different postures and shooting angles. Data enhancement can effectively prevent the model from "remembering" the details of specific images in the training set, thereby showing better generalization ability on the test data.
[0078] In step S103 of some embodiments, the step of inputting the training data set into a pre-built classification prediction network for training to obtain a classification prediction model includes the following steps:
[0079] Inputting the training data set into the multimodal learning module for feature extraction processing to obtain an initial concept feature map; the initial concept feature map is used to characterize the response intensity of the concept label at the image space position;
[0080] Inputting the initial concept feature graph into the concept optimization module for whitening and gradient optimization to obtain a target concept feature graph;
[0081] Inputting the target concept feature graph into the classification module for classification processing to obtain a classification result;
[0082] Perform loss calculation processing on the classification result according to the classification prediction loss function to obtain a classification prediction loss value;
[0083] The classification prediction network is subjected to parameter update processing according to the classification prediction loss value to obtain the classification prediction model.
[0084] In an embodiment of the present application, the classification prediction network includes a multimodal learning module, a concept optimization module and a classification module, wherein the training data set is input into the multimodal learning module for feature extraction processing to obtain an initial concept feature map, which is used to characterize the response intensity of the concept label at the image space position. Among them, the multimodal learning module can be improved based on the single-classification deep neural network model structure to enhance the model's interpretation ability and location perception ability. The multimodal learning module includes a backbone network and a text encoder. The network structure of the backbone network can adopt the network structure of the residual neural network (ResNet), and can also adopt other deep network model structures.
[0085] In some embodiments, the embodiment of the present application inputs the initial concept feature map into the concept optimization module for whitening and gradient optimization processing to obtain a target concept feature map, which is an optimized concept feature map, wherein the whitening process uses ZCA whitening. Compared with other principal component analysis methods, ZCA whitening retains the feature distribution of the original image, so it can be used inside the model and will not cause the whitening output to change dramatically and fail to converge as the features change during training. In order to further increase the simple difference between concepts, the embodiment of the present application introduces an orthogonal matrix, and iteratively optimizes the orthogonal matrix through a gradient optimization algorithm, so that it can orthogonalize the feature components of the concept, increase the variance between concepts in the concept feature map, and thus obtain the final target concept map.
[0086] In some embodiments, the embodiment of the present application inputs the target concept feature map into the classification module for classification processing to obtain a classification result. Among them, the classification module includes a global average pooling layer and a fully connected layer. The global average pooling layer performs global average pooling on the target concept feature map to obtain the specific activation value of each concept, and then the fully connected layer summarizes these activation values, which are finally used for disease classification prediction. The embodiment of the present application performs global average pooling on the concept map G to obtain the average activation value of each concept. The calculation formula for the overall activation value of each concept is as follows:
[0087]
[0088] in, represents the overall activation value of the kth concept, H′ and W′ represent the spatial dimensions of the feature map, and i, j represent independent variables. Input to the fully connected layer to get the final classification result The calculation formula for the classification result is as follows:
[0089]
[0090] Among them, W and b are the weight matrix and bias term of the fully connected layer respectively, and ReLU is the activation function for the probability output of single-label classification.
[0091] In some embodiments, the classification result is subjected to loss calculation processing according to the classification prediction loss function to obtain a classification prediction loss value. In order to calculate the concept loss and prediction loss in the present embodiment of the application, two different loss functions can be used: the binary cross entropy loss function (BinaryCrossEntropyLoss) is used to calculate the concept loss, and the cross entropy loss function (CrossEntropyLoss) is used to calculate the prediction loss. These two loss functions are suitable for processing multi-label and multi-category classification tasks, respectively. Concept loss l u The calculation formula is:
[0092]
[0093] Among them, K is the number of concepts, and log is the natural logarithm. Here, we need to Normalize, for example, using the sigmoid function. Classification loss l c The calculation formula is:
[0094]
[0095] Among them, K is the number of concepts, and log is the natural logarithm. Here, we need to Normalize the data, for example using the sigmoid function.
[0096] The classification prediction loss function is:
[0097] l=l c +αl u ;
[0098] Among them, l is the classification prediction loss value, and α is the hyperparameter for balancing concepts and classifications.
[0099] In some embodiments, the embodiments of the present application perform parameter update processing on the classification prediction network according to the classification prediction loss value to obtain a classification prediction model. The classification prediction loss value is calculated through the total loss function, the loss function of the classification prediction network is optimized according to the classification prediction loss value, the network loss of the loss function is back-propagated, and the network parameters are continuously adjusted until the update threshold is reached, and the optimization of the classification prediction network is stopped to obtain a classification prediction model that meets the requirements.
[0100] One of the above technical solutions has the following advantages or beneficial effects: the embodiment of the present application inputs the training data set into a pre-constructed classification prediction network for training processing, performs multimodal learning on the data through the multimodal learning module in the classification prediction network, and performs principal component extraction on the extracted features through the concept optimization module and increases the variance between concepts to obtain the target concept feature map, and finally classifies and predicts the target concept feature map through the classification module to obtain the classification result, and updates the parameters of the classification prediction network based on the classification prediction loss function to obtain a trained classification prediction model, thereby improving the accuracy of the classification prediction model.
[0101] In some embodiments, the step of inputting the training data set into the multimodal learning module for feature extraction to obtain an initial concept feature map comprises the following steps:
[0102] Inputting the training data set into the multimodal learning module for network training processing to obtain a feature extraction layer;
[0103] Inputting the training data set into the feature extraction layer for feature extraction processing to obtain a high-level feature map;
[0104] The high-level feature map is converted through a convolutional layer to obtain the initial concept feature map.
[0105] In an embodiment of the present application, the training data set is input into the multimodal learning module for network training processing, and the backbone network in the multimodal learning module is trained to obtain a feature extraction layer, which is used to extract features from the image data. The image data in the training data set is then input into the feature extraction layer for feature extraction processing to obtain a high-level feature map. Since the traditional deep convolutional network contains a large number of extracted high-level features in the last few layers, and the spatial correspondence between these features and the original input image is retained through the convolution operation. The embodiment of the present application takes advantage of this feature and changes the last global average pooling of the backbone network to a 1x1 convolutional layer to convert the high-level features extracted from the previous level into multiple concept maps containing location information. Each concept map represents the response intensity of a specific disease concept in different areas of the image, and these response values retain the spatial distribution information. Suppose the input image is X∈R H×W×C , where H, W and C represent the height, width and number of channels of the image respectively. Through several layers of convolution, pooling and activation operations in the backbone network, the high-level feature map F∈R is obtained. H ′×W′×D , where H′ and W′ are the spatial dimensions of the feature map, and D is the number of feature channels. In the last layer of the model, the embodiment of the present application uses a 1×1 convolution operation to convert the high-level feature map F into a concept map G∈R H′×W′×K , where K is the number of concepts. Each channel G kRepresents the response of the k-th concept in the image space position:
[0106] G k =Conv 1×1 (F) k ;
[0107] Among them, Conv 1×1 Represents a 1×1 convolution operation.
[0108] One of the above technical solutions has the following advantages or beneficial effects: The embodiment of the present application converts the high-level feature map through the convolution layer to obtain the initial concept feature map, which can make the model pay more attention to the learning of concepts and improve the model's ability to learn concepts.
[0109] In some embodiments, the multimodal learning module includes a backbone network and a text encoder, and the inputting of the training data set into the multimodal learning module for network training processing to obtain a feature extraction layer includes the following steps:
[0110] Inputting the training data set into the backbone network and the text encoder for feature extraction processing to obtain image representation and text representation;
[0111] Performing linear mapping conversion processing and contrast loss calculation processing on the image representation and the text representation to obtain a contrast loss value;
[0112] The parameters of the backbone network are updated according to the contrast loss value to obtain the feature extraction layer.
[0113] In the embodiment of the present application, in order to obtain better initialization weights, the embodiment of the present application introduces a multimodal learning method to pre-train the backbone network. The training of the backbone network through multimodal learning can use medical text containing more image information to replace labels to supervise the training of the image encoder, thereby improving the image encoder's ability to extract image detail features, improving model performance, and obtaining excellent initialization weights. Please refer to Figure 2The multimodal learning module includes a backbone network and a text encoder. The backbone network uses the MobileNet-v2 network. The text encoder is based on the Transformer encoder. The image and text are respectively input into the backbone network and the text encoder for feature extraction to obtain image representation and text representation. Then, the image and text representations are converted into a unified dimension through linear mapping. Finally, the image representation and text representation are subjected to contrast loss calculation. The loss value is back-propagated to update the weight parameters of the backbone network and the text encoder to obtain the feature extraction layer. Among them, the calculation of contrast loss realizes the alignment of image and text modalities. The calculation process is to calculate the cosine similarity between each image and all texts and the cosine similarity between each text and all images, perform weighted summation, and finally maximize the similarity between the real image-text pairs.
[0114] One of the above technical solutions has the following advantages or beneficial effects: the embodiment of the present application improves the image encoder's ability to extract image detail features through multimodal learning, thereby improving the performance of the model.
[0115] In some embodiments, the step of inputting the initial concept feature graph into the concept optimization module for whitening and gradient optimization to obtain a target concept feature graph comprises the following steps:
[0116] Performing iterative whitening processing on the initial concept feature map to obtain a whitening matrix;
[0117] Constructing an initial orthogonal matrix according to the output result of the multimodal learning module;
[0118] Performing iterative optimization processing on the initial orthogonal matrix to obtain an optimized orthogonal matrix;
[0119] The whitening matrix is subjected to feature space orthogonalization processing according to the optimized orthogonal matrix to obtain the target concept feature map.
[0120] In an embodiment of the present application, a whitening matrix is obtained by iteratively whitening the initial concept feature map, and then an initial orthogonal matrix is constructed based on the output results of the multimodal learning module. In order to increase the variance between concepts, the embodiment of the present application needs to iteratively optimize the initial orthogonal matrix so that it can orthogonalize the feature components of the concepts, and an optimized orthogonal matrix is obtained by iteratively optimizing the initial orthogonal matrix. The feature space of the whitened matrix is orthogonalized according to the optimized orthogonal matrix, and the transformed matrix can be obtained by multiplying the whitened matrix with the optimized orthogonal matrix. This transformation process converts the feature space of the whitened matrix into an orthogonal feature space. According to the transformed matrix, the feature vector or feature map related to the target concept can be extracted to obtain the target concept feature map, and the target concept feature map can be directly input into the classification module for classification prediction to obtain the classification result.
[0121] One of the above technical solutions has the following advantages or beneficial effects: the embodiment of the present application further performs a ZCA operation on the result of feature extraction, which can increase the model's ability to extract feature concepts, ignore unnecessary data information as much as possible, and improve the accuracy of the model. And the embodiment of the present application minimizes the variance of concept classes and maximizes the variance between concepts by optimizing the orthogonal matrix, so that the concept features learned by the model are clear and accurate. On this basis, the classification prediction model can separate concepts through a simple classifier and achieve a higher classification level.
[0122] In some embodiments, the iterative whitening process is performed on the initial concept feature map to obtain a whitening matrix, comprising the following steps:
[0123] Decentralizing the initial concept feature graph to obtain an initial matrix;
[0124] Performing covariance calculation processing on the initial matrix to obtain a covariance matrix;
[0125] Initializing the initial matrix to obtain a unit matrix;
[0126] Perform Newton iterative calculation processing on the unit matrix according to the covariance matrix to obtain an iterative matrix;
[0127] The whitening matrix is obtained by performing calculation processing according to the iteration matrix and the covariance matrix.
[0128] In the embodiment of the present application, let x be the input data, f be the feature extractor, C = {c1, c2, ..., c k} is the concept output, Y is the diagnosis output, and f(x) is defined here as the concept feature extracted by the model. Obviously, the features extracted here are diverse and not specified, which include the concept features required by the embodiment of the present application (c1, c2, ..., ck ), and also includes other features, such as hair, moles and other interference factors. Obviously, these features are useless for concept prediction in the embodiment of the present application. Therefore, the embodiment of the present application performs a ZCA operation on the result of feature extraction, in order to increase the model's ability to extract feature concepts and ignore unnecessary data information as much as possible.
[0129] In normal operations, ZCA whitening requires calculating the covariance matrix of the data and performing eigenvalue decomposition on the covariance matrix to obtain a whitening matrix. This algorithm can be used for general image preprocessing, but when processing the feature output of the model, the eigenvalue decomposition takes a lot of time to calculate. At the same time, the uncertainty of the model output leads to the frequent appearance of highly ill-conditioned covariance matrices during training, and these matrices cannot be eigenvalue decomposed. To address this problem, the embodiment of the present application improves the ZCA whitening process and uses the Newton algorithm to calculate the ZCA decomposition. Specifically, the embodiment of the present application sets a ZCA whitening operation after the feature extraction layer. in The calculation formula is as follows:
[0130]
[0131] Among them, n = b × h × w represents the dimension of the image, and W represents the feature vector Z extracted by the feature extraction layer. b×d×h×w The whitening operation, that is, the whitening matrix, μ 1×n is the average value of the channel. In order to further improve the stability of the whitening operation and the efficiency of calculating the W matrix, the embodiment of the present application adopts an iterative whitening algorithm to accelerate the calculation. The specific steps of the algorithm are as follows: First, the initial concept feature map is decentralized to obtain the initial matrix. The calculation formula for decentralizing the input data is as follows:
[0132] Z c =Z-μ 1×n ;
[0133] Among them, Z represents the initial concept feature map, Z c represents the initial matrix;
[0134] Then the initial matrix is processed for covariance calculation to obtain the covariance matrix. The formula for calculating the covariance matrix is as follows:
[0135]
[0136] Initialize the identity matrix according to the dimension of the initial matrix, that is, initialize P1 to the same identity matrix of dimension Z, and for k = 2 to T, use the following Newton iteration formula to update P k :
[0137]
[0138] Finally, calculate the whitened output:
[0139]
[0140] Among them, tr(Σ) is the trace of matrix Σ (the sum of all diagonal elements), which is used for normalization. The final W is the whitening matrix, and the whitening operation is
[0141] One of the above technical solutions has the following advantages or beneficial effects: The embodiment of the present application obtains a whitening matrix by iteratively whitening the initial concept feature map, which can increase the model's ability to extract feature concepts, ignore unnecessary data information as much as possible, and improve the accuracy of the model.
[0142] In some embodiments, the iterative optimization process of the initial orthogonal matrix to obtain an optimized orthogonal matrix comprises the following steps:
[0143] Performing binarization calculation processing on the initial concept feature map to obtain a concept segmentation map;
[0144] Performing maximization regularization processing on the initial orthogonal matrix according to the concept segmentation graph to construct a target optimization function;
[0145] The target optimization function is iteratively calculated according to a conjugate gradient search algorithm to obtain the optimized orthogonal matrix.
[0146] In the embodiment of the present application, in order to further enhance the conceptual simplicity difference, the embodiment of the present application introduces an orthogonal matrix Q, let is the matrix after orthogonal transformation, and f is the part of the model before ZCA whitening, that is, the feature extraction layer. In order to increase the variance between concepts, the embodiment of the present application needs to iteratively optimize the orthogonal matrix so that it can orthogonalize the feature components of the concept. The model finally outputs the cth k Concept feature map is the concept attention map. The embodiment of the present application further performs a binarization operation to obtain a concept segmentation map In order to align k concepts and dimensions, the present embodiment uses To construct the target optimization function:
[0147]
[0148] Q T Q=I d ;
[0149] Among them, Q is the orthogonal matrix used to orthogonalize the conceptual data, Indicates that K is the number of concepts and n is the dimension of Q; Indicates the c k The number of data partitions corresponding to the concept, x i For the corresponding c k The specific data in the division is used in the embodiment of the present application. It represents the feature map obtained after the data passes through the feature extraction layer and uses the ZCA whitening algorithm. There are still many features that are irrelevant to the concept, such as the features of the area without lesions. In order to further simplify the complexity of the optimization problem and improve the accuracy, the embodiment of the present application uses the model's attention hotspot map for the concept After binarization, we get It can be roughly regarded as a segmentation map of conceptual features. In the process of solving the regularization problem, the embodiment of the present application removes the data irrelevant to the concept. And processed x i The average is obtained by multiplication, and then the different concept data are regularized to obtain the optimization function. Since the embodiment of the present application needs to distinguish different concepts as much as possible, the embodiment of the present application needs to find the maximum value of the optimization problem.
[0150] This optimization function is NP-hard in the calculation process, and it is difficult to solve it directly. In fact, for the feasible set is a Stiefel flow type. In particular, the optimization function constructed in the embodiment of the present application is a differentiable function, and it is Therefore, it can be further simplified to the unit sphere flow pattern
[0151] For the unit sphere flow pattern Sp n-1 Optimization, initially, the embodiment of the present application defines a feasible point Q (t) And the current optimization function In Q (t) The gradient The embodiment of the present application defines an oblique symmetrical rectangle A and a smooth curve Y(τ), and the calculation of A is given by the following formula:
[0152] A=G(Q (t) ) T -Q (t) G T ;
[0153] Y(τ) is is a smooth function, and is related to Q by the following relationship: Y(τ)=Q (t+1) Q (t) ,YT (τ)Y(τ)=Q T Q,. For Y(τ), the embodiment of the present application adopts the conjugate gradient search algorithm for iteration, where τ is the step size. The following formula can be obtained:
[0154]
[0155] According to the property of Y and the characteristic that A is a skew-symmetric matrix, the embodiment of the present application can further eliminate Y to obtain the following formula:
[0156]
[0157] Among them, τ represents the learning rate, G represents the gradient of the Q objective function, and the final optimized orthogonal matrix can be obtained.
[0158] One of the above technical solutions has the following advantages or beneficial effects: The embodiment of the present application minimizes the variance of concept classes and maximizes the variance between concepts by optimizing the orthogonal matrix, so that the concept features learned by the model are clear and accurate, thereby improving the accuracy of classification prediction.
[0159] The following is a detailed description of the embodiments of the present invention with reference to specific application examples:
[0160] The embodiment of the present application is based on a convolutional neural network and is applied to the field of artificial intelligence technology. It can be used for the recognition of all medical images, such as X-ray imaging, CT scan images, ultrasound imaging, endoscopy, dermatoscope and other medical image recognition scenarios. Some medical images are relatively complex, and doctors who are not experienced enough may make wrong judgments. For example, in dermoscopy, black spots on the skin are normal melanin deposition, or melanoma. The boundary between the two may be relatively vague, and it is easy to make misjudgments, especially for hospitals in backward areas. Due to the lack of professional medical personnel, the treatment time may be delayed, resulting in more serious consequences. The embodiment of the present application can be used to assist doctors in diagnosis, provide highly accurate medical diagnosis results, and give a certain concept of diagnosis to help doctors make their own judgments. On the other hand, the diagnosis accumulation of medical personnel requires a lot of practical accumulation. If there is no more professional medical personnel to guide, wrong experience may be accumulated. The medical image interpretable classification prediction method provided by the embodiment of the present application can give relevant concept distributions while giving diagnostic results, so that doctors can learn the characteristics of specific diseases in the concepts, helping doctors to learn better autonomously. The present application embodiment obtains a medical data set and performs concept label conversion and image preprocessing to obtain a training data set. The training data set can be divided into a training set and a test set. The data in the training set is used to train the model, and the data in the test set is used to test the learning effect of the model. The training set is then input into a pre-built classification prediction network for training. Figure 3, the structure of the classification prediction network is as follows Figure 3 As shown in the figure, the classification prediction network includes a multimodal learning module, a concept optimization module and a classification module. The multimodal learning module includes a backbone network and a text encoder. The backbone network adopts a deep neural network model structure. The backbone network is used to extract features from image data. The extracted high-level features are converted into concept graphs through convolution layers. The text encoder is used to extract features from text data. The extracted text representation and the image representation extracted by the backbone network are compared and loss is calculated, so that the parameters of the backbone network are adjusted according to the comparison loss value. Finally, the backbone network and convolution layer with adjusted parameters are used as feature extraction layers. Then, the extracted concept graph is subjected to concept feature extraction and regularization processing through the concept optimization module, and finally input into the classification module for classification prediction to obtain the classification result.
[0161] See also Figure 4 The embodiment of the present application also provides a medical image interpretable classification prediction system, which can implement the above-mentioned medical image interpretable classification prediction method, and the system includes:
[0162] The first module 401 is used to obtain a medical data set;
[0163] The second module 402 is used to perform concept label conversion and image preprocessing on the medical data set to obtain a training data set;
[0164] The third module 403 is used to input the training data set into a pre-built classification prediction network for training processing to obtain a classification prediction model; the classification prediction network includes a multimodal learning module, a concept optimization module and a classification module;
[0165] The fourth module 404 is used to input the medical image to be processed into the classification prediction model for classification prediction processing to obtain a classification prediction result.
[0166] It can be understood that the contents of the above method embodiments are all applicable to the present system embodiments, the functions specifically implemented by the present system embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0167] The embodiment of the present application also provides an electronic device, the electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the above-mentioned medical image interpretable classification prediction method when executing the computer program. The electronic device can be any intelligent terminal including a tablet computer, a car computer, etc.
[0168] It can be understood that the contents of the above method embodiments are all applicable to the present device embodiments, the functions specifically implemented by the present device embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0169] See also Figure 5 , Figure 5 The hardware structure of an electronic device of another embodiment is illustrated, and the electronic device includes:
[0170] The processor 501 may be implemented by a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application;
[0171] The memory 502 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 502 can store an operating system and other application programs. When the technical solution provided in the embodiment of this specification is implemented by software or firmware, the relevant program code is stored in the memory 502, and the processor 501 calls and executes the medical image interpretable classification prediction method of the embodiment of this application;
[0172] Input / output interface 503, used to implement information input and output;
[0173] Communication interface 504, used to realize communication interaction between the device and other devices, which can be realized by wired mode (such as USB, network cable, etc.) or wireless mode (such as mobile network, WIFI, Bluetooth, etc.);
[0174] A bus 505 that transmits information between the various components of the device (e.g., the processor 501, the memory 502, the input / output interface 503, and the communication interface 504);
[0175] The processor 501 , the memory 502 , the input / output interface 503 and the communication interface 504 are connected to each other in communication within the device via the bus 505 .
[0176] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned medical image interpretable classification prediction method.
[0177] It can be understood that the contents of the above method embodiments are all applicable to the present storage medium embodiments, the functions specifically implemented by the present storage medium embodiments are the same as those of the above method embodiments, and the beneficial effects achieved are also the same as those achieved by the above method embodiments.
[0178] The memory, as a non-transient computer-readable storage medium, can be used to store non-transient software programs and non-transient computer executable programs. In addition, the memory may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories may be connected to the processor via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0179] The embodiments of the present application provide a method, system, device and medium for interpretable classification prediction of medical images. The scheme obtains a medical data set, performs concept label conversion and image preprocessing on the medical data set to obtain a training data set, and can convert the text data in the medical data set into concept labels for subsequent training processing of the model, thereby improving the model's ability to learn concepts. In addition, the embodiments of the present application input the training data set into a pre-built classification prediction network for training processing to obtain a classification prediction model, and input the medical image to be processed into the classification prediction model for classification prediction processing to obtain a classification prediction result. The embodiments of the present application can extract the principal components of the extracted features through the concept optimization module in the classification prediction network, reduce interference information, and maximize the concept difference so that the concept features learned by the model are clear and accurate, thereby improving the accuracy of the classification prediction results.
[0180] The embodiment of the present application can deepen the model's learning of concepts: through the intermediate whitening matrix and maximizing the concept difference, the model can actively retain the main components in the extracted features, ignore interference information, minimize the concept class variance, and maximize the concept variance, so that the concept features learned by the model are clear and accurate. On this basis, the concepts can be separated through a simple classifier and a higher classification level can be achieved.
[0181] The embodiment of the present application can also reduce the black box effect of the model, making the model more reliable: the main components of the model are divided into three parts, the first part is the data information extraction part, the second part is the concept information attention concentration part, and the third part is the auxiliary judgment part from concept information to diagnostic information. These three parts are optimized by the loss function to ensure that each part can achieve the expected results. At the same time, each part can visualize the results as a separate existence. The three parts have clear division of labor, which reduces the black box effect of the model and allows staff to have more basis when testing the reliability of the model output results. At the same time, the well-organized model division results can be split and used separately, which is convenient for deepening and optimizing the model.
[0182] The embodiments of the present application improve the anti-interference and robustness of the model: in the model concept extraction, only the main components of the model will be extracted through the ZCA whitening method and trade-offs will be made. In this step, the impact of interfering factors in the image can be reduced, thereby alleviating the problems of insufficient robustness and low clinical trust in the medical image-assisted diagnosis model.
[0183] In addition, the embodiments of the present application also enhance the interpretability and generalization performance of the model: the embodiments of the present application integrate the diagnostic report text into the model training process, and improve the feature extraction capability of the image encoder by providing a wider range of supervision sources, so that the model can better capture the lesion area and improve the interpretability; when the model is migrated to other data sets, it can obtain better results under the same fine-tuning conditions, and the generalization performance of the model is improved.
[0184] The embodiments described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided in the embodiments of the present application. Those skilled in the art will appreciate that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided in the embodiments of the present application are also applicable to similar technical problems.
[0185] Those skilled in the art will appreciate that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or a combination of certain steps, or different steps.
[0186] The system embodiments described above are merely illustrative, and the units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.
[0187] Those skilled in the art will appreciate that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices may be implemented as software, firmware, hardware, or a suitable combination thereof.
[0188] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0189] It should be understood that in the present application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the objects associated before and after are in an "or" relationship. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0190] In the several embodiments provided in the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are only schematic. For example, the division of the above units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of systems or units, which can be electrical, mechanical or other forms.
[0191] The units described above as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0192] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0193] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including multiple instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), disk or optical disk and other media that can store programs.
[0194] The preferred embodiments of the present invention are described above with reference to the accompanying drawings, but the scope of the rights of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by a person skilled in the art without departing from the scope and essence of the present invention should be within the scope of the rights of the present invention.
Claims
1. A medical image interpretable classification prediction method, characterized in that: The method comprises the following steps: Access to medical datasets; Performing concept label conversion and image preprocessing on the medical data set to obtain a training data set; Inputting the training data set into a pre-built classification prediction network for training processing to obtain a classification prediction model; the classification prediction network includes a multimodal learning module, a concept optimization module and a classification module; The medical image to be processed is input into the classification prediction model for classification prediction processing to obtain a classification prediction result.
2. The method according to claim 1, characterized in that The step of inputting the training data set into a pre-built classification prediction network for training to obtain a classification prediction model comprises the following steps: Inputting the training data set into the multimodal learning module for feature extraction processing to obtain an initial concept feature map; the initial concept feature map is used to characterize the response intensity of the concept label at the image space position; Inputting the initial concept feature graph into the concept optimization module for whitening and gradient optimization to obtain a target concept feature graph; Inputting the target concept feature graph into the classification module for classification processing to obtain a classification result; Perform loss calculation processing on the classification result according to the classification prediction loss function to obtain a classification prediction loss value; The classification prediction network is subjected to parameter update processing according to the classification prediction loss value to obtain the classification prediction model.
3. The method according to claim 2, characterized in that The step of inputting the training data set into the multimodal learning module for feature extraction to obtain an initial concept feature graph comprises the following steps: Inputting the training data set into the multimodal learning module for network training processing to obtain a feature extraction layer; Inputting the training data set into the feature extraction layer for feature extraction processing to obtain a high-level feature map; The high-level feature map is converted through a convolutional layer to obtain the initial concept feature map.
4. The method according to claim 3, characterized in that The multimodal learning module includes a backbone network and a text encoder, and the training data set is input into the multimodal learning module for network training processing to obtain a feature extraction layer, including the following steps: Inputting the training data set into the backbone network and the text encoder for feature extraction processing to obtain image representation and text representation; Performing linear mapping conversion processing and contrast loss calculation processing on the image representation and the text representation to obtain a contrast loss value; The parameters of the backbone network are updated according to the contrast loss value to obtain the feature extraction layer.
5. The method according to claim 2, characterized in that: The step of inputting the initial concept feature graph into the concept optimization module for whitening and gradient optimization to obtain a target concept feature graph comprises the following steps: Performing iterative whitening processing on the initial concept feature map to obtain a whitening matrix; Constructing an initial orthogonal matrix according to the output result of the multimodal learning module; Performing iterative optimization processing on the initial orthogonal matrix to obtain an optimized orthogonal matrix; The whitening matrix is subjected to feature space orthogonalization processing according to the optimized orthogonal matrix to obtain the target concept feature map.
6. The method according to claim 5, characterized in that The iterative whitening process is performed on the initial concept feature map to obtain a whitening matrix, comprising the following steps: Decentralizing the initial concept feature graph to obtain an initial matrix; Performing covariance calculation processing on the initial matrix to obtain a covariance matrix; Initializing the initial matrix to obtain a unit matrix; Perform Newton iterative calculation processing on the unit matrix according to the covariance matrix to obtain an iterative matrix; The whitening matrix is obtained by performing calculation processing according to the iteration matrix and the covariance matrix.
7. The method according to claim 5, characterized in that The iterative optimization process of the initial orthogonal matrix to obtain an optimized orthogonal matrix comprises the following steps: Performing binarization calculation processing on the initial concept feature map to obtain a concept segmentation map; Performing maximization regularization processing on the initial orthogonal matrix according to the concept segmentation graph to construct a target optimization function; The target optimization function is iteratively calculated according to a conjugate gradient search algorithm to obtain the optimized orthogonal matrix.
8. A medical image interpretable classification prediction system, characterized in that: The system comprises: The first module is used to obtain medical data sets; The second module is used to perform concept label conversion and image preprocessing on the medical data set to obtain a training data set; The third module is used to input the training data set into a pre-built classification prediction network for training processing to obtain a classification prediction model; the classification prediction network includes a multimodal learning module, a concept optimization module and a classification module; The fourth module is used to input the medical image to be processed into the classification prediction model for classification prediction processing to obtain the classification prediction result.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.