Method, electronic device and program product for image processing
By introducing an image interpretation generator into the deep learning model, extracting medical image features and generating descriptive text, the problem of the lack of interpretability in medical image processing in deep learning models is solved, and more credible classification results and improved user experience is achieved.
Patent Information
- Application Number
- CN202380077515.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-11-08
- Filing Date
- 2023-11-07
- Publication Date
- 2025-06-13
AI Technical Summary
Deep learning models lack interpretability and confidence in medical image processing, making them difficult to provide a clear explanation of their decision-making process, and cannot meet the needs of regulators and doctors.
An image interpretation generator is proposed that extracts image features from medical images and generates descriptive texts by using a deep learning model with two module structures, providing an explanation of image classification results.
The classification and interpretation of medical images are realized, the generated classification results are more trustworthy, the user experience is improved, and they are in line with the regulatory requirements of relevant departments.
Smart Images

Figure CN120153407A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of computers, and more particularly, to methods, electronic devices, and computer program products for image processing. Background Art
[0002] With the development of image processing technology, some image processing models are used for image classification with good accuracy. Some image processing models can output the classification of an image, while some image processing models can output simple text of what objects are included in the image. These image processing models usually have neural units that simulate the human ability to process information. However, a large amount of pre-annotated data is required as training data to train the image processing model so that these neural units can learn the relationship between the input and the output through model parameters. Summary of the Invention
[0003] According to an embodiment of the present disclosure, there is provided a method, an electronic device, and a computer program product for image processing.
[0004] According to a first aspect of the present disclosure, there is provided a method for image processing. The method includes obtaining a training data set including a medical image and a corresponding annotation text, where the annotation text includes classification information of the medical image and descriptive information associated with the classification information. The method further includes training an image interpretation model using the training data set, where the image interpretation model includes a first computational model for extracting image features from the medical image and a second computational model for generating text from the image features.
[0005] According to a second aspect of the present disclosure, there is provided an electronic device including: at least one processing unit and at least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for being run by the at least one processing unit, where the instructions, when run by the at least one processing unit, cause a computing device to execute a method. The method includes obtaining a training data set including a medical image and a corresponding annotation text, where the annotation text includes classification information of the medical image and descriptive information associated with the classification information. The method further includes training an image interpretation model using the training data set, where the image interpretation model includes a first computational model for extracting image features from the medical image and a second computational model for generating text from the image features.
[0006] According to a third aspect of the present disclosure, there is provided a computer program product including machine-executable instructions, where the machine-executable instructions, when executed by a device, cause the device to execute the method according to the first aspect of the present disclosure.
[0007] The Summary of the Invention is provided to introduce a selection of concepts in a simplified form, and these concepts will be further described in the Detailed Description below. The Summary of the Invention does not identify the key features or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] The above and other features, advantages, and aspects of various embodiments of the present disclosure will become more apparent from the following detailed description in conjunction with the accompanying drawings. In the drawings, the same or similar reference numerals denote the same or similar elements, wherein:
[0009] Figure 1 A block diagram showing an example environment in which some embodiments of the present disclosure can be implemented;
[0010] Figure 2 A schematic flowchart showing a method for image processing according to an embodiment of the present disclosure;
[0011] Figure 3A A schematic diagram showing an image feature extraction process in a first computing model according to an embodiment of the present disclosure;
[0012] Figure 3B A schematic diagram showing an image feature output process in a first computing model according to an embodiment of the present disclosure;
[0013] Figure 4 A schematic diagram showing a process for determining attention in a first computing model according to an embodiment of the present disclosure;
[0014] Figure 5 A schematic diagram showing a process for generating a text sequence in a second computing model according to an embodiment of the present disclosure;
[0015] Figure 6 A schematic diagram showing a process for determining a dictionary according to an embodiment of the present disclosure;
[0016] Figure 7 A schematic diagram showing a comparison between an embodiment of the present disclosure and image captions; and
[0017] Figure 8 A schematic block diagram showing an example device that can be used to implement some embodiments according to the present disclosure.
[0018] In all the drawings, the same or similar reference numerals denote the same or similar elements. DETAILED DESCRIPTION
[0019] In the following, embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure have been shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided to help understand the present disclosure more thoroughly and completely. It should be understood that the drawings and embodiments of the present disclosure are for illustrative purposes only and are not intended to limit the scope of protection of the present disclosure.
[0020] In the description of the embodiments of the present disclosure, the term "including" and its like terms should be understood as open-ended inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. Other explicit and implicit definitions may also be included below. In addition, the specific numerical values herein are examples and are only intended to assist understanding and not to limit the scope.
[0021] With the rapid development of artificial intelligence (AI) in the field of medical care, regulatory agencies have begun to formulate and issue new regulations and standards for AI-based medical devices, which has had a significant impact on the market launch of AI medical products. Although AI, especially deep learning (DL), has achieved great success in various fields and applications compared with conventional machine learning methods such as decision trees and support vector machines (SVM), DL-based methods usually rely on black-box algorithms and are relatively weak in explaining their reasoning processes.
[0022] DL algorithms belonging to the black box usually learn the mapping function between the input and output by training a neural network with multiple hidden layers. DL algorithms are based on a large amount of training data and high computing power. Through the training process, features are automatically learned, and these features are difficult to be explained by the professional knowledge in the medical field. Traditional white-box machine learning methods manually design and extract features based on the expert's domain knowledge, thus having better interpretability.
[0023] Therefore, relying on traditional techniques, we only know that deep learning techniques can significantly improve the performance of various applications and tasks through experimental verification. However, there is still no very clear answer as to why the DL model can provide correct results and what information the model is based on to provide such correct decisions, and there is also a lack of a theoretical basis to support and confirm the capabilities of deep learning. In particular, in the field of medical care, not only do regulatory agencies gradually begin to require AI medical devices to provide algorithm interpretability, but doctors also begin to expect trustworthy explanations for AI products. Therefore, the lack of interpretability and confidence of DL methods is a key issue that needs to be solved.
[0024] In view of this, embodiments of the present disclosure provide an image processing solution. In this solution, in the present invention, an image interpretation generator is proposed, which can provide the results and descriptions for classifying medical images to assist in diagnosis. The images are annotated by experienced experts using text paragraphs, and the words or phrases in the paragraphs are extracted and grouped into a dataset as a standardized dictionary (also known as a lexicon). A deep learning model with a two-module structure is used as the basis for training. The output paragraphs are generated word by word or phrase by phrase, and each word or phrase is selected from the previously defined dictionary. In this way, not only can the input medical images be classified, but also the reasons for generating the classification results are provided. Therefore, the classification results are more trustworthy, the user experience is improved, and it also complies with the regulatory requirements of relevant departments.
[0025] Reference will continue to be made below to Figures 1 to 8 Describe some example embodiments of the present disclosure. It should be noted that for ease of understanding, the embodiments of the present disclosure are described below using medical images as an example, but the embodiments of the present disclosure are also applicable to any other type of image. In this case, the annotation information will be given by a person with domain knowledge. In addition, the classification task is regarded as an example described below, but the embodiments of the present disclosure are also applicable to other types of image processing tasks, such as object detection, object recognition, object tracking, etc. The present disclosure is not limited to the above aspects.
[0026] Figure 1 A block diagram of an example environment 100 in which some embodiments of the present disclosure can be implemented is shown. The example environment 100 generally depicts various exemplary elements involved in the methods proposed in the present disclosure. The environment 100 includes a computing device 110. The computing device 110 can be, for example, a computing system or a server. The computing device 110 includes an image interpretation model 111 to provide image processing functions. In some embodiments, the computing device 110 can store code with instructions to provide image processing functions.
[0027] The example environment 100 also includes a training dataset 120. The training dataset 120 includes medical images 122 and corresponding annotation texts 124. It can be understood that for the sake of brevity, the medical image 122 is a general term for multiple medical images. The annotation text 124 is also a general term, which includes multiple annotation texts corresponding to multiple medical images.
[0028] Medical images are generally understood to be images of the human body or specific parts obtained by medical imaging devices. For example, images of the stomach, kidneys, liver, lungs, etc. are obtained using techniques such as X-rays, ultrasound, computed tomography (CT), or magnetic resonance imaging (MRI).
[0029] Medical images are annotated by experienced doctors to obtain the correct classification results and the reasons for classifying them in this way, and the medical images are used as the annotation text. For example, for a medical image with a liver lesion, the annotation text can be "LI-RADS (Liver Imaging Reporting and Data System) Liver Lesion Grade 4, because of the large size, there is an enhanced envelope and non-peripheral washout".
[0030] A training data set 120 is provided to train the image interpretation model 111. The image interpretation model 111 includes a first computational model 112 for extracting features from the image, and a second computational model 116 for generating classification information and corresponding descriptive text based on the extracted image features. Specifically, the first computational model 112 extracts image features 114-1 and determines the corresponding attention 114-2. It should be noted that for the sake of illustration, the image features 114-1 and attention 114-2 in this article are also abstract concepts, which include multiple image features and multiple corresponding attentions. Attention can be understood as a weight, which reflects the degree of attention to the corresponding image features. Generally, the attention to unimportant image features (such as the background) is relatively low. However, the attention to lesions is relatively high.
[0031] The image features 114-1 and attention 114-2 are input into the second computational model 116. Based on the image features 114-1 and attention 114-2, the second computational model 116 generates words and phrases 118-1 that make up the descriptive text and the corresponding probability distribution 118-2. Sentences are formed by the words or phrases with the highest probability, and then paragraphs are formed. The paragraph includes classification information 130 and descriptive information 132. The descriptive information 132 can explain why the image is classified as the result.
[0032] The example environment 100 may also include a medical image 140 to be processed. The medical image 140 and the medical image 122 are basically not different in terms of the physical acquisition mode, and both are medical images of human organs, but there are differences in application. When the image interpretation model 111 is trained and applied in the inference stage, the medical image 140 is used to perform the image classification task. Generally speaking, the training data set 120 is used for the training stage, and the medical image 140 is used for the inference stage. That is to say, the training data set and the medical image are not used simultaneously.
[0033] The above refers to Figure 1 The environment 100 in which the embodiments of the present disclosure can be implemented has been described. It should be understood that the environment 100 is merely exemplary, and the embodiments of the present disclosure can be implemented in other environments different from this. For example, the training data set 120 and the medical image 140 can be implemented in the same or different devices.
[0034] Figure 2FIG. 200 is a schematic flow chart of a method for image processing according to an embodiment of the present disclosure. For ease of description, method 200 may be implemented in Figure 1 the computing device 110 shown. It should be understood that method 200 may also include additional actions not shown and / or may omit the actions shown, and the scope of the present disclosure is not limited in this regard. For ease of understanding, in conjunction with Figure 1 FIG. 200 is shown.
[0035] At block 202, a training data set is obtained, where the training data set includes medical images and corresponding annotation texts, and the annotation texts include classification information of the medical images and descriptive information associated with the classification information. As an example, the computing device 110 may obtain a training data set 120, where the training data set 120 may include medical images 122. The medical images 122 may be images of each type of lesion, and each image is formed by a medical imaging device. For each medical image that has been imaged, it may be annotated with expert experience as the annotation text 124 for training the image interpretation model. The annotation text 124 includes the classification result of the medical image 122 and the descriptive information of the classification result. The descriptive information is not a simple sentence, but details the reason for generating the classification result. For example, for liver tumors, the descriptive information may include, for example, the size and form information of the tumor.
[0036] In some embodiments, the descriptive information is determined based on medical standards. As an example, assume that liver tumors have a 5-level classification, and each level has corresponding tumor size and form. Then, the classification and description of the medical image of the liver should be determined according to the annotation. For the descriptive information, standard medical vocabulary and correct grammar should also be used to ensure the applicability of the image interpretation model to the public.
[0037] At block 204, an image interpretation model is trained using the training data set, where the image interpretation model includes a first computing model for extracting image features from medical images and a second computing model for generating text from the image features. As an example, the first computing model 112 may be used to extract image features, and the second computing model 116 may be used to generate the classification information of a specific medical image and the descriptive text for explaining why the specific medical image corresponds to the classification result.
[0038] As an example, in the LI-RADS system, liver lesions can be classified into five levels. The higher the level, the more severe the disease. There are criteria in conventional diagnosis, that is, size, the presence of an enhancement envelope, and non-peripheral washout can directly lead to a diagnostic decision. Terms such as "size", "large", "small", "medium", "presence of an enhancement envelope", "presence of non-peripheral washout" are defined as standardized words. According to standardized medical criteria, each input image is translated by a doctor into a text explanation.
[0039] Reference will be made to Figure 3A 、 Figure 3B and Figure 4 to describe the detailed structure and application process of the first computational model 112, and reference will be made to Figure 5 to describe the detailed structure and application process of the second computational model 116, so the first computational model 112 and the second computational model 116 will not be described in detail here.
[0040] By means of the method 200, the input medical images can be classified, and the reasons for generating the classification results can be provided. Since the descriptive information of the classification results is human-readable text, not simple sentences, but includes specific reasons, the classification results are more acceptable to people, making the classification results more trustworthy, thus providing a more user-friendly experience.
[0041] Figure 3A FIG. shows a schematic diagram of the image feature extraction process 300 in the first computational model according to an embodiment of the present disclosure. As Figure 3A shown, the medical image 302 can be divided into small blocks, such as image block 304 and image block 306. These small blocks can have N×N pixels, which can have the same or different sizes. Image features can be extracted from the image block 304 by the convolutional layer 320 of the first computational model 112. Image features can also be extracted from the image block 306 by the convolutional layer 320 of the first computational model 112. The convolutional layer 320 has a convolutional kernel that is sensitive to specific features, so that the convolutional kernel can be used to extract the features of interest. The first computational model 112 can have multiple convolutional layers, such as convolutional layer 320, convolutional layer 322, and convolutional layer 324. Each convolutional layer can have a convolutional kernel that is interested in different features.
[0042] After the image patch 304 is subjected to feature extraction by the convolutional layer 320, the latent vector 308 can be obtained. The convolutional layer 322 can continuously extract features from the latent vector 308 to generate the latent vector 312, and this process can be performed multiple times, thereby finally generating the image feature 316. Similarly, after the image patch 306 is subjected to feature extraction by the convolutional layer 320, the latent vector 310 can be obtained. The convolutional layer 324 can continuously extract features from the latent vector 310 to generate the latent vector 314, and this process can be performed multiple times, thereby finally generating the image feature 318.
[0043] In this way, after the medical image 302 is processed by multiple convolutional layers, corresponding image features can be extracted to generate multiple image features, such as the vector 316 and the vector 318. It can be understood that in the field of deep learning, features are abstract concepts, which do not necessarily correspond to some or some physical meanings of the target object, and features are usually represented by vectors.
[0044] Figure 3B A schematic diagram of the image feature output process 330 in the first computational model according to an embodiment of the present disclosure is shown. In the field of image processing, an RGB or YUV color system is usually used to represent an image. In this way, an image usually has multiple image channels. As Figure 3B shown, the features of each image channel can be superimposed to form the final image feature. For example, the image feature 332 of the first image channel and the image feature 334 of the second image channel are concatenated together, and then concatenated with the third feature 336 of the third image channel.
[0045] In some embodiments, these image features also have corresponding attention. For example, the attention can be the weight set 338. The weight set 338 and the concatenated vector matrix are output to the second computational model 116 as a whole. The process and module for determining attention will be described below with reference to Figure 4 specifically describe the process and module for determining attention.
[0046] In some embodiments, based on the image channels of the medical image, an image channel weight set associated with the image channels is determined. As an example, the image features of three YUV image channels can be determined based on the medical image 302.
[0047] In some embodiments, based on the spatial distribution of the objects in the medical image, a spatial weight set associated with the space is determined. As an example, if the lesion area of the liver is of interest, the weight of the lesion can be adjusted to be larger, while the weights of the areas of other organs and the image background are adjusted to be smaller.
[0048] In some embodiments, image features are determined based on an image channel weight set and a spatial weight set. As an example, the weight set can be concatenated with the image features extracted by a convolutional layer to form the image features output to the second computing model 116.
[0049] Through process 300 and process 330, the features of the region of interest related to the lesion and the weights of the features can be determined, so that the features are not interfered by other noises, and thus the classification and description of the image are more accurate.
[0050] Figure 4 FIG. shows a schematic diagram of process 400 for determining attention in a first computing model according to an embodiment of the present disclosure. In some embodiments, process 400 can be configured to be executed in process 300. As an example, process 400 can be embedded in each convolutional layer (e.g., Figure 3A convolutional layer 320, convolutional layer 322, and convolutional layer 324) in the CNN network, and the corresponding attention is extracted by the convolutional layer. In some embodiments, process 400 can be executed by a dedicated attention module, and the attention module extracts the corresponding attention.
[0051] As Figure 4 shown, the image features 114-1 can be respectively input into, for example, three fully connected layers (which can be referred to as a multi-layer perceptron (MLP)): fully connected layer 402, fully connected layer 404, and fully connected layer 406. Therefore, three vectors can be obtained respectively, namely vector Q408, vector K410, and vector V412, where vector Q408 and vector K410 can be multiplied and normalized (softmax) at block 414. At block 416, the normalized vector is multiplied by vector V412 to generate the weight set 420.
[0052] In process 400, vector Q acts as a query vector, vector 410 acts as a key vector, and vector 412 acts as a value vector. The importance of the query vector is determined by the similarity of the query vector and the key vector relative to the value vector, and is reflected in its weight. It can be understood that process 400 can be repeated multiple times to complement each other in order to prevent missing details, so as to achieve the purpose of fully focusing on the details that should be noted.
[0053] Figure 5 FIG. shows a schematic diagram of process 500 for generating a text sequence in a second computing model according to an embodiment of the present disclosure. As Figure 5 described, the second computing model 116 can include several prediction units. For example, the computing model 116 includes a sequence-to-sequence model, and the sequence-to-sequence model includes a plurality of prediction units 502, 504, 506, and 508 connected in series, and each prediction unit is configured to output a predicted word or phrase. It can be understood that, Figure 5Merely for example, the second computing model may have more prediction units.
[0054] In some embodiments, the image features output from the first computing model 112 are input into the first prediction unit, such as prediction unit 502. As an example, the prediction unit may be a long short-term memory network (LSTM). In other examples, the prediction unit may also be a Transformer or BERT.
[0055] In some embodiments, for a prediction unit among a plurality of serially connected prediction units, the prediction unit may receive a word or phrase generated by the previous prediction unit as input; and the prediction unit may output the predicted word or phrase to the next prediction unit.
[0056] As an example, [START] is the default input of prediction unit 502 and is used as a start. Prediction unit 502 outputs a probability distribution of the first token according to the image features. In some embodiments, prediction unit 502 outputs a probability distribution of the first token 1 according to the image features and attention (e.g., a set of weights).
[0057] In some embodiments, tokens are determined according to a dictionary. The probability distribution represents the probability of each word in the dictionary. The determination of the dictionary will be referred to below Figure 6 and will not be described in detail herein.
[0058] In prediction unit 504, the token 1 output from prediction unit 502 is processed to generate the second token 2. In some embodiments, the token 1 output from prediction unit 502 and its attention are processed to generate the second token 2.
[0059] In prediction unit 506, the token 2 output from prediction unit 504 is processed to generate the third token 3. In some embodiments, the token 2 output from prediction unit 504 and its attention are processed to generate the third token 3.
[0060] In some embodiments, the prediction unit receiving a word or phrase generated by the previous prediction unit as input includes: in the first prediction unit, based on the image features and the attention to the image features, a first semantic feature associated with the image features is determined. For example, in prediction unit 502, a first semantic feature associated with the image features is determined based on the image feature 114-1 and the attention to the image feature.
[0061] In some embodiments, based on the first semantic feature, a probability distribution of the word or phrase output from the first prediction unit is determined. For example, based on the first semantic feature, a probability distribution of the word or phrase output from prediction unit 502 is determined, and the token 1 is determined according to the probability distribution.
[0062] In some embodiments, in the second prediction unit, based on the first semantic feature and the attention to the first semantic feature, a second semantic feature associated with the first semantic feature is generated. Based on the second semantic feature, the probability distribution of the words or phrases output from the second prediction unit is determined. For example, in the prediction unit 504, a second semantic feature associated with the first semantic feature can be generated based on the first semantic feature and the attention to the first semantic feature. Based on the second semantic feature, the probability distribution of the words or phrases output from the prediction unit 504 is determined, and the token 2 is determined according to the probability distribution.
[0063] Such a serial process can continue until the end, at which point the prediction unit will output [End]. For example, when the user inputs a medical image, words are generated one by one until the paragraph is completed. The generated paragraph consists of two parts. The first sentence can summarize the diagnosis result, and the remaining sentences explain why the result is obtained. It can be seen that the classification result of the medical image and the corresponding descriptive text are generated through the image features and attention using this serial structure. This is a complete paragraph, including the classification result and the reason, so it has better interpretability and is more credible to humans. This also simplifies the workload of medical staff and can provide effective auxiliary diagnosis.
[0064] Figure 6 FIG. shows a schematic diagram of a process 600 for determining a dictionary according to an embodiment of the present disclosure. As Figure 6 shown, the explanatory text or sentence is split into some words or phrases, and then these words or phrases are collected to build a dictionary. Once all the images are annotated, all the words or phrases in the paragraph will be split, extracted, and then collected into the dictionary, which is a set composed of all the words or phrases without repetition. It should be noted that the dictionary needs to include <Start> and <End> to indicate the start and end of the paragraph.
[0065] In some embodiments, training an image interpretation model using a training data set includes: performing word segmentation on the text in the training data set to obtain a dictionary for generating text; and training a second computing model based on the dictionary such that the words or phrases in the text generated by the second computing model are included in the dictionary.
[0066] For example, word segmentation is performed on text 1 to obtain tokens 110 to 1nn, and word segmentation is performed on text 2 to obtain tokens 210 to 2nn. These tokens are added to the dictionary 602, and the image classification levels (e.g., numbers 1 to 5) are also added to the dictionary 602.
[0067] In some embodiments, the granularity of the dictionary is a character. In some embodiments, the granularity of the dictionary is a word. In some embodiments, the granularity of the dictionary is a phrase or a word group. It can be understood that these granularities can be determined according to the training effect. For example, the tokens used for training are generated using different word segmentation criteria. These tokens with different granularities can be used in combination. For example, the granularity of a word or a phrase can be compatible with the granularity of a character. Corresponding to English, the minimum granularity is a single word, and letters are not used as granularities unless a letter is a single word, such as the article "a".
[0068] It can be seen that the dictionary established in this way is specifically for the medical field, so the vocabulary included therein is specific, the quantity is relatively reduced, and the vocabulary is relatively accurate. Such a process 600 can reduce the computational overhead and improve the efficiency of outputting classification results and descriptive text through the computational speed of the image interpretation model.
[0069] In some embodiments, the first computational model includes at least a part of a pre-trained model, and training the image interpretation model using a training data set includes: freezing the parameters of the first computational model; and updating the parameters of the second computational model.
[0070] As an example, the first training model 112 can be some image classification models that have been commercially pre-trained, which can be specifically trained based on the training data set 120 to conform to the specific scenario of medical image classification. In some embodiments, the first computational model can be trained first, its parameters are frozen, and then the second computational model 116 is trained, and the parameters of the second computational model are updated until the requirements are met. Alternatively, in some embodiments, the first computational model can be trained first, its parameters are frozen, and then the second training model is trained, and the parameters of the second computational model are updated. Then, the parameters of the second model are frozen, and the parameters of the first model are updated based on the loss function of the second model. The iterative alternating training is carried out by analogy until the requirements are met.
[0071] In some embodiments, the classification information includes auxiliary diagnostic information for a medical image, and the descriptive information includes a description of one or more objects in the medical image, for example, a description of lesions of different organs of the human body.
[0072] In some embodiments, a trained image interpretation model is used to generate classification information for an input medical image and descriptive text for interpreting the classification information from the input medical image. For example, after the image interpretation model is trained and put into actual use, the trained image interpretation model generates classification information 130 for the input medical image 140 and descriptive information 132 for interpreting the classification information 130 from the input medical image 140.
[0073] Figure 7A schematic diagram showing a comparison 700 between an embodiment of the present disclosure and image captions is shown. As Figure 7 shown, images 702 and 706 are similar medical images. By using the traditional image captioning mode, a sentence such as "LI-RADS Liver Grade 4 Lesion" shown in 704 can be obtained, but the classification reason is unknown. However, by using the image processing solution proposed in the present disclosure, a sentence paragraph such as "LI-RADS Liver Grade 4 Lesion, due to large size, there is an enhanced envelope and non-peripheral washout" shown in 708 can be obtained, where the first sentence summarizes the diagnostic result, and the rest explains why this result is obtained. It can be seen that the effect of description 708 is much better than that of description 704. Description 708 has a detailed comparison reason, and its classification is based on medical criteria.
[0074] Figure 8 A schematic block diagram showing an example device 800 that can be used to implement some embodiments according to the present disclosure is shown. As Figure 8 shown, device 800 includes a central processing unit (CPU) 801, which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) 802 or computer program instructions loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of device 800 are also stored. The CPU 801, ROM 802, and RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0075] A plurality of components in device 800 are connected to the I / O interface 805, including: (one or more) input units 806, such as keyboards, mice, etc.; (one or more) output units 807, such as various types of displays, speakers, etc.; (one or more) storage units 808, such as disks, optical discs, etc.; and (one or more) communication units 809, such as network cards, modems, wireless communication transceivers, etc. The (one or more) communication units 809 allow device 800 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0076] The various processes and treatments described above (such as method 200 or processes 300, 330, 400, 500, and 600) can be executed by processing unit 801. For example, in some embodiments, one or more of method 200, processes 300, 330, 400, 500, and 600 can be implemented as a computer software program tangibly embodied in a machine-readable medium (such as storage unit 808). In some embodiments, all or all of the computer programs can be loaded and / or installed onto device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by CPU 801, one or more of the above-described method 200, processes 300, 330, 400, 500, and 600 can be executed.
[0077] The present disclosure can be a method, apparatus, system, and / or computer program product. The computer program product can include a computer-readable storage medium having thereon loaded computer-readable program instructions for performing various aspects of the present disclosure.
[0078] The computer-readable storage medium can be a tangible device that can store and hold instructions for use by an instruction execution device. The computer-readable storage medium can be, by way of example and not limitation, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (a non-exhaustive list) of the computer-readable storage medium include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disc (DVD), a memory stick, a floppy disk, a mechanically encoded device having instructions stored thereon (such as punched cards in a groove or raised structures), and any suitable combination thereof. As used herein, a computer-readable storage medium should not be construed as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., an optical pulse propagating through an optical fiber cable), or an electrical signal transmitted through a wire.
[0079] The computer-readable program instructions described herein can be downloaded to various computing / processing devices from a computer-readable storage medium or downloaded to an external computer or external storage device via a network such as the Internet, a local area network, a wide area network, and / or a wireless network. The network may include copper transmission cables, optical fiber transmission, wireless transmission, routers, firewalls, switches, gateway computers, and / or edge servers. A network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in a computer-readable storage medium in each computing / processing device.
[0080] The computer program instructions for performing the operations of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-dependent instructions, microcode, firmware instructions, state-setting data, or source code or object code compiled in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The computer-readable program instructions may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network connection, including a local area network (LAN) or a wide area network (WAN), or it may be connected to an external computer (e.g., using an Internet service provider via the Internet). In some embodiments, the state information of the computer-readable program instructions may be used to customize an electronic circuit, such as a programmable logic circuit, a field-programmable gate array (FPGA), or a programmable logic array (PLA). The electronic circuit may execute the computer-readable program instructions to implement various aspects of the present disclosure.
[0081] Here, various aspects of the present disclosure are described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and the combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0082] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, when executed by the processing unit of the computer or other programmable data processing apparatus, generate means for implementing the specified functions / acts in one or more blocks of the flowchart and / or block diagram. These computer-readable program instructions may also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer-readable storage medium storing the instructions comprises an article of manufacture including instructions for implementing various aspects of the specified functions / acts in one or more blocks of the flowchart and / or block diagram.
[0083] The computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, such that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the specified functions / acts in one or more blocks of the flowchart and / or block diagram.
[0084] The flowcharts and block diagrams in the figures illustrate the system architectures, functions, and operations that can be implemented by systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowcharts and block diagrams may represent a module, a segment of a program, or a part of an instruction, and the module, segment of a program, or part of an instruction contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the blocks may occur in an order different from that noted in the figures. For example, depending on the functions involved, two consecutive blocks may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system for performing the specified functions or acts, or it can be implemented by a combination of dedicated hardware and computer instructions.
[0085] Various embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and not limited to the various disclosed embodiments. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the various described embodiments. The selection of the terms used herein is intended to best explain the principles of the various embodiments, the practical application, or the improvement of the technology in the market, or to enable other ordinary skill in the art to understand the various embodiments disclosed herein.
Claims
1. A method for image processing, comprising: obtaining a training dataset including medical images and corresponding annotation texts, wherein the annotation texts include classification information of the medical images and descriptive information associated with the classification information; and using the training dataset to train an image interpretation model, wherein the image interpretation model includes a first computational model for extracting image features from the medical images and a second computational model for generating text from the image features.
2. The method according to claim 1, wherein, the descriptive information is applicable to interpreting the classification information based on medical criteria.
3. The method according to claim 1, wherein, using the first computational model to extract the image features from the medical images includes: determining a set of image channel weights associated with the image channels based on the image channels of the medical images; determining a set of spatial weights associated with space based on the spatial distribution of objects in the medical images; and determining the image features based on the set of image channel weights and the set of spatial weights.
4. The method according to claim 1, wherein, the second computational model includes a sequence-to-sequence model, and wherein the sequence-to-sequence model includes a plurality of prediction units connected in series, and each of the prediction units is configured to output a predicted word or phrase.
5. The method according to claim 4, wherein, using the training dataset to train the image interpretation model includes: for a prediction unit among the plurality of prediction units connected in series, receiving, by the prediction unit, a word or phrase generated by a previous prediction unit as input; and outputting, by the prediction unit, the predicted word or phrase to a next prediction unit.
6. The method according to claim 5, wherein, receiving, by the prediction unit, a word or phrase generated by a previous prediction unit as input includes: in a first prediction unit, determining a first semantic feature associated with the image features based on the image features and attention to the image features; determining a probability distribution of the word or phrase output from the first prediction unit based on the first semantic feature; in a second prediction unit, generating a second semantic feature associated with the first semantic feature based on the first semantic feature and attention to the first semantic feature; and determining a probability distribution of the word or phrase output from the second prediction unit based on the second semantic feature.
7. The method according to claim 4, wherein, using the training dataset to train the image interpretation model includes: performing word segmentation on the texts in the training dataset to obtain a dictionary for generating the texts; and training the second computational model based on the dictionary so that words or phrases in the texts generated by the second computational model are included in the dictionary.
8. The method according to claim 1, wherein, the first computational model includes at least a part of a pre-trained model, and using the training dataset to train the image interpretation model includes: freezing the parameters of the first computational model; and Update the parameters of the second computing model.
9. The method according to claim 1, wherein, the classification information includes auxiliary diagnosis information of the medical image, and the descriptive information includes a description of one or more objects in the medical image.
10. The method according to claim 1, further comprising: using a trained image interpretation model to generate classification information of the input medical image and descriptive text for interpreting the classification information from the input medical image.
11. An electronic device, comprising: at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for being run by the at least one processing unit, wherein the instructions, when run by the at least one processing unit, cause the computing device to execute a method comprising the following operations: obtaining a training data set including a medical image and corresponding annotation text, wherein the annotation text includes classification information of the medical image and descriptive information associated with the classification information; and using the training data set to train an image interpretation model, wherein the image interpretation model includes a first computing model for extracting image features from the medical image and a second computing model for generating text from the image features.
12. The electronic device according to claim 11, wherein, the descriptive information is suitable for interpreting the classification information based on medical criteria.
13. The electronic device according to claim 11, wherein, using the first computing model to extract the image features from the medical image includes: determining a set of image channel weights associated with the image channels based on the image channels of the medical image; determining a set of spatial weights associated with space based on the spatial distribution of objects in the medical image; and determining the image features based on the set of image channel weights and the set of spatial weights.
14. The electronic device according to claim 11, wherein, the second computing model includes a sequence-to-sequence model, and wherein the sequence-to-sequence model includes a plurality of prediction units connected in series, and each of the prediction units is configured to output a predicted word or phrase.
15. The electronic device according to claim 14, wherein, using the training data set to train the image interpretation model includes: for a prediction unit among the plurality of prediction units connected in series, receiving, by the prediction unit, a word or phrase generated by a previous prediction unit as input; and outputting, by the prediction unit, the predicted word or phrase to a next prediction unit.
16. The electronic device according to claim 15, wherein, receiving, by the prediction unit, a word or phrase generated by a previous prediction unit as input includes: in a first prediction unit, determining a first semantic feature associated with the image features based on the image features and attention to the image features; determining a probability distribution of the word or phrase output from the first prediction unit based on the first semantic feature; In a second prediction unit, a second semantic feature associated with the first semantic feature is generated based on the first semantic feature and attention to the first semantic feature; and Based on the second semantic feature, a probability distribution of the word or phrase output from the second prediction unit is determined.
17. The electronic device according to claim 14, wherein, Training the image interpretation model using the training data set includes: Performing word segmentation on the text in the training data set to obtain a dictionary for generating the text; and Training the second computing model based on the dictionary so that words or phrases in the text generated by the second computing model are included in the dictionary.
18. The electronic device according to claim 11, wherein, The first computing model includes at least a part of a pre-trained model, and training the image interpretation model using the training data set includes: Freezing the parameters of the first computing model; and Updating the parameters of the second computing model.
19. The electronic device according to claim 11, wherein, The classification information includes auxiliary diagnosis information of the medical image, and the descriptive information includes a description of one or more objects in the medical image.
20. The electronic device according to claim 11, further comprising instructions that, when run by the at least one processing unit, cause the computing device to perform: Generating classification information of the input medical image and descriptive text for interpreting the classification information from the input medical image using the trained image interpretation model.
21. A computer program product comprising machine-executable instructions, wherein, The machine-executable instructions, when run by a device, cause the device to perform the method according to any one of claims 1-10.