Image Processing Method, Apparatus, Storage Medium, and Electronic Device
By converting the regression prediction layer of the convolutional neural network model into a classification prediction layer and using the class activation mapping algorithm for feature visualization, the parameter characteristics problem that is difficult to explain in the application of convolutional neural network in the medical image field is solved, and the visual interpretation of the regression prediction model and the universality of the feature visualization scheme is achieved.
Patent Information
- Application Number
- CN202210693405.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-17
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-06-17
AI Technical Summary
The application of convolutional neural networks in the field of medical images is limited by its difficult-to-explain parameter characteristics, and its training basis cannot be evaluated, which limits its implementation value in more fields.
By converting the regression prediction layer of the regression prediction model into a classification prediction layer, a second convolutional neural network model is obtained, and the target image is forward propagated according to the model to obtain the first feature map output from the target convolution layer. Then, the first feature map is visualized according to the class activation mapping algorithm to obtain the class activation mapping map of the target image.
The visual interpretation of the regression prediction model is realized, the universality of the feature visualization scheme is improved, and the key areas of the model's decision-making basis is reflected.
Smart Images

Figure CN115049886B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular, to an image processing method, apparatus, storage medium, and electronic device. Background Art
[0002] Convolutional neural networks have achieved great success in tasks such as image classification, object detection, and semantic segmentation. However, due to their difficult-to-interpret parameter characteristics and the inability to evaluate their training basis, etc., factors have hindered their implementation value in more fields. For example, in the field of medical images, clinical scenarios require evaluating the prediction stability of the model, the decision-making basis, and even the confidence interval of the prediction results.
[0003] In order to achieve visual interpretation of convolutional neural networks, some current feature visualization schemes have been proposed. However, conventional feature visualization schemes have many restrictions on the model and can only be used for classification tasks. Summary of the Invention
[0004] Embodiments of this application provide an image processing method, apparatus, storage medium, and electronic device, which can achieve visual interpretation of a regression prediction model and improve the universality of the feature visualization scheme.
[0005] In a first aspect, an embodiment of this application provides an image processing method, including:
[0006] Determine a trained first convolutional neural network model, where the first convolutional neural network model includes multiple network layers that can be differentiated backward;
[0007] If the first convolutional neural network model is a regression prediction model, convert the regression prediction layer of the first convolutional neural network model into a classification prediction layer to obtain a second convolutional neural network model;
[0008] Perform forward propagation calculation on a target image according to the second convolutional neural network model to obtain a first feature map output by a target convolutional layer;
[0009] Perform feature visualization processing on the first feature map according to the class activation mapping algorithm to obtain a class activation mapping map of the target image.
[0010] In a second aspect, an embodiment of this application further provides an image processing apparatus, including:
[0011] A determination module, configured to determine a trained first convolutional neural network model, where the first convolutional neural network model includes multiple network layers that can be differentiated backward;
[0012] A conversion module, configured to, if the first convolutional neural network model is a regression prediction model, convert the regression prediction layer of the first convolutional neural network model into a classification prediction layer to obtain a second convolutional neural network model;
[0013] A computing module, configured to perform forward propagation calculation on a target image according to the second convolutional neural network model to obtain a first feature map output by a target convolutional layer;
[0014] An interpretation module, configured to perform feature visualization processing on the first feature map according to a class activation mapping algorithm to obtain a class activation mapping map of the target image.
[0015] In a third aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program runs on a computer, the computer is enabled to execute the image processing method provided in any embodiment of the present application.
[0016] In a fourth aspect, an embodiment of the present application further provides an electronic device, including a processor and a memory. The memory has a computer program, and the processor is configured to execute the image processing method provided in any embodiment of the present application by calling the computer program.
[0017] For the first convolutional neural network model obtained by training provided in the embodiments of the present application, the model includes multiple network layers that can be traced back. If the model is a regression prediction model, the regression prediction layer of the model is converted into a classification prediction layer to obtain a second convolutional neural network model. Then, forward propagation calculation is performed on the target image according to the second convolutional neural network model to obtain a first feature map output by a target convolutional layer. Feature visualization processing is performed on the first feature map according to a class activation mapping algorithm to obtain a class activation mapping map of the target image. The class activation mapping map can reflect the key area of the model's decision-making basis. The solution of the embodiment of the present application can realize the visualization explanation of the regression prediction model and improve the universality of the feature visualization solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0019] Figure 1 It is a schematic flowchart of an image processing method provided in an embodiment of the present application.
[0020] Figure 2 It is a schematic diagram of an application scenario of the image processing method provided in an embodiment of the present application.
[0021] Figure 3 It is a schematic structural diagram of an image processing device provided in an embodiment of the present application.
[0022] Figure 4 The first structural schematic diagram of the electronic device provided by the embodiment of the present application.
[0023] Figure 5 The second structural schematic diagram of the electronic device provided by the embodiment of the present application. Detailed implementation manners
[0024] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.
[0025] Referring to "embodiment" herein means that a specific feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the present application. The phrase appears in various places in the specification and does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein can be combined with other embodiments.
[0026] The embodiment of the present application provides an image processing method. The execution subject of the image processing method can be the image processing device provided by the embodiment of the present application, or an electronic device integrated with the image processing device, where the image processing device can be implemented in a hardware or software manner. Among them, the electronic device can be a smart phone, a tablet computer, a handheld computer, a notebook computer, a desktop computer, or other devices.
[0027] Please refer to Figure 1 , Figure 1 which is a flowchart of an image processing method provided by the embodiment of the present application. The specific process of the image processing method provided by the embodiment of the present application can be as follows:
[0028] 101. Determine a trained first convolutional neural network model, where the first convolutional neural network model includes multiple network layers that can be reversely differentiated.
[0029] The first convolutional neural network model is a pre-trained model for processing a certain task. This task can be a classification task or a regression task. The first convolutional neural network model includes multiple backpropagation-capable network layers, such as convolutional layers, prediction layers for processing classification tasks or regression tasks, pooling layers, etc. Among them, for a model used to process a classification task, the output result of its output layer is discrete, which is the category to which the object belongs, such as a cat, a dog, etc. For a model used to process a regression task, the output result of its output layer is continuous, which is the value of the object, and this value is generally within a certain range. For example, the first convolutional neural network model is trained to estimate the height of a person in an image through the image. Since the height of a person is generally within a specific range and is a continuous value, this task is a regression task. Then the output layer of the first convolutional neural network model is a mapping unit for outputting the regression result.
[0030] In addition, the first convolutional neural network model may further include a pooling layer, a fully connected layer, or other backpropagation-capable modules, etc.
[0031] 102. If the first convolutional neural network model is a regression prediction model, then convert the regression prediction layer of the first convolutional neural network model into a classification prediction layer to obtain a second convolutional neural network model.
[0032] After obtaining the first convolutional neural network model, determine the type of this model. If this model is a regression prediction model, then the regression prediction layer of this model needs to be converted into a classification prediction layer. As mentioned above, the regression prediction layer generally uses a regression algorithm to output a value within a specific range. Based on this principle, in some embodiments, the specific range can be divided into intervals, and then the calculation of the regression value can be converted into the calculation of the probability that the regression value falls within the corresponding interval to increase the output dimension, and then the regression prediction layer is converted into a classification prediction layer.
[0033] For example, in one embodiment, if the first convolutional neural network model is a regression prediction model, then converting the regression prediction layer of the first convolutional neural network model into a classification prediction layer to obtain a second convolutional neural network model includes: if the first convolutional neural network model is a regression prediction model, then determine the output value interval of each output dimension of the regression prediction layer of the first convolutional neural network model; for each output dimension, divide the output value interval of the output dimension into multiple sub-intervals, and correspond each sub-interval to a classification category of the output dimension to obtain multiple classification categories; generate a classification prediction layer according to the multiple classification categories; use the classification prediction layer to replace the regression prediction layer to obtain a second convolutional neural network model.
[0034] For example, the output of the regression prediction layer is a vector of size L×1, where L represents the total number of dimensions of the regression. The elements in this vector represent the predicted regression values for each output dimension. Determine the output value range for each output dimension, and then divide the output value range of this output dimension into multiple sub-ranges. Among them, one sub-range corresponds to the first classification category of this output dimension, and multiple classification categories of this output dimension are obtained. For example, for the predicted regression value y i There is a maximum value y max and a minimum value y min , then y i ∈[y min ,y max , that is, the output value range. Divide this output value range into K sub-ranges: y i ∈{[y min ,y t1 )∪[y t1 ,y t2 )∪…∪[y t(K-1) ,y max}. Among them, the value of K is an empirical value, and a reasonable value can be determined through testing. Among them, the above-mentioned division of K sub-ranges can be a uniform division, or can be scaled proportionally according to the probability density function of the predicted regression value, or can also be a custom sampling rule. After the above division, the original regression prediction layer is converted into a classification prediction layer, and the output size is converted from a vector of L×1 to a tensor of L×K. Each element in the tensor represents the probability that the regression value output by the model in the corresponding output dimension falls into the corresponding sub-range. That is, for the L output dimensions, each dimension changes from originally outputting one regression value to outputting the probability that this regression value falls into the K sub-ranges.
[0035] After obtaining multiple classification categories for each output dimension in the above manner, for each output dimension, generate a classification prediction layer according to these multiple classification categories. For example, construct an initial classification prediction layer based on multiple classification categories. For example, modify the regression prediction layer into a fully connected layer and set a softmax function behind it to generate the probability of the corresponding category, and obtain the initial classification prediction layer.
[0036] After constructing the initial classification prediction layer in the above manner, the weight parameters in this initial classification prediction layer are initial parameters, and the classification prediction layer needs to be trained to update the weight parameters. For example, in one embodiment, generating a classification prediction layer according to multiple classification categories includes: constructing an initial classification prediction layer based on multiple classification categories; training the initial classification prediction layer according to the sample images to obtain the classification prediction layer.
[0037] In this embodiment, the initial classification prediction layer is trained based on the sample image to obtain the classification prediction layer. For example, after replacing the regression prediction layer in the first convolutional neural network model with the above initial classification prediction layer, the weight parameters of other layers except the initial classification prediction layer are frozen. The sample image is used as the training data, and the model is retrained using the loss function to determine the weight parameters of the initial classification prediction layer, thereby obtaining the classification prediction layer and the second convolutional neural network model. That is to say, the weight parameters of other network layers in the second convolutional neural network model except the classification prediction layer are the same as those in the first convolutional neural network model.
[0038] 103. Perform forward propagation calculation on the target image according to the second convolutional neural network model to obtain the first feature map output by the target convolutional layer.
[0039] After the output dimension conversion process described above, the second convolutional neural network model is obtained. Next, the target image is acquired, and forward propagation calculation is performed on the target image according to the second convolutional neural network model. Among them, the target image refers to the object to be subjected to feature visualization processing. The target image can be an image captured by the camera of the electronic device itself or an image received from other terminals. After the target image is acquired, the target image is input into the second convolutional neural network model for calculation. After the calculation is completed, the first feature map output by the target convolutional layer is obtained.
[0040] In addition, it should be noted that the target image in the embodiment of the present application can be a single-frame image or an image sequence. For example, the target image can be a CT (Computed Tomography) three-dimensional image, a video frame sequence, etc. Denote the size of the target image as D×H×W×C, where D represents the number of image layers. If the target image is a single-frame 2D image, then D = 1. If the target image is a 3D image such as a CT three-dimensional image or a video frame sequence, then D≥2. H represents the image height; W represents the image width. C represents the number of image channels. For example, if the target image is a grayscale image, then C = 1. If the target image is an RGB image, then C = 3. If the target image is an RGBA image, then C = 4.
[0041] Among them, the target convolutional layer can be set as needed and can be one or more. For example, in one embodiment, the last convolutional layer of the second convolutional neural network model is determined as the target convolutional layer. Or, in another embodiment, the last three convolutional layers of the second convolutional neural network model are all used as the target convolutional layer, and the first feature map output by each target convolutional layer is obtained.
[0042] 104. Perform feature visualization processing on the first feature map according to the class activation mapping algorithm to obtain the class activation mapping of the target image.
[0043] After feature extraction processing through the convolutional layer, the obtained feature map contains information of different categories. The purpose of feature visualization is to present this information in a visual way so that users can intuitively see it.
[0044] After obtaining the first feature map output by the target convolutional layer, next, feature visualization processing is performed based on the obtained first feature map. In the embodiments of the present application, the class activation mapping algorithm is used to implement feature visualization processing. For example, in one embodiment, the LayerCAM (Layer Class Activation Mapping) algorithm is used to perform feature visualization processing on the first feature map to obtain the class activation mapping map of the target image.
[0045] In addition, the classification prediction layer generally has multiple categories. After the model calculates the target image, the probability of the target image in each category can be obtained, and the probability values in different categories generally vary. Therefore, the visualization effects presented by the first feature map in different categories will also be different. One or more categories among the multiple classifications of the classification prediction layer can be used as the target category.
[0046] Then, for the target category, obtain the prediction probability of the classification prediction layer of the second convolutional neural network model on the target category; calculate the partial derivative of the prediction probability with respect to each pixel of the first feature map, and input the partial derivative of each pixel into the activation function for processing to obtain the channel weight corresponding to each pixel; according to the pixel value and channel weight of each pixel, calculate the class activation mapping map of the target image on the target category.
[0047] Based on the class activation mapping algorithm, a forward propagation operation is performed on the target image with a size of D×H×W×C, and the prediction probability of each category in the classification prediction layer can be obtained, and the prediction probability on the target category is determined from them. Among them, obtain the value Y of the target category C before the softmax operation in the classification prediction layer c As the prediction probability corresponding to the target category C. After obtaining the prediction probability, perform backpropagation differentiation according to the chain rule to obtain the weights corresponding to each pixel point in each feature map, and then calculate the class activation mapping map according to the weights. The specific calculation formula is as follows:
[0048] Among them, the value of the pixel at coordinate P on the class activation mapping map of the target image on the target category C Is expressed as:
[0049]
[0050] Among them, A q(P) represents the pixel value at the P position in the q-th channel of the feature map A, where P = (i, j, k) and P ∈ R H ′×W′×D′ , represents the weight corresponding to the pixel at the P position in the q-th channel of the feature map A. H′ represents the height of the feature map, W′ represents the width of the feature map, and D′ represents the number of layers of the feature map.
[0051] Among them, the weight is:
[0052]
[0053] Calculate each pixel value on the class activation map according to the above formula, thereby obtaining the class activation map.
[0054] Among them, it can be understood that for each first feature map, the corresponding class activation map can be calculated. In one embodiment, these multiple class activation maps can be output and displayed. Or, in another embodiment, when there are multiple target convolutional layers, the first feature maps output by each target convolutional layer are respectively subjected to feature visualization processing according to the class activation mapping algorithm to obtain multiple class activation maps; an average operation is performed on the multiple class activation maps to obtain the class activation map of the target image.
[0055] The class activation map calculated according to the above method has a size of D′×H′×W′. In order to better display the class activation map subsequently, the size of the class activation map can be enlarged to the size D×H×W of the original input image through an upsampling operation, obtaining Among them, the upsampling operation can be bilinear interpolation upsampling, nearest neighbor interpolation upsampling, or cubic interpolation upsampling.
[0056] This solution can be applied not only to models for processing regression tasks but also to models for processing classification tasks. For example, in some embodiments, after step 101, the method further includes: if the first convolutional neural network model is a classification prediction model, then perform the step of forward propagation calculation on the target image according to the second convolutional neural network model to obtain the first feature map output by the target convolutional layer.
[0057] In this embodiment, if it is determined that the first convolutional neural network model is a classification prediction model, there is no need to perform conversion processing on the output dimension anymore, but directly execute step 103 to perform forward propagation calculation on the target image to obtain the first feature map output by the target convolutional layer.
[0058] In specific implementation, the present application is not limited by the execution order of the described steps. Without conflict, some steps can be performed in other orders or simultaneously.
[0059] As can be seen from the above, for the first convolutional neural network model obtained by training in the image processing method provided by the embodiments of the present application, if the model is a regression prediction model, the regression prediction layer of the model is converted into a classification prediction layer to obtain a second convolutional neural network model, and then forward propagation calculation is performed on the target image according to the second convolutional neural network model to obtain a first feature map output by the target convolutional layer. Feature visualization processing is performed on the first feature map according to the class activation mapping algorithm to obtain a class activation mapping map of the target image. The class activation mapping map can reflect the key area of the model's decision-making basis. The solution of the embodiments of the present application can realize the visualization explanation of the regression prediction model and improve the universality of the feature visualization solution.
[0060] Among them, in one embodiment, after forward propagation calculation is performed on the target image according to the second convolutional neural network model to obtain a first feature map output by the target convolutional layer, the method further includes: performing augmentation processing on the target image to obtain an augmented image; performing forward propagation calculation on the augmented image according to the second convolutional neural network model to obtain a second feature map output by the target convolutional layer; performing feature visualization processing on the first feature map according to the class activation mapping algorithm to obtain a class activation mapping map of the target image, including: performing feature visualization processing on the first feature map according to the class activation mapping algorithm to obtain a first class activation mapping map; performing feature visualization processing on the second feature map according to the class activation mapping algorithm to obtain a second class activation mapping map; determining the class activation mapping map of the target image according to the first class activation mapping map and the second class activation mapping map.
[0061] In this embodiment, in order to improve the positioning ability of the key area in the image of the present solution, augmentation processing is performed on the target image to simulate the interference brought by multi-scale and multi-noise conditions. For example, after obtaining the target image, augmentation processing is performed on the relevant target image. For example, the values of each pixel point of the target image are normalized; Gaussian blur processing and / or Gaussian noise addition processing are performed on the normalized target image to obtain an augmented image. For example, it can be adding random Gaussian noise that satisfies the distribution following N(0,σ 2 ) pixel by pixel, where σ = 1, or performing Gaussian blur operation using a Gaussian filter with a filter size of (k1,…,k G ) and filter parameters set to follow a multivariate Gaussian distribution N(0,∑), where Σ represents the covariance matrix of the multivariate Gaussian distribution, which can be a randomly generated diagonal matrix or a matrix with irrelevant H and W dimensions and relevant D and H, W dimensions.
[0062] Suppose the original target image undergoes N kinds of augmentation processes to obtain N augmented images. Together with the original target image, they form an image set consisting of N + 1 images. For the augmented images, following the same processing procedure as above, the corresponding second feature maps of the augmented images are obtained, and the corresponding second type of activation mapping maps of the augmented images are calculated based on the second feature maps. Finally, N + 1 activation mapping maps of this type are obtained. Among them, the generation method of the second type of activation mapping map is similar to that of the first type of activation mapping map in the above text, and will not be elaborated here.
[0063] After obtaining the second type of activation mapping maps, the N + 1 activation mapping maps of this type can be averaged as the final activation mapping map of the target image for output.
[0064] Among them, in some embodiments, after performing feature visualization processing on the first feature map according to the activation mapping algorithm to obtain the activation mapping map of the target image, the method further includes: performing upsampling processing on the activation mapping map to adjust the size of the activation mapping map to the size of the target image; adjusting the transparency of the upsampled activation mapping map according to a preset transparency, and superimposing and displaying the adjusted activation mapping map on the target image. Refer to Figure 2 shown Figure 2 is a schematic diagram of the application scenario of the image processing method provided by the embodiment of the present application.
[0065] This embodiment provides a way to present the activation map. After obtaining the activation mapping map, normalize the pixel values of each activation mapping map to the interval [0, 1]. Then, adjust the transparency of the normalized activation mapping map. For example, adjust its transparency to 50%. Then, when superimposing and displaying the activation mapping map with a certain transparency on the target image, users can intuitively see the key areas of the target image. Or, in another embodiment, a color map can be specified, map the normalized activation mapping map to the color map, and superimpose it on the target image for display.
[0066] Through the solution of the embodiment of the present application, the discriminative explanation of the model prediction result and the ability to locate the key areas of the decision-making basis are realized, and the key areas are visually displayed.
[0067] In one embodiment, an image processing device is also provided. Please refer to Figure 3 , Figure 3 is a schematic structural diagram of the image processing device provided by the embodiment of the present application. The image processing device 300 is applied to an electronic device. The image processing device 300 includes a determination module 301, a conversion module 302, a calculation module 303, and an interpretation module 304, as follows:
[0068] A determination module 301, configured to determine a trained first convolutional neural network model, where the first convolutional neural network model includes multiple network layers that are differentiable backward;
[0069] A conversion module 302, configured to, if the first convolutional neural network model is a regression prediction model, convert the regression prediction layer of the first convolutional neural network model into a classification prediction layer to obtain a second convolutional neural network model;
[0070] A calculation module 303, configured to perform forward propagation calculation on a target image according to the second convolutional neural network model to obtain a first feature map output by a target convolutional layer;
[0071] An interpretation module 304, configured to perform feature visualization processing on the first feature map according to a class activation mapping algorithm to obtain a class activation mapping map of the target image.
[0072] In some embodiments, the image processing apparatus 300 further includes:
[0073] An augmentation module, configured to perform augmentation processing on a target image to obtain an augmented image;
[0074] The calculation module 303 is further configured to: perform forward propagation calculation on the augmented image according to the second convolutional neural network model to obtain a second feature map output by a target convolutional layer;
[0075] The interpretation module 304 is further configured to: perform feature visualization processing on the first feature map according to a class activation mapping algorithm to obtain a first class activation mapping map; perform feature visualization processing on the second feature map according to a class activation mapping algorithm to obtain a second class activation mapping map; and determine the class activation mapping map of the target image according to the first class activation mapping map and the second class activation mapping map.
[0076] In some embodiments, the augmentation module is further configured to: normalize the values of each pixel point of the target image; and perform Gaussian blur processing and / or Gaussian noise addition processing on the normalized target image to obtain an augmented image.
[0077] In some embodiments, the conversion module 302 is further configured to: if the first convolutional neural network model is a regression prediction model, determine the output value interval of each output dimension of the regression prediction layer of the first convolutional neural network model; for each output dimension, divide the output value interval of the output dimension into multiple sub-intervals, and correspond each sub-interval to a classification category of the output dimension to obtain multiple classification categories; generate a classification prediction layer according to the multiple classification categories; and use the classification prediction layer to replace the regression prediction layer to obtain a second convolutional neural network model.
[0078] In some embodiments, the conversion module 302 is further configured to: construct an initial classification prediction layer based on the multiple classification categories; and train the initial classification prediction layer according to the sample images to obtain a classification prediction layer.
[0079] In some embodiments, the explanation module 304 is further configured to: when there are multiple target convolutional layers, perform feature visualization processing on the first feature maps output by each target convolutional layer respectively according to the class activation mapping algorithm to obtain multiple class activation mapping graphs; and perform an averaging operation on the multiple class activation mapping graphs to obtain the class activation mapping graph of the target image.
[0080] In some embodiments, the calculation module 303 is further configured to: if the first convolutional neural network model is a classification prediction model, perform the step of performing forward propagation calculation on the target image according to the second convolutional neural network model to obtain the first feature map output by the target convolutional layer.
[0081] In some embodiments, the explanation module 304 is further configured to: obtain the prediction probability of the classification prediction layer of the second convolutional neural network model on the target category; calculate the partial derivative of the prediction probability with respect to each pixel of the first feature map, and input the partial derivative of each pixel into an activation function for processing to obtain the channel weight corresponding to each pixel; and calculate the class activation mapping graph of the target image on the target category according to the pixel value of each pixel and the channel weight.
[0082] In some embodiments, the value of the pixel at the coordinate P on the class activation mapping graph of the target image on the target category C is expressed as:
[0083]
[0084] where, A q (P) represents the pixel value at the position of the q-th channel and coordinate P of the feature map A, P = (i, j, k), P ∈ R H ′×W′×D′ , represents the weight corresponding to the pixel at the position of the q-th channel and coordinate P of the feature map A, H' represents the height of the feature map, W' represents the width of the feature map, and D' represents the number of layers of the feature map.
[0085] In some embodiments, the image processing apparatus 300 further includes:
[0086] A display module, configured to perform upsampling processing on the class activation mapping graph to adjust the size of the class activation mapping graph to the size of the target image; and, adjust the transparency of the upsampled class activation mapping graph according to a preset transparency, and superimpose and display the adjusted class activation mapping graph on the target image.
[0087] It should be noted that the image processing device provided in the embodiments of the present application and the image processing method in the above embodiments belong to the same concept. Through this image processing device, any method provided in the embodiments of the image processing method can be implemented. The specific implementation process is detailed in the embodiments of the image processing method and will not be elaborated here.
[0088] As can be seen from the above, for the first convolutional neural network model obtained by training in the embodiments of the present application, if the model is a regression prediction model, the regression prediction layer of the model is converted into a classification prediction layer to obtain a second convolutional neural network model, and then the target image is forward propagated according to the second convolutional neural network model to calculate the first feature map output by the target convolutional layer. Feature visualization processing is performed on the first feature map according to the class activation mapping algorithm to obtain the class activation mapping graph of the target image. This class activation mapping graph can reflect the key area of the model's decision-making basis. The solution of the embodiments of the present application can realize the visualization explanation of the regression prediction model and improve the universality of the feature visualization scheme.
[0089] The embodiments of the present application also provide an electronic device. The electronic device can be a smart phone, a tablet computer, or other devices. Please refer to Figure 4 , Figure 4 which is the first structural schematic diagram of the electronic device provided in the embodiments of the present application. The electronic device 400 includes a processor 401 and a memory 402. Among them, the processor 401 is electrically connected to the memory 402.
[0090] The processor 401 is the control center of the electronic device 400, connecting various parts of the entire electronic device through various interfaces and lines, and executing various functions of the electronic device and processing data by running or calling the computer program stored in the memory 402 and calling the data stored in the memory 402, so as to perform overall monitoring of the electronic device.
[0091] The memory 402 can be used to store computer programs and data. The computer program stored in the memory 402 contains instructions that can be executed in the processor. The computer program can form various functional modules. The processor 401 executes various functional applications and data processing by calling the computer program stored in the memory 402.
[0092] In this embodiment, the processor 401 in the electronic device 400 will load the instructions corresponding to the processes of one or more computer programs into the memory 402 according to the following steps, and the processor 401 will run the computer programs stored in the memory 402 to implement various functions:
[0093] Determine the trained first convolutional neural network model, where the first convolutional neural network model includes multiple reversely differentiable network layers;
[0094] If the first convolutional neural network model is a regression prediction model, convert the regression prediction layer of the first convolutional neural network model into a classification prediction layer to obtain a second convolutional neural network model;
[0095] Perform forward propagation calculation on the target image according to the second convolutional neural network model to obtain a first feature map output by the target convolutional layer;
[0096] Perform feature visualization processing on the first feature map according to the class activation mapping algorithm to obtain the class activation mapping diagram of the target image.
[0097] In some embodiments, please refer to Figure 5 , Figure 5 This is the second structural schematic diagram of the electronic device provided by the embodiment of the present application. The electronic device 400 further includes: a radio frequency circuit 403, a display screen 404, a control circuit 405, an input unit 406, an audio circuit 407, a sensor 408, and a power supply 409. Among them, the processor 401 is electrically connected to the radio frequency circuit 403, the display screen 404, the control circuit 405, the input unit 406, the audio circuit 407, the sensor 408, and the power supply 409 respectively.
[0098] The radio frequency circuit 403 is used to receive and transmit radio frequency signals to communicate with network devices or other electronic devices through wireless communication.
[0099] The display screen 404 can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the electronic device. These graphical user interfaces can be composed of images, texts, icons, videos, and any combination thereof.
[0100] The control circuit 405 is electrically connected to the display screen 404 and is used to control the display screen 404 to display information.
[0101] The input unit 406 can be used to receive input digital, character information, or user feature information (such as fingerprints), and generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function controls. Among them, the input unit 406 can include a fingerprint recognition module.
[0102] The audio circuit 407 can provide an audio interface between the user and the electronic device through a speaker and a microphone. Among them, the audio circuit 407 includes a microphone. The microphone is electrically connected to the processor 401. The microphone is used to receive the voice information input by the user.
[0103] The sensor 408 is used to collect external environmental information. The sensor 408 may include one or more of sensors such as an ambient light sensor, an acceleration sensor, and a gyroscope.
[0104] The power supply 409 is used to supply power to each component of the electronic device 400. In some embodiments, the power supply 409 may be logically connected to the processor 401 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system.
[0105] Although not shown in the figure, the electronic device 400 may further include a camera, a Bluetooth module, etc., which will not be elaborated here.
[0106] In this embodiment, the processor 401 in the electronic device 400 will load the instructions corresponding to the processes of one or more computer programs into the memory 402 according to the following steps, and the processor 401 will run the computer programs stored in the memory 402 to implement various functions:
[0107] Determine a trained first convolutional neural network model, where the first convolutional neural network model includes multiple network layers that can be differentiated backward;
[0108] If the first convolutional neural network model is a regression prediction model, convert the regression prediction layer of the first convolutional neural network model into a classification prediction layer to obtain a second convolutional neural network model;
[0109] Perform forward propagation calculation on the target image according to the second convolutional neural network model to obtain a first feature map output by the target convolutional layer;
[0110] Perform feature visualization processing on the first feature map according to the class activation mapping algorithm to obtain the class activation mapping diagram of the target image.
[0111] As described above, the embodiment of the present application provides an electronic device. For the first convolutional neural network model obtained by training, if the model is a regression prediction model, the regression prediction layer of the model is converted into a classification prediction layer to obtain a second convolutional neural network model, and then the forward propagation calculation is performed on the target image according to the second convolutional neural network model to obtain the first feature map output by the target convolutional layer. The class activation mapping algorithm is used to perform feature visualization processing on the first feature map to obtain the class activation mapping diagram of the target image. The class activation mapping diagram can reflect the key area of the model decision basis. The solution of the embodiment of the present application can realize the visualization explanation of the regression prediction model and improve the universality of the feature visualization solution.
[0112] The embodiment of the present application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program runs on a computer, the computer executes the image processing method described in any one of the above embodiments.
[0113] It should be noted that those of ordinary skill in the art can understand that all or part of the steps in the above various methods can be completed by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer-readable storage medium can include but is not limited to: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.
[0114] In addition, the terms "first", "second", "third", etc. in the present application are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or modules is not limited to the listed steps or modules, but some embodiments also include steps or modules that are not listed, or some embodiments also include other steps or modules inherent to these processes, methods, products or devices.
[0115] The above has introduced in detail the image processing method, device, storage medium and electronic device provided by the embodiment of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. An image processing method, characterized in that, Including: Determine a trained first convolutional neural network model, where the first convolutional neural network model includes multiple network layers that are differentiable backward; If the first convolutional neural network model is a regression prediction model, convert the regression prediction layer of the first convolutional neural network model into a classification prediction layer to obtain a second convolutional neural network model; Perform forward propagation calculation on the target image according to the second convolutional neural network model to obtain a first feature map output by the target convolutional layer; Normalize the values of each pixel point of the target image; perform Gaussian blur processing and / or add Gaussian noise processing on the normalized target image to obtain an augmented image; Perform forward propagation calculation on the augmented image according to the second convolutional neural network model to obtain a second feature map output by the target convolutional layer; Perform feature visualization processing on the first feature map according to the class activation mapping algorithm to obtain the class activation mapping map of the target image, including: performing feature visualization processing on the first feature map according to the class activation mapping algorithm to obtain a first class activation mapping map; performing feature visualization processing on the second feature map according to the class activation mapping algorithm to obtain a second class activation mapping map; performing averaging processing according to the first class activation mapping map and the second class activation mapping map to determine the class activation mapping map of the target image.
2. The method according to claim 1, characterized in that, The step of, if the first convolutional neural network model is a regression prediction model, converting the regression prediction layer of the first convolutional neural network model into a classification prediction layer to obtain a second convolutional neural network model includes: If the first convolutional neural network model is a regression prediction model, determine the output value interval of each output dimension of the regression prediction layer of the first convolutional neural network model; For each output dimension, divide the output value interval of the output dimension into multiple sub-intervals, and correspond each sub-interval to a classification category of the output dimension to obtain multiple classification categories; Generate a classification prediction layer according to the multiple classification categories; Use the classification prediction layer to replace the regression prediction layer to obtain a second convolutional neural network model.
3. The method according to claim 2, characterized in that, The step of generating a classification prediction layer according to the multiple classification categories includes: Construct an initial classification prediction layer based on the multiple classification categories; Train the initial classification prediction layer according to the sample image to obtain a classification prediction layer.
4. The method according to claim 1, characterized in that, The step of performing feature visualization processing on the first feature map according to the class activation mapping algorithm to obtain the class activation mapping map of the target image includes: When there are multiple target convolutional layers, perform feature visualization processing on the first feature map output by each target convolutional layer according to the class activation mapping algorithm to obtain multiple class activation mapping maps; Perform an averaging operation on the multiple class activation mapping maps to obtain the class activation mapping map of the target image.
5. The method according to claim 1, characterized in that, After determining the trained first convolutional neural network model, the method further includes: If the first convolutional neural network model is a classification prediction model, execute the step of performing forward propagation calculation on the target image according to the second convolutional neural network model to obtain a first feature map output by the target convolutional layer.
6. The method according to claim 1, characterized in that, Performing feature visualization processing on the first feature map according to the class activation mapping algorithm to obtain the class activation mapping map of the target image includes: Obtaining the prediction probability of the classification prediction layer of the second convolutional neural network model on the target class; Calculating the partial derivatives of the prediction probability with respect to each pixel of the first feature map, and inputting the partial derivatives of each pixel into an activation function for processing to obtain the channel weights corresponding to each pixel; Calculating the class activation mapping map of the target image on the target class according to the pixel values of each pixel and the channel weights.
7. The method according to claim 6, characterized in that, The value of the pixel at coordinate P on the class activation map of the target image for the target class C is denoted as: Among them, A q (P) represents the pixel value at the P position in the q-th channel of the feature map A, P = (i, j, k), P ∈ R H ′×W′×D′ , represents the weight corresponding to the pixel at the P position in the q-th channel of the feature map A, H ′ represents the height of the feature map, W ′ represents the width of the feature map, D ′ represents the number of layers of the feature map.
8. The method according to any one of claims 1 to 7, characterized in that, After performing feature visualization processing on the first feature map according to the class activation mapping algorithm to obtain the class activation mapping map of the target image, the method further includes: Performing upsampling processing on the class activation mapping map to adjust the size of the class activation mapping map to the size of the target image; Adjusting the transparency of the upsampled class activation mapping map according to a preset transparency, and superimposing and displaying the adjusted class activation mapping map on the target image.
9. An image processing apparatus, characterized in that, Including: A determination module for determining a trained first convolutional neural network model, where the first convolutional neural network model includes multiple network layers that can be reversely differentiated; A conversion module for converting the regression prediction layer of the first convolutional neural network model into a classification prediction layer to obtain a second convolutional neural network model if the first convolutional neural network model is a regression prediction model; A calculation module for performing forward propagation calculation on a target image according to the second convolutional neural network model to obtain a first feature map output by a target convolutional layer; An augmentation module for normalizing the values of each pixel point of the target image; performing Gaussian blur processing and / or adding Gaussian noise processing on the normalized target image to obtain an augmented image; The calculation module is further configured to perform forward propagation calculation on the augmented image according to the second convolutional neural network model to obtain a second feature map output by the target convolutional layer; An interpretation module for performing feature visualization processing on the first feature map according to the class activation mapping algorithm to obtain the class activation mapping map of the target image, including: performing feature visualization processing on the first feature map according to the class activation mapping algorithm to obtain a first class activation mapping map; performing feature visualization processing on the second feature map according to the class activation mapping algorithm to obtain a second class activation mapping map; and performing averaging processing according to the first class activation mapping map and the second class activation mapping map to determine the class activation mapping map of the target image.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program runs on a computer, the computer is caused to execute the image processing method according to any one of claims 1 to 8.
11. An electronic device comprising a processor and a memory, the memory storing a computer program, characterized in that, The processor is configured to execute the image processing method according to any one of claims 1 to 8 by calling the computer program.
Citation Information
Patent Citations
Class activation mapping target positioning method and system based on convolutional neural network
CN112465909A