Image processing method, device, equipment and storage medium
By acquiring the high-response area and target foreground of the image and adjusting the pixel value of the interfering area in the image, the problem of degradation of target recognition accuracy caused by interference information in image information retrieval is solved, and a higher target recognition accuracy is achieved.
Patent Information
- Application Number
- CN202210707348.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-21
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-06-21
AI Technical Summary
In image information retrieval, when the image is input into the deep learning target recognition model, it is easily affected by interference information, resulting in a decrease in the accuracy of target recognition.
By acquiring the high-response area and the target foreground of the image to be processed, if the preset conditions are met, the pixel value of the target area outside the target foreground is adjusted so that its influence on feature modeling is less than the first threshold value, thereby removing interference information.
It effectively improves the accuracy of target recognition and avoids the impact of interfering information on the target recognition model.
Smart Images

Figure CN115063587B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to an image processing method, apparatus, device and storage medium. Background Art
[0002] With the rapid development of computer computing, image information retrieval has been widely used in various scenarios. Image information retrieval takes a given picture as input, uses deep learning algorithms to extract the features of the input picture, and then further queries one or more pictures with specific similar features from the database based on the extracted features, as well as related description information. Image information retrieval is often used in scenarios such as face recognition and image search.
[0003] However, in actual use, image information retrieval generally inputs images into deep learning target recognition models for target recognition. The input images often contain interference information (such as irrelevant background parts, etc.), and the target recognition model may be disturbed by this interference information and fail to focus on the target itself, which in turn reduces the accuracy of target recognition. Summary of the invention
[0004] The present application provides an image processing method, apparatus, device and storage medium, which can effectively improve the target recognition accuracy of images.
[0005] In a first aspect, the present application provides an image processing method, the method comprising: obtaining a high response area of an image to be processed, wherein the influence of each pixel in the high response area on feature modeling is greater than a first threshold; obtaining a target foreground of the image to be processed; the target foreground is used to indicate the area where a target object is located in the image to be processed; if the target foreground and the high response area meet preset conditions, adjusting the pixel value of a target area outside the target foreground in the image to be processed, so that the influence of each pixel in the target area on feature modeling is less than the first threshold, and the target area includes a portion of the high response area that does not belong to the target foreground.
[0006] The image processing method provided by the present application obtains the high response area and target foreground in the image to be processed, and when the target foreground and the high response area meet the preset conditions, determines that there is interference in the image to be processed, and then adjusts the pixel value of the target area outside the target foreground in the image to be processed, so that the influence of each pixel point in the target area on the feature modeling is less than the first threshold, so as to remove the interference in the image to be processed. This method improves the situation that the target recognition model is easily interfered by interference information during modeling, and effectively improves the accuracy of target recognition.
[0007] In a possible implementation, obtaining a high response area of the image to be processed includes: obtaining a feature response visualization map of the image to be processed, the feature response visualization map is used to reflect the influence of each pixel point in the image to be processed on the feature modeling; the feature response visualization map includes a response activation value of each pixel point; and the area of the pixel points whose response activation value is greater than a first threshold is regarded as a high response area. It can be understood that the high response area indicates the part of the image to be processed that has a greater influence on the feature modeling, which is conducive to analyzing whether these parts with greater influence are the areas that need to be paid attention to for target recognition, so as to subsequently determine whether there is interference information in the image to be processed.
[0008] In another possible implementation, the preset condition is that the ratio of the area of the overlapped area between the high response area and the target foreground to the area of the high response area is less than the second threshold. It can be understood that the high response area is used to indicate which areas in the image to be processed have a greater impact on feature modeling, and the target foreground is used to indicate which areas in the image to be processed are the areas where feature modeling is mainly desired. Therefore, by combining the high response area and the target foreground, it is possible to effectively analyze whether there is interference information that affects feature modeling, that is, there are too many high response areas outside the target foreground, so as to eliminate these interference information in a targeted manner.
[0009] In another possible implementation, the pixel values of the target area outside the target foreground in the image to be processed are adjusted so that the influence of each pixel point in the target area on the feature modeling is less than a first threshold, including: using a mean blur method to adjust the pixel values of the target area to be consistent so that the influence of each pixel point in the target area on the feature modeling is less than the first threshold.
[0010] In another possible implementation, the target area is: the portion of the high response area that does not belong to the target foreground; or, outside the target foreground in the image to be processed, the area where the boundary of the high response area is expanded outward by a third threshold; or, all areas outside the target foreground in the image to be processed.
[0011] In another possible implementation, obtaining a feature response visualization graph of the image to be processed includes: modeling the image to be processed through deep learning to obtain multiple nodes; one node is the data of a modeling feature vector obtained by modeling the image to be processed in one dimension; determining the activation value response result of each node; the activation value response result is used to indicate the contribution of each area of the image to be processed to the node; determining the weight of each node; the weight of a node is determined by the activation value response result of the node and the activation value response results of all nodes; the activation value response result of each node and the weight of each node are processed by a normalization function and an upsampling function to obtain a feature response visualization graph.
[0012] In another possible implementation, determining the activation value response result of each node includes: modeling the image to be processed through deep learning to obtain a target feature map; the target feature map is any one of the multiple feature maps obtained by modeling the image to be processed; based on the target feature map and the node, using a gradient-weighted class activation map (Grad-CAM) method to determine the activation value response result of each node.
[0013] In another possible implementation, obtaining the target foreground of the image to be processed includes: segmenting the image to be processed based on a semantic segmentation model based on deep learning, or a traditional segmentation algorithm to obtain the target foreground.
[0014] In a second aspect, the present application provides an image processing device, the device comprising: an acquisition module and an adjustment module.
[0015] The acquisition module is used to acquire a high response area of the image to be processed, where the influence of each pixel point in the high response area on feature modeling is greater than a first threshold.
[0016] The acquisition module is also used to acquire the target foreground of the image to be processed; the target foreground is used to indicate the area where the target object is located in the image to be processed.
[0017] The adjustment module is used to adjust the pixel value of the target area outside the target foreground in the image to be processed if the target foreground and the high response area meet preset conditions, so that the influence of each pixel point in the target area on the feature modeling is less than a first threshold, and the target area includes the part of the high response area that does not belong to the target foreground.
[0018] In one possible implementation, the acquisition module is specifically used to obtain a feature response visualization map of the image to be processed, where the feature response visualization map is used to reflect the influence of each pixel point in the image to be processed on feature modeling; the feature response visualization map includes a response activation value of each pixel point; and the area of pixels whose response activation values are greater than a first threshold is regarded as a high response area.
[0019] In another possible implementation, the preset condition may be that a ratio of an area of an overlapped region between the high response region and the target foreground to an area of the high response region is smaller than a second threshold.
[0020] In another possible implementation, the adjustment module is specifically used to adjust the pixel values of the target area to be consistent by using a mean fuzzy method, so that the influence of each pixel point in the target area on the feature modeling is less than a first threshold.
[0021] In another possible implementation, the target area is: the portion of the high response area that does not belong to the target foreground; or, outside the target foreground in the image to be processed, the area where the boundary of the high response area is expanded outward by a third threshold; or, all areas outside the target foreground in the image to be processed.
[0022] In another possible implementation, the acquisition module is specifically used to model the image to be processed through deep learning to obtain multiple nodes; one node is the data of the modeling feature vector obtained by modeling the image to be processed in one dimension; the activation value response result of each node is determined; the activation value response result is used to indicate the contribution of each area of the image to be processed to the node; the weight of each node is determined; the weight of a node is determined by the activation value response result of the node and the activation value response results of all nodes; the activation value response result of each node and the weight of each node are processed by normalization function and upsampling function to obtain a feature response visualization graph.
[0023] In another possible implementation, the device further includes: a determination module. The determination module is used to model the image to be processed through deep learning to obtain a target feature map; the target feature map is any one of the multiple feature maps obtained by modeling the image to be processed; based on the target feature map and the node, the Grad-CAM method is used to determine the activation value response result of each node.
[0024] In another possible implementation, the acquisition module is specifically used to segment the image to be processed based on a semantic segmentation model based on deep learning or a traditional segmentation algorithm to obtain a target foreground.
[0025] In a third aspect, the present application provides an image processing device, comprising: a processor and a memory; the memory stores instructions executable by the processor; when the processor is configured to execute the instructions, the image processing device implements the method of the first aspect above.
[0026] In a fourth aspect, the present application provides a computer-readable storage medium, which includes: computer software instructions; when the computer software instructions are executed in a computer, the computer implements the method of the first aspect.
[0027] In a fifth aspect, the present application provides a computer program product. When the computer program product is run on a computer, it enables the computer to execute the steps of the related method described in the first aspect above to implement the method of the first aspect above.
[0028] In a sixth aspect, the present application provides a chip comprising a processor and an interface, wherein the processor is coupled to a memory via the interface, and when the processor executes a computer program in the memory or an image processing device executes instructions, the method described in the first aspect above is executed.
[0029] The beneficial effects of the second to sixth aspects mentioned above can be referred to the corresponding description of the first aspect and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 A schematic diagram of an application environment of an image processing method provided in this application;
[0031] Figure 2 A flowchart of an image processing method provided in this application;
[0032] Figure 3 A flowchart of the specific steps of obtaining a feature response visualization graph provided in this application;
[0033] Figure 4 A schematic diagram of an image to be processed and a corresponding feature response visualization diagram provided in this application;
[0034] Figure 5 A schematic diagram of a high response area provided for this application;
[0035] Figure 6 A schematic diagram of a target prospect provided for this application;
[0036] Figure 7 A schematic diagram of an overlap area provided for this application;
[0037] Figure 8 A schematic diagram of a target area provided for this application;
[0038] Fig. 9 A schematic diagram of another target area provided for this application;
[0039] Fig.10 A schematic diagram of a processed image provided for this application;
[0040] Fig.11 A schematic diagram of an image processing process provided by this application;
[0041] Fig.12 A schematic diagram of the composition of an image processing device provided in this application;
[0042] Fig.13 A schematic diagram of the composition of an image processing device provided in this application. DETAILED DESCRIPTION
[0043] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0044] It should be noted that, in the embodiments of the present application, words such as "exemplarily" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplarily" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplarily" or "for example" is intended to present related concepts in a specific way.
[0045] In order to facilitate the clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, the words "first", "second", etc. are used to distinguish the same items or similar items with basically the same functions and effects. Those skilled in the art can understand that the words "first", "second", etc. are not limiting the quantity and execution order.
[0046] Since the embodiments of the present application involve a large number of models, each model can be implemented using a deep learning model or a machine learning model, for example, a neural network. For ease of understanding, related concepts such as neural networks are introduced below.
[0047] (1) Neural network (NN)
[0048] A neural network is a machine learning model that simulates the human brain to achieve artificial intelligence-like machine learning technology. The input and output of a neural network can be configured according to actual needs, and the neural network can be trained through sample data to minimize the error between its output and the actual output corresponding to the sample data. A neural network can be composed of neural units, which can be x-shaped. s The output of the operation unit with the intercept 1 as input can be:
[0049]
[0050] Where, s = 1, 2, ... n, n is a natural number greater than 1, W s For x sThe weight of the neural unit, b is the bias of the neural unit. f is the activation function of the neural unit, which is used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into the output signal. The output signal of the activation function can be used as the input of the next convolutional layer. The activation function can be a sigmoid function. A neural network is a network formed by connecting many of the above-mentioned single neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field. The local receptive field can be an area composed of several neural units.
[0051] (2) Deep Neural Networks
[0052] A deep neural network (DNN), also known as a multi-layer neural network, can be understood as a neural network with many hidden layers. There is no special metric for "many" here. From the position of different layers of DNN, the neural network inside DNN can be divided into three categories: input layer, hidden layer, and output layer. Generally speaking, the first layer is the input layer, the last layer is the output layer, and the layers in between are all hidden layers. The layers are fully connected, that is, any neuron in the i-th layer must be connected to any neuron in the i+1-th layer. Although DNN looks complicated, the work of each layer is actually not complicated. Simply put, it is the following linear relationship expression: in, is the input vector, is the output vector, b is the offset vector, W is the weight matrix (also called coefficient), and α() is the activation function. Each layer is just an input vector After such a simple operation, the output vector Since there are many DNN layers, the number of coefficients W and offset vectors b is also large. The definitions of these parameters in DNN are as follows: Take coefficient W as an example: Assume that in a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as The superscript 3 represents the layer number of the coefficient W, while the subscripts correspond to the output third layer index 2 and the input second layer index 4. In summary, the coefficients from the kth neuron in the L-1th layer to the jth neuron in the Lth layer are defined as It should be noted that the input layer does not have a W parameter. In a deep neural network, more hidden layers allow the network to better describe complex situations in the real world. Theoretically, the more parameters a model has, the higher its complexity and the greater its "capacity", which means it can complete more complex learning tasks. Training a deep neural network is the process of learning the weight matrix, and its ultimate goal is to obtain the weight matrix of all layers of the trained deep neural network (the weight matrix formed by many layers of vector W).
[0053] (3) Convolutional neural networks (CNN)
[0054] CNN is a deep neural network with a convolutional structure. Convolutional neural networks contain a feature extractor consisting of a convolutional layer and a subsampling layer. The feature extractor can be regarded as a filter, and the convolution process can be regarded as using a trainable filter to convolve an input image or convolution feature plane (feature map). The convolutional layer refers to the neuron layer in the convolutional neural network that performs convolution processing on the input signal. In the convolutional layer of the convolutional neural network, a neuron can only be connected to some neurons in the adjacent layer. A convolutional layer usually contains several feature planes, and each feature plane can be composed of some rectangularly arranged neural units. The neural units in the same feature plane share weights, and the shared weights here are the convolution kernels. Shared weights can be understood as the way to extract image information is independent of position. The implicit principle is that the statistical information of a part of the image is the same as that of other parts. This means that the image information learned in a part can also be used in another part. Therefore, the same learned image information can be used for all positions on the image. In the same convolution layer, multiple convolution kernels can be used to extract different image information. Generally speaking, the more convolution kernels there are, the richer the image information reflected by the convolution operation.
[0055] The convolution kernel can be initialized in the form of a matrix of random size, and the convolution kernel can obtain reasonable weights through learning during the training process of the convolutional neural network. In addition, the direct benefit of shared weights is to reduce the connections between the layers of the convolutional neural network, while reducing the risk of overfitting.
[0056] As described in the background technology, in the process of target recognition, the image is generally input into the deep learning target recognition model to extract the key target features in the image for target recognition. However, due to the presence of interference information in the image (such as irrelevant background, other objects, etc.), the target recognition model is affected by the interference information during the feature extraction process, resulting in a decrease in the accuracy of target recognition.
[0057] In summary, how to eliminate the influence of interference information in images to improve the accuracy of target recognition is an urgent problem to be solved.
[0058] Under this background technology, an embodiment of the present application provides an image processing method, which obtains a high response area in the image to be processed and a target foreground in the image to be processed for indicating the area where the target object is located, and then, based on the high response area and the target foreground, when it is determined that there is interference information in the image to be processed, adjusts the pixel value of the target area outside the target foreground in the image to be processed, and the target area includes the part of the high response area that does not belong to the target foreground, so that the influence of each pixel point in the target area on feature modeling is greatly reduced, thereby achieving the purpose of removing interference information in the image to be processed, thereby avoiding the situation in which the target recognition model easily extracts interference information as a valid feature during the target recognition process and causing data pollution, thereby improving the accuracy of target recognition.
[0059] The solution provided by this application is described below in conjunction with the accompanying drawings of the specification.
[0060] The image processing method provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Figure 1 As shown, the application environment includes two parts: the target recognition model processes the image and outputs the features. Exemplarily, the image processing method can be implemented in the process of the target recognition model processing the image.
[0061] The image processing method provided in the embodiment of the present application can be performed by an image processing device. For example, the image processing device can be a server or an electronic device or other. The embodiment of the present application does not limit the product form of the device that executes the scheme of the present application. Among them, the server mentioned here can be a server cluster composed of multiple servers, or a single server, or a computer. The image processing device can specifically be a processor or a processing chip in the server. The embodiment of the present application does not limit the specific device form of the above-mentioned server. For another example, the image processing device can be a mobile phone, a tablet computer, a desktop, a laptop, a handheld computer, a notebook computer, an ultra-mobile personal computer (ultra-mobile personal computer, UMPC), a netbook, as well as a cellular phone, a personal digital assistant (personal digital assistant, PDA), an augmented reality (augmented reality, AR)\virtual reality (virtual reality, VR) device, etc. The present application does not impose any special restrictions on the specific form of the image processing device. The image processing device is used as an example for explanation below.
[0062] Figure 2 The following is a flowchart of an image processing method provided in an embodiment of the present application. Exemplarily, the image processing method provided in the present application can be applied in the above application environment.
[0063] like Figure 2 As shown, the image processing method provided by the present application may specifically include the following steps:
[0064] S201: An image processing apparatus obtains a high response area of an image to be processed.
[0065] The influence of each pixel point in the high response area on the feature modeling is greater than a first threshold value. The first threshold value may be set in advance based on actual experience values, and is a threshold value at which the influence of a certain pixel point on the feature modeling cannot be ignored.
[0066] Exemplarily, the influence of a pixel point on feature modeling may be a response activation value in a feature response visualization graph.
[0067] In some embodiments, when it is necessary to perform target recognition on a certain image to be processed, the image processing device can obtain a high response area in the image to be processed where the influence of the pixel points on feature modeling is greater than a first threshold. The influence of the pixel points in the image to be processed can be determined using a feature response visualization technology.
[0068] It should be noted that if the influence is greater than the first threshold, it is determined as a high response area. If the influence is less than the first threshold, it is determined as a low response area. If the influence is equal to the first threshold, it can be configured according to actual needs. Whether it is equal to the first threshold belongs to the category of determining a high response area or the category of determining a low response area. The embodiment of the present application does not make specific restrictions on this.
[0069] It can be understood that the high response area indicates the part of the image to be processed that has a greater impact on feature modeling, which is helpful for analyzing whether these parts with greater impact are the areas that need to be paid attention to for target recognition, so as to subsequently determine whether there is interference information in the image to be processed.
[0070] Specifically, Figure 3 As shown, the specific steps of the image processing device in S201 acquiring the high response area of the image to be processed may include the following S201a-S201b.
[0071] S201a, the image processing device obtains a feature response visualization graph of the image to be processed.
[0072] The feature response visualization graph is used to reflect the influence of each pixel in the image to be processed on the feature modeling. The feature response visualization graph includes the response activation value of each pixel, which directly reflects the influence on the feature modeling.
[0073] In some embodiments, the image processing device may obtain a feature response visualization diagram of the image to be processed. Figure 4 The image to be processed and its corresponding feature response visualization diagram provided in the embodiment of the present application are as follows: Figure 4 As shown in the figure, it can be seen that the feature response visualization is similar to a heat map. In the feature response visualization, the brightness of each area reflects the influence of each area. The closer to the center of the image, the brighter the color (that is, the closer to white), indicating a greater influence. The closer to the edge of the image, the darker the color (that is, the closer to black), indicating a smaller influence.
[0074] Exemplarily, the specific steps of obtaining a feature response visualization graph of the image to be processed may include the following S1-S4.
[0075] S1. The image processing device models the image to be processed through deep learning to obtain multiple nodes.
[0076] Among them, a node is the data of a modeling feature vector in one dimension obtained by modeling the image to be processed.
[0077] The image processing device can use the target recognition model in deep learning to model the image to be processed to obtain a modeling feature vector. The modeling feature vector includes multiple nodes, such as N nodes. N is the feature dimension output by the target recognition model, so one node represents the data of the modeling feature vector in one dimension.
[0078] S2. The image processing device determines the activation value response result of each node.
[0079] Among them, the activation value response result is used to indicate the contribution of each area in the image to be processed to the node.
[0080] Specifically, the image processing device models the processed image through deep learning to obtain a target feature map. The target feature map is any one of the multiple feature maps obtained by processing the processed image using a convolutional neural network in deep learning. Since each layer of feature maps extracted using a convolutional neural network can retain the spatial information of the image, and as the number of convolutions increases, the more abstract the feature maps of the later layers are, the richer the semantic information is. Therefore, the embodiment of the present application uses the last layer of feature maps output by the convolutional neural network to determine the feature response visualization map.
[0081] Furthermore, the image processing device can use the Grad-CAM method to determine the activation value response result of each node based on the target feature map and the nodes, so as to determine the contribution of each area in the image to be processed to the node.
[0082] Among them, the calculation formula of the Grad-CAM method satisfies the following expression:
[0083]
[0084] In the above formula, L(f a ) represents the activation value response result of the node, f represents the aforementioned modeling feature vector, f a represents the ath node. Norm is the normalization function in the field of deep learning, and Relu is the linear rectification function. m and n represent the length and width of the target feature map respectively. represents the partial derivative of the corresponding value, which can be obtained by calculating the gradient through the neural network back propagation method. k is the kth channel of the target feature map, is the value at position (i, j) on the k-th channel feature map.
[0085] For the nodes representing each dimension data in the modeling feature vector, the above Grad-CAM formula is used to determine the activation value response result of each node.
[0086] S3. The image processing device determines the weight of each node.
[0087] The weight of a node is determined by the activation value response result of the node and the activation value response results of all nodes. The weight indicates the importance of each node for feature modeling.
[0088] After determining the activation value response result of each node, the image processing device can determine the weight of each node. The process of determining the weight satisfies the following expression:
[0089]
[0090] Among them, w a represents the weight of node a, e is a natural constant, and f p represents the p-th node.
[0091] S4. The image processing device processes the activation value response result of each node and the weight of each node through a normalization function and an upsampling function to obtain a feature response visualization graph.
[0092] After determining the activation value response result and weight of each node, the image processing device processes the activation value response result of each node and the weight of each node through a normalization function and an upsampling function to obtain a feature response visualization graph.
[0093] Specifically, the processing process satisfies the following expression:
[0094]
[0095] Among them, F is the feature response visualization, which includes the activation value of each pixel, ranging from 0 to 1. The larger the value, the greater the influence of the pixel in the image to be processed on the feature modeling. Upsample is an upsampling function, such as nearest neighbor upsampling or bilinear upsampling. It can be understood that the feature response visualization is a digital matrix with the same resolution as the image to be processed, and each number is between 0 and 1.
[0096] S201b: The image processing apparatus regards a region of pixels whose response activation values are greater than a first threshold as a high response region.
[0097] In some embodiments, after obtaining the feature response visualization graph, the image processing device can, according to a pre-set first threshold, take the area consisting of pixels whose response activation values are greater than the first threshold as a high response area, which is subsequently used to determine whether there is interference information in the image to be processed.
[0098] For example, combined with Figure 4 The characteristic response visualization diagram shown in FIG. 1 is used to illustrate that the image processing device sets a first threshold value, and the determined high response area can be as follows Figure 5 The remaining darker areas indicate areas with low influence on feature modeling and can be ignored.
[0099] S202: The image processing device obtains a target foreground of the image to be processed.
[0100] The target foreground is used to indicate the area where the target object is located in the image to be processed. The so-called foreground is a concept in the field of images, which is opposite to the background. It can be said that the image is composed of the foreground part and the background part. For example, in an image containing a portrait, the portrait is the foreground, and the part outside the portrait is the background.
[0101] In some embodiments, the image processing device may obtain a target foreground in the image to be processed where a target object is located. The target object is an object to be identified in the image to be processed, such as a portrait.
[0102] There are many ways to obtain the target foreground. Specifically, in the embodiment of the present application, the image processing device can segment the image to be processed based on a semantic segmentation model of deep learning or a traditional segmentation algorithm to obtain the target foreground.
[0103] For example, Figure 6 As shown in FIG, the target foreground segmented from the image to be processed is the white area in the foreground image shown in the figure. It can be seen that the white area is approximately the area surrounded by the outline of the portrait part in the image to be processed.
[0104] The high response area is used to indicate which areas in the image to be processed have a greater impact on feature modeling, and the target foreground is used to indicate which areas in the image to be processed are the areas where feature modeling is mainly desired. Therefore, by combining the high response area and the target foreground, it is possible to effectively analyze whether there is interference information that affects feature modeling, so as to eliminate the interference information in a targeted manner. The specific determination of whether there is interference information is shown in S203 below.
[0105] S203. If the target foreground and the high response area meet preset conditions, the image processing device adjusts the pixel value of the target area outside the target foreground in the image to be processed so that the influence of each pixel point in the target area on the feature modeling is less than a first threshold.
[0106] The preset condition is used to indicate that interference information exists in the image to be processed.
[0107] It should be noted that the preset conditions can be used to indicate that there are too many high-response areas outside the target foreground area in the image to be processed. The content of the preset conditions can be configured according to actual needs. The embodiments of the present application are not limited to this, as long as it can indicate that there are too many high-response areas outside the target foreground.
[0108] In some embodiments, if the target foreground and the high response area meet the preset conditions, it is determined that there is interference information in the image to be processed, and the operation of adjusting the pixel value in S203 below is performed. Conversely, if the target foreground and the high response area do not meet the preset conditions, it is determined that there is no interference information in the image to be processed, and the image to be processed does not need to be processed, and the original features can be directly used for target recognition.
[0109] In a possible implementation, the preset condition may be that the ratio of the area of the overlapped area between the high response area and the target foreground to the area of the high response area is less than a second threshold. Figure 7 For the lighter colored area shown, the second threshold is a proportional threshold value set based on actual experience. The area of the overlapping area and the area of the high response area can be determined using a related area function.
[0110] It should be noted that, when the ratio of the area of the overlapping area between the high response area and the target foreground to the area of the high response area is equal to the second threshold, it can be determined according to actual needs that there is no interference information in the image to be processed, or it can be determined that there is interference information in the image to be processed and the operation S203 is performed. The embodiment of the present application does not impose any specific restrictions on this.
[0111] In another possible implementation, the preset condition may be that the ratio of the area of the non-overlapping area with the target foreground in the high response area to the area of the high response area is greater than a third threshold, and the third threshold is a ratio threshold value set based on actual experience.
[0112] It should be noted that, for the case where the ratio of the area of the non-overlapping area with the target foreground in the high response area to the area of the high response area is equal to the third threshold, it can be determined according to actual needs that there is no interference information in the image to be processed, or it can be determined that there is interference information in the image to be processed and the operation S203 is performed. The embodiments of the present application do not impose specific restrictions on this.
[0113] It can be understood that the closer the ratio of the area of the overlapping area of the target foreground and the high response area to 1 (the closer the ratio of the area of the non-overlapping area with the target foreground to the area of the high response area in the high response area is to 0), the closer the high response area is to the target foreground, and there is basically no high response area outside the target foreground in the image to be processed, so there is basically no interference information that will affect the feature modeling during the modeling process. Conversely, the smaller the ratio of the area of the overlapping area to the high response area (the larger the ratio of the area of the non-overlapping area with the target foreground to the area of the high response area in the high response area), it means that there are more high response areas outside the target foreground. If the image to be processed is modeled directly, there may be factors such as background noise that will have a greater impact on the modeling process, thereby affecting the accuracy of target recognition.
[0114] Furthermore, if the target foreground and the high response area meet preset conditions (that is, when it is determined that interference information exists), the image processing device can adjust the pixel value of the target area outside the target foreground in the image to be processed, so that the influence of each pixel point in the target area on the feature modeling is less than the first threshold, so that the target area becomes the aforementioned low response area.
[0115] In a possible implementation, the image processing device may use a mean fuzzy method to adjust the pixel values of the target area to be consistent, so that the influence of each pixel point in the target area on the feature modeling is less than a first threshold. Specifically, the pixel values of each pixel point in the target area are summed and averaged, and the average value is used as the new pixel value of each pixel point to overwrite the original pixel value, so that the pixel values of the target area are consistent, thereby realizing the conversion of the target area into a low response area.
[0116] It can be understood that the consistency of pixel values is intuitively manifested as a consistent background tone, which can clearly distinguish the target object from other parts in the image. Therefore, in the target recognition process, adjusting the tone of the area outside the target object to be consistent is conducive to the model's recognition and extraction of the target object. Therefore, the embodiment of the present application adopts the mean fuzzy method to adjust the pixels of the target area to be consistent.
[0117] In another possible implementation, the image processing device may set the pixel values of the target area to preset values to adjust the pixel values of the target area to be consistent, so that the influence of each pixel point in the target area on the feature modeling is less than the first threshold value. For example, all pixel values in the target area are set to 0 or 255, so that the target area appears pure black or pure white, so as to avoid the influence of the target area on the feature modeling.
[0118] Among them, the target area includes the part of the high response area that does not belong to the target foreground.
[0119] In a possible implementation, the target region may be a portion of the high response region that does not belong to the target foreground, that is, the high response region outside the target foreground is taken as the target region for processing.
[0120] In another possible implementation, considering the error of the actual area in the determination process, the target area can also be an area outside the target foreground in the image to be processed, where the boundary of the high response area is expanded by a third threshold. Figure 8 Provide a schematic diagram of the target area. Figure 8 As shown in the figure, the black area is the target foreground. Outside the target foreground, the area selected by the dotted circle is the high response area. In addition, Figure 8 There is also an irregular area selected by a solid coil, which includes the high response area. The irregular area is the target area of the third threshold value extending outside the high response area.
[0121] The third threshold can be set according to actual experience, for example, it is set to 5 pixels. Then, for any pixel point in the high response area outside the target foreground, all pixels whose distance to the pixel point is less than or equal to 5 are considered as the target area, and the pixel values are adjusted to be consistent. For example, the following Table 1 shows the positions of all pixels whose distance to the pixel point (x, y) is less than or equal to 1.
[0122] Table 1
[0123] (x-1, y+1) (x, y+1) (x+1, y+1) (x-1, y) (x, y) (x+1, y) (x-1, y-1) (x, y-1) (x+1, y-1)
[0124] In another possible implementation, since all areas outside the target foreground are generally irrelevant background areas, the above-mentioned target area can also be all areas outside the target foreground in the image to be processed, and then all areas outside the target foreground are blurred by the mean or set to the same preset value.
[0125] Take the above target area as the part that does not belong to the target foreground in the high response area as an example. Fig. 9A schematic diagram of a target area is provided, showing target area 1 and target area 2 circled by white lines. It can be seen that the high response area outside the target foreground is circled to adjust its pixel value. Taking the pixel mean blur adjustment method as an example, Fig. 9 The target area is adjusted, and the final image is obtained after filtering out the high response area outside the target foreground, such as Fig.10 As shown, the color of the target area in the processed image has been adjusted to be consistent.
[0126] After determining that interference information exists in the image to be processed and executing S204 to process the image to be processed, the image processing apparatus may re-input the processed image into the target recognition model to perform new feature extraction and target recognition.
[0127] Fig.11 A schematic diagram of a process of image processing provided for an embodiment of the present application. First, an image is input (i.e., the aforementioned image to be processed), and for the image, the target foreground (specifically see S202 above) and the high response area (specifically see S201 above) are extracted. Then, it is determined whether there is interference (specifically see S203 above). If so, the interference information in the image is deleted (specifically see S204 above), and then new features are extracted through the target recognition model. If not, the original features of the image can be used.
[0128] The technical solution provided by the above-mentioned embodiment brings at least the following beneficial effects. The image processing method provided by the embodiment of the present application obtains the high response area and the target foreground in the image to be processed. When the target foreground and the high response area meet the preset conditions, it is determined that there is interference in the image to be processed, and then the pixel value of the target area outside the target foreground in the image to be processed is adjusted so that the influence of each pixel point in the target area on the feature modeling is less than the first threshold value, so as to remove the interference in the image to be processed. This method improves the situation where the target recognition model is easily interfered by interference information during modeling. After detecting the existence of interference information, the interference information can be automatically screened out to improve the accuracy of target recognition.
[0129] It can be seen that the above mainly introduces the solution provided by the embodiment of the present application from the perspective of the method. In order to achieve the above functions, the embodiment of the present application provides a hardware structure and / or software module corresponding to each function. It should be easily appreciated by those skilled in the art that, in combination with the modules and algorithm steps of each example described in the embodiment disclosed herein, the embodiment of the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.
[0130] In an exemplary embodiment, the present application further provides an image processing device, which may include one or more functional modules for implementing the image processing method of the above method embodiment.
[0131] For example, Fig.12 The following is a schematic diagram of the composition of an image processing device provided in an embodiment of the present application. Fig.12 As shown, the image processing device includes: an acquisition module 1201 and an adjustment module 1202. The acquisition module 1201 and the adjustment module 1202 are connected to each other.
[0132] The acquisition module 1201 is used to acquire a high response area of the image to be processed, where the influence of each pixel point in the high response area on feature modeling is greater than a first threshold.
[0133] The acquisition module 1201 is further used to acquire a target foreground of the image to be processed; the target foreground is used to indicate the area where the target object is located in the image to be processed.
[0134] The adjustment module 1202 is used to adjust the pixel value of the target area outside the target foreground in the image to be processed if the target foreground and the high response area meet the preset conditions, so that the influence of each pixel point in the target area on the feature modeling is less than a first threshold, and the target area includes the part of the high response area that does not belong to the target foreground.
[0135] In some embodiments, the acquisition module 1201 is specifically used to obtain a feature response visualization map of the image to be processed, where the feature response visualization map is used to reflect the influence of each pixel point in the image to be processed on the feature modeling; the feature response visualization map includes a response activation value of each pixel point; and the area of pixels whose response activation values are greater than a first threshold value is regarded as a high response area.
[0136] In some embodiments, the preset condition is that the ratio of the area of the overlapped area between the high response area and the target foreground to the area of the high response area is less than a second threshold.
[0137] In some embodiments, the adjustment module 1202 is specifically used to adjust the pixel values of the target area to be consistent by using a mean fuzzy method, so that the influence of each pixel point in the target area on the feature modeling is less than a first threshold.
[0138] In some embodiments, the target area is: the portion of the high response area that does not belong to the target foreground; or, the area outside the target foreground in the image to be processed, where the boundary of the high response area is expanded by a third threshold; or, all areas outside the target foreground in the image to be processed.
[0139] In some embodiments, the acquisition module 1201 is specifically used to model the image to be processed through deep learning to obtain multiple nodes; one node is the data of the modeling feature vector obtained by modeling the image to be processed in one dimension; the activation value response result of each node is determined; the activation value response result is used to indicate the contribution of each area of the image to be processed to the node; the weight of each node is determined; the weight of a node is determined by the activation value response result of the node and the activation value response results of all nodes; the activation value response result of each node and the weight of each node are processed by normalization function and upsampling function to obtain a feature response visualization graph.
[0140] In some embodiments, the above device further includes: a determination module 1203. The determination module 1202 is used to model the image to be processed by deep learning to obtain a target feature map; the target feature map is any one of the multiple feature maps obtained by modeling the image to be processed; based on the target feature map and the node, the Grad-CAM method is used to determine the activation value response result of each node.
[0141] In some embodiments, the acquisition module 1201 is specifically used to segment the image to be processed based on a semantic segmentation model based on deep learning or a traditional segmentation algorithm to obtain a target foreground.
[0142] In the case of implementing the functions of the above-mentioned integrated modules in the form of hardware, the embodiment of the present application provides a schematic diagram of the composition of an image processing device, which may be the above-mentioned image processing apparatus. Fig.13 As shown, the image processing device 1300 includes: a processor 1302 , a communication interface 1303 , and a bus 1304 . Optionally, the image processing device 1300 may further include a memory 1301 .
[0143] The processor 1302 may be a processor that implements or executes various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of the present application. The processor 1302 may be a central processing unit, a general-purpose processor, a digital signal processor, an application-specific integrated circuit, a field programmable gate array, or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of the present application. The processor 1302 may also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like.
[0144] The communication interface 1303 is used to connect with other devices via a communication network, such as Ethernet, wireless access network, wireless local area network (WLAN), etc.
[0145] The memory 1301 may be a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, or an electrically erasable programmable read-only memory (EEPROM), a disk storage medium or other magnetic storage device, or any other medium that can be used to carry or store desired program codes in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0146] As a possible implementation, the memory 1301 may exist independently of the processor 1302, and the memory 1301 may be connected to the processor 1302 via a bus 1304 for storing instructions or program codes. When the processor 1302 calls and executes the instructions or program codes stored in the memory 1301, the image processing method provided in the embodiment of the present application can be implemented.
[0147] In another possible implementation, the memory 1301 may also be integrated with the processor 1302 .
[0148] The bus 1304 may be an extended industry standard architecture (EISA) bus, etc. The bus 1304 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Fig.13 Only one thick line is used in the diagram, but this does not mean that there is only one bus or only one type of bus.
[0149] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the image processing device can be divided into different functional modules to complete all or part of the functions described above.
[0150] The embodiment of the present application also provides a computer-readable storage medium. All or part of the processes in the above method embodiments can be completed by computer instructions to instruct the relevant hardware, and the program can be stored in the above computer-readable storage medium. When the program is executed, it may include the processes of the above method embodiments. The computer-readable storage medium can be the memory or memory of any of the above embodiments. The above computer-readable storage medium can also be an external storage device of the above image processing device, such as a plug-in hard disk, a smart memory card (smart media card, SMC), a secure digital (secure digital, SD) card, a flash card (flash card), etc. equipped on the above XXXX device. Further, the above computer-readable storage medium can also include both the internal storage unit of the above image processing device and an external storage device. The above computer-readable storage medium is used to store the above computer program and other programs and data required by the above image processing device. The above computer-readable storage medium can also be used to temporarily store data that has been output or is to be output.
[0151] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program product is run on a computer, the computer is enabled to execute any one of the image processing methods provided in the above embodiments.
[0152] An embodiment of the present application also provides a chip, including: a processor and an interface, wherein the processor is coupled to a memory via the interface, and when the processor executes a computer program in the memory or an image processing device executes instructions, any one of the methods provided in the above embodiments is executed.
[0153] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other changes to the disclosed embodiments by viewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "one" or "an" does not exclude multiple situations. A single processor or other unit can implement several functions listed in the claims. Certain measures are recorded in different dependent claims, but this does not mean that these measures cannot be combined to produce good results.
[0154] Although the present application has been described in conjunction with specific features and embodiments thereof, it is obvious that various modifications and combinations may be made thereto without departing from the spirit and scope of the present application. Accordingly, this specification and the drawings are merely exemplary illustrations of the present application as defined by the appended claims, and are deemed to have covered any and all modifications, variations, combinations or equivalents within the scope of the present application. Obviously, those skilled in the art may make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
[0155] The above are only specific implementations of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. An image processing method, It is characterized in that The method comprises: Acquire a high response area of the image to be processed, wherein the influence of each pixel point in the high response area on the feature modeling is greater than a first threshold; the influence of the pixel point on the feature modeling is a response activation value in a feature response visualization diagram; Acquire a target foreground of the image to be processed; the target foreground is used to indicate an area where a target object is located in the image to be processed; If the target foreground and the high response area meet preset conditions, adjust the pixel value of the target area outside the target foreground in the image to be processed so that the influence of each pixel in the target area on the feature modeling is less than the first threshold, and the target area is: the part of the high response area that does not belong to the target foreground; or, in the image to be processed, outside the target foreground, the boundary of the high response area is expanded to an area of the third threshold; or, all areas outside the target foreground in the image to be processed.
2. The method according to claim 1, It is characterized in that The step of obtaining a high response area of the image to be processed includes: Obtaining a feature response visualization map of the image to be processed, wherein the feature response visualization map is used to reflect the influence of each pixel point in the image to be processed on feature modeling; the feature response visualization map includes a response activation value of each pixel point; The region of pixels whose response activation values are greater than the first threshold is taken as the high response region.
3. The method according to claim 1, It is characterized in that The preset condition is that the ratio of the area of the overlapping area between the high response area and the target foreground to the area of the high response area is less than a second threshold.
4. The method according to any one of claims 1 to 3, It is characterized in that The step of adjusting the pixel value of the target area outside the target foreground in the image to be processed so that the influence of each pixel point in the target area on the feature modeling is less than the first threshold value includes: The pixel values of the target area are adjusted to be consistent by using a mean fuzzy method, so that the influence of each pixel point in the target area on the feature modeling is less than the first threshold.
5. The method according to claim 2, It is characterized in that The step of obtaining a feature response visualization graph of the image to be processed includes: The image to be processed is modeled by deep learning to obtain a plurality of nodes; one node is data in one dimension of a modeling feature vector obtained by modeling the image to be processed; Determine an activation value response result of each of the nodes; the activation value response result is used to indicate the contribution of each area of the image to be processed to the node; Determine the weight of each of the nodes; the weight of a node is determined by the activation value response result of the node and the activation value response results of all nodes; The activation value response result of each of the nodes and the weight of each of the nodes are processed by a normalization function and an upsampling function to obtain the feature response visualization graph.
6. The method according to claim 5, It is characterized in that The determining of the activation value response result of each of the nodes includes: Modeling the image to be processed by deep learning to obtain a target feature map; the target feature map is any one of the multiple feature maps obtained by modeling the image to be processed; Based on the target feature map and the nodes, the class activation heat map Grad-CAM method is used to determine the activation value response result of each node.
7. The method according to any one of claims 1 to 3, claim 5, or claim 6, It is characterized in that The step of obtaining a target foreground of the image to be processed includes: The image to be processed is segmented based on a semantic segmentation model of deep learning or a traditional segmentation algorithm to obtain the target foreground.
8. An image processing device, It is characterized in that The device comprises: an acquisition module and an adjustment module; The acquisition module is used to acquire a high response area of the image to be processed, wherein the influence of each pixel point in the high response area on the feature modeling is greater than a first threshold; the influence of the pixel point on the feature modeling is a response activation value in a feature response visualization diagram; The acquisition module is further used to acquire a target foreground of the image to be processed; the target foreground is used to indicate an area where a target object is located in the image to be processed; The adjustment module is used to adjust the pixel value of the target area outside the target foreground in the image to be processed, if the target foreground and the high response area meet preset conditions, so that the influence of each pixel point in the target area on the feature modeling is less than the first threshold value, and the target area is: the part of the high response area that does not belong to the target foreground; or, in the image to be processed, outside the target foreground, the boundary of the high response area is expanded to an area of a third threshold value; or, all areas outside the target foreground in the image to be processed.
9. The device according to claim 8, It is characterized in that The device further comprises: a determination module; The acquisition module is specifically used to acquire a feature response visualization map of the image to be processed, wherein the feature response visualization map is used to reflect the influence of each pixel point in the image to be processed on the feature modeling; the feature response visualization map includes a response activation value of each pixel point; and the area of the pixel points whose response activation value is greater than the first threshold is taken as the high response area; The ratio of the area of the overlapping area between the high response area and the target foreground to the area of the high response area is less than a second threshold; The adjustment module is specifically used to adjust the pixel values of the target area to be consistent by using a mean fuzzy method, so that the influence of each pixel point in the target area on the feature modeling is less than the first threshold; The acquisition module is specifically used to model the image to be processed through deep learning to obtain multiple nodes; one node is the data of the modeling feature vector obtained by modeling the image to be processed in one dimension; determine the activation value response result of each node; the activation value response result is used to indicate the contribution of each area of the image to be processed to the node; determine the weight of each node; the weight of a node is determined by the activation value response result of the node and the activation value response results of all nodes; the activation value response result of each node and the weight of each node are processed by a normalization function and an upsampling function to obtain the feature response visualization graph; The determination module is used to model the image to be processed through deep learning to obtain a target feature map; the target feature map is any one of the multiple feature maps obtained by modeling the image to be processed; based on the target feature map and the nodes, the Grad-CAM method is used to determine the activation value response result of each node; The acquisition module is specifically used to segment the image to be processed based on a semantic segmentation model of deep learning or a traditional segmentation algorithm to obtain the target foreground.
10. An image processing device, It is characterized in that The image processing device comprises: a processor and a memory; The memory stores instructions executable by the processor; When the processor is configured to execute the instructions, the image processing device implements the method according to any one of claims 1 to 7.
11. A computer-readable storage medium, It is characterized in that The computer-readable storage medium includes: computer software instructions; When the computer software instructions are executed in a computer, the computer is enabled to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Multimedia resource detection method and device based on artificial intelligence, equipment and medium
CN111178343A
Image processing method and device thereof, computer readable medium and electronic equipment
CN113516735A