Image semantic segmentation method, device and electronic device
By constructing the edge loss function and the edge consistency loss function, and combining the original loss function to generate a composite loss function, the problem of insufficient accuracy of edge pixel segmentation in the prior art is solved, and the accuracy of semantic segmentation and edge processing are improved.
Patent Information
- Application Number
- CN202010852460.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-08-21
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2040-08-21
AI Technical Summary
The existing semantic segmentation algorithms have shortcomings in the accuracy of edge pixel segmentation.
By obtaining the semantic probability map and truth map of the image to be detected, the object edge detection map and the mask edge detection map are determined, the edge loss function and edge consistency loss function are constructed, and the composite loss function is generated in combination with the original loss function, which is used to train the image semantic segmentation model.
The segmentation effect of object edges and the accuracy of semantic segmentation are improved, making the processing of edges more accurate by semantic segmentation.
Smart Images

Figure CN114170251B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to an image semantic segmentation method, device and electronic equipment. Background Art
[0002] Image semantic segmentation refers to the process of segmenting an image based on its semantics. It segments different objects in the image from the perspective of pixels, defines the characteristics of different regions as the semantic category information of the object, and labels each pixel in the original image. This allows for segmentation results that are more valuable and easier for people to understand. Since deep learning has made great breakthroughs in image classification tasks, deep convolutional neural networks have also been applied to various computer vision tasks, such as object detection, saliency detection, and image super-resolution, and have achieved good results. The first end-to-end prediction of semantic segmentation tasks using neural networks was a fully convolutional network, which has now matured on some data sets, such as Pascal VOC. Semantic segmentation solves basic problems in computer vision, so its application scenarios are also very broad, such as autonomous driving, scene understanding, medical image processing, and remote sensing image analysis.
[0003] Existing pixel-level segmentation methods that use supervised learning for semantic segmentation generally use a global loss function to evaluate the quality of model training during model training. It lacks the ability to fine-tune local features and does not pay attention to the accuracy of edge pixel segmentation of objects.
[0004] In the process of implementing the present invention, the inventors found that the prior art has at least the following problems: the existing semantic segmentation algorithm has insufficient accuracy in edge pixel segmentation. Summary of the invention
[0005] An object of an embodiment of the present invention is to provide an image semantic segmentation method, device and electronic device, which can improve the segmentation effect of object edges and improve the accuracy of semantic segmentation.
[0006] In a first aspect, an embodiment of the present invention provides an image semantic segmentation method, which is applied to an electronic device, and the method includes:
[0007] Acquire an image to be detected, and determine a semantic probability map and a truth map corresponding to the image to be detected;
[0008] Determining an object edge detection map according to the semantic probability map;
[0009] Determining a mask edge detection map according to the true value map;
[0010] Determining an edge loss function according to the object edge detection map and the mask edge detection map;
[0011] Performing edge detection on the image to be detected to determine an original image edge detection map, and determining an edge consistency loss function according to the original image edge detection map and the mask edge detection map;
[0012] Determine a composite loss function according to an original loss function, an edge loss function and an edge consistency loss function, wherein the original loss function is used to measure the gap between the semantic probability map corresponding to the image to be detected and the true probability map;
[0013] Training is performed based on the composite loss function to obtain an image semantic segmentation model, and semantic segmentation is performed on the image to be detected based on the image semantic segmentation model.
[0014] In some embodiments, determining the object edge detection map according to the semantic probability map includes:
[0015] Performing horizontal filtering on the semantic probability map to determine a horizontal object edge detection map;
[0016] Performing vertical filtering on the semantic probability map to determine a vertical object edge detection map;
[0017] The horizontal object edge detection map and the vertical object edge detection map are combined along the channel direction to determine the object edge detection map.
[0018] In some embodiments, determining a mask edge detection map according to the true value map includes:
[0019] Performing horizontal filtering on the true value image to determine a horizontal mask edge detection image;
[0020] Performing vertical filtering on the true value image to determine a vertical mask edge detection image;
[0021] The horizontal mask edge detection map and the vertical mask edge detection map are combined along a channel direction to determine the mask edge detection map.
[0022] In some embodiments, determining the edge loss function according to the object edge detection map and the mask edge detection map includes:
[0023] The object edge detection image and the mask edge detection image are subjected to norm processing to construct an edge loss function.
[0024] In some embodiments, the edge loss function is:
[0025]
[0026] Among them, L Edge is the marginal loss function, is the value of the i-th position in the true probability map of the image to be detected, For The value after filtering, is the value of the i-th position in the semantic probability map of the image to be detected predicted by the neural network, For The value after filtering, p is a positive integer and p≥2.
[0027] In some embodiments, performing edge detection on the image to be detected to determine an edge detection map of the original image includes:
[0028] Performing horizontal filtering on the image to be detected to determine a horizontal original image edge detection image;
[0029] Performing vertical filtering on the image to be detected to determine a vertical original image edge detection image;
[0030] The horizontal original image edge detection map and the vertical original image edge detection map are merged along the channel direction to determine the original image edge detection map.
[0031] In some embodiments, determining the edge consistency loss function according to the original image edge detection map and the mask edge detection map includes:
[0032] The edge consistency loss function is determined according to the horizontal original image edge detection map and the vertical original image edge detection map.
[0033] In some embodiments, the edge consistency loss function is:
[0034]
[0035] Among them, L Mag is the edge consistency loss function, I x is the horizontal original image edge detection image, I y is the vertical original image edge detection image, M x is the horizontal mask edge detection map, M y It is the vertical mask edge detection map.
[0036] In some embodiments, determining the composite loss function according to the original loss function, the edge loss function and the edge consistency loss function includes:
[0037] Determine a first scaling factor and a second scaling factor corresponding to the edge loss function and the edge consistency loss function;
[0038] The composite loss function is determined according to the first scaling factor and the second scaling factor.
[0039] In some embodiments, the composite loss function is: L = L CE +λL Edge +βL Mag ,
[0040] Among them, L is the composite loss function, L CE is the original loss function, λ is the first scaling factor, L Edge is the edge loss function, β is the second scale factor, L Mag is the edge consistency loss function.
[0041] In some embodiments, the original loss function is a cross entropy loss function.
[0042] In some embodiments, the cross entropy loss function is:
[0043]
[0044] Among them, L CE is the cross entropy loss function, is the value of the i-th position in the true probability map of the image to be detected, It is the value of the i-th position in the semantic probability map of the image to be detected predicted by the neural network.
[0045] In a second aspect, an embodiment of the present invention provides an image semantic segmentation device, applied to an electronic device, the device comprising:
[0046] A semantic probability map unit, used to obtain an image to be detected and determine a semantic probability map corresponding to the image to be detected;
[0047] A truth map unit, used to determine a truth map corresponding to the image to be detected;
[0048] An object edge detection map unit, used to determine an object edge detection map according to the semantic probability map;
[0049] A mask edge detection map unit, used to determine a mask edge detection map according to the true value map;
[0050] An edge loss function unit, used to determine an edge loss function according to the object edge detection map and the mask edge detection map;
[0051] An original image edge detection map unit is used to perform edge detection on the image to be detected and determine an original image edge detection map;
[0052] An edge consistency loss function unit, used to determine an edge consistency loss function according to the original image edge detection map and the mask edge detection map;
[0053] A composite loss function unit, used to determine a composite loss function according to an original loss function, an edge loss function and an edge consistency loss function, wherein the original loss function is used to measure the gap between the semantic probability map corresponding to the image to be detected and the true probability map;
[0054] The semantic segmentation unit is used to perform training based on the composite loss function to obtain an image semantic segmentation model, and perform semantic segmentation on the image to be detected based on the image semantic segmentation model.
[0055] In some embodiments, the object edge detection image unit is specifically used to:
[0056] Performing horizontal filtering on the semantic probability map to determine a horizontal object edge detection map;
[0057] Performing vertical filtering on the semantic probability map to determine a vertical object edge detection map;
[0058] The horizontal object edge detection map and the vertical object edge detection map are combined along the channel direction to determine the object edge detection map.
[0059] In some embodiments, the mask edge detection image unit is specifically used to:
[0060] Performing horizontal filtering on the true value image to determine a horizontal mask edge detection image;
[0061] Performing vertical filtering on the semantic probability map to determine a vertical mask edge detection map;
[0062] The horizontal mask edge detection map and the vertical mask edge detection map are combined along a channel direction to determine the mask edge detection map.
[0063] In some embodiments, the edge loss function unit is specifically used to:
[0064] The object edge detection image and the mask edge detection image are subjected to norm processing to construct an edge loss function.
[0065] In some embodiments, the edge loss function is:
[0066]
[0067] Among them, L Edge is the marginal loss function, is the value of the i-th position in the true probability map of the image to be detected, For The value after filtering, is the value of the i-th position in the semantic probability map of the image to be detected predicted by the neural network, For The value after filtering, p is a positive integer and p≥2.
[0068] In some embodiments, the original image edge detection unit is specifically used to:
[0069] Performing horizontal filtering on the image to be detected to determine a horizontal original image edge detection image;
[0070] Performing vertical filtering on the image to be detected to determine a vertical original image edge detection image;
[0071] The horizontal original image edge detection map and the vertical original image edge detection map are merged along the channel direction to determine the original image edge detection map.
[0072] In some embodiments, the edge consistency loss function unit is specifically used to:
[0073] The edge consistency loss function is determined according to the horizontal original image edge detection map and the vertical original image edge detection map.
[0074] In some embodiments, the edge consistency loss function is:
[0075]
[0076] Among them, L Mag is the edge consistency loss function, I x is the horizontal original image edge detection image, I y is the vertical original image edge detection image, M x is the horizontal mask edge detection map, M y It is the vertical mask edge detection map.
[0077] In some embodiments, the composite loss function unit is specifically used to:
[0078] Determine a first scaling factor and a second scaling factor corresponding to the edge loss function and the edge consistency loss function;
[0079] The composite loss function is determined according to the first scaling factor and the second scaling factor.
[0080] In some embodiments, the composite loss function is: L = L CE +λL Edge +βL Mag ,
[0081] Among them, L is the composite loss function, λ is the first scale factor, and L Edge is the edge loss function, β is the second scale factor, L Mag is the edge consistency loss function.
[0082] In some embodiments, the original loss function is a cross entropy loss function.
[0083] In some embodiments, the cross entropy loss function is:
[0084]
[0085] Among them, L CE is the cross entropy loss function, is the value of the i-th position in the true probability map of the image to be detected, It is the value of the i-th position in the semantic probability map of the image to be detected predicted by the neural network.
[0086] In a third aspect, an embodiment of the present invention provides an electronic device, including:
[0087] at least one processor; and
[0088] a memory communicatively connected to the at least one processor; wherein,
[0089] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can perform the above-mentioned image semantic segmentation method.
[0090] In a fourth aspect, an embodiment of the present invention provides a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable an electronic device to execute the above-mentioned image semantic segmentation method.
[0091] In a fifth aspect, an embodiment of the present invention provides a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by one or more processors in an electronic device, the electronic device executes the above-mentioned image semantic segmentation method.
[0092] The beneficial effect of the embodiment of the present invention is that, different from the prior art, the embodiment of the present invention provides an image semantic segmentation method, which is applied to electronic equipment, and the method includes: acquiring an image to be detected, determining a semantic probability map and a true value map corresponding to the image to be detected; determining an object edge detection map according to the semantic probability map; determining a mask edge detection map according to the true value map; determining an edge loss function according to the object edge detection map and the mask edge detection map; performing edge detection on the image to be detected, determining an original image edge detection map, and determining an edge consistency loss function according to the original image edge detection map and the mask edge detection map; determining a composite loss function according to the original loss function, the edge loss function and the edge consistency loss function, wherein the original loss function is used to measure the gap between the semantic probability map corresponding to the image to be detected and the true probability map; training based on the composite loss function to obtain an image semantic segmentation model, and performing semantic segmentation on the image to be detected based on the image semantic segmentation model. On the one hand, by adding the edge loss function and the edge consistency loss function to generate a composite loss function, the semantic segmentation processing of the edge is made more accurate. On the other hand, by training the composite loss function to obtain an image semantic segmentation model, the present invention can improve the accuracy of semantic segmentation. BRIEF DESCRIPTION OF THE DRAWINGS
[0093] One or more embodiments are exemplarily described by pictures in the corresponding drawings, and these exemplified descriptions do not constitute limitations on the embodiments. Elements with the same reference numerals in the drawings represent similar elements, and unless otherwise stated, the figures in the drawings do not constitute proportional limitations.
[0094] Figure 1 It is a flowchart of an image semantic segmentation method provided by an embodiment of the present invention;
[0095] Figure 2 yes Figure 1 A detailed flowchart of step S20 in FIG.
[0096] Figure 3 yes Figure 1 A detailed flowchart of step S30 in FIG.
[0097] Figure 4 yes Figure 1 A detailed flowchart of step S50 in FIG.
[0098] Figure 5 yes Figure 1 A detailed flowchart of step S60 in FIG.
[0099] Figure 6 yes Figure 5 A detailed flowchart of step S61 in FIG.
[0100] Figure 7 is a structural schematic diagram of an image semantic segmentation device provided by an embodiment of the present invention;
[0101] Figure 8 It is a structural schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0102] In order to make the purpose, technical scheme and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0103] It should be noted that, if there is no conflict, the various features in the embodiments of the present invention can be combined with each other, all within the scope of protection of the present invention. In addition, although the functional modules are divided in the device schematic diagram and the logical order is shown in the flow chart, in some cases, the steps shown or described can be performed in a different order from the module division in the device or the flow chart. Furthermore, the words "first", "second", "third", etc. used in the present invention do not limit the data and execution order, but only distinguish the same items or similar items with basically the same functions and effects.
[0104] Before describing the present invention in detail, the nouns and terms involved in the embodiments of the present invention are described. The nouns and terms involved in the embodiments of the present invention are subject to the following explanations.
[0105] (1) Semantic segmentation: refers to the computer segmentation based on the semantics of the image. Different objects in the image are segmented from the perspective of pixels. The characteristics of different areas are defined as the semantic category information of the object. Each pixel in the original image is labeled. The image is classified at the pixel level, and pixels belonging to the same category are classified into one category.
[0106] (2) Loss function: refers to the function used to measure the degree of inconsistency between the model's predicted value and the true value.
[0107] (3) Semantic probability map: refers to the probability map of pixels at corresponding positions in the image.
[0108] (4) True value image: refers to the binary image.
[0109] Semantic segmentation defines the characteristics of different regions as the semantic category information of objects. It is usually implemented using a fully convolutional neural network. For example, the portrait segmentation task is implemented using a neural network. It can be applied to mobile terminals, such as photo filters, background replacement and other photo effects, and can also be applied to non-mobile terminals, such as monitoring analysis, video, and image editing. Portrait segmentation can be regarded as a binary semantic segmentation task, so it is very relevant to saliency detection or cutout tasks. The cutout task assumes that an image can be decomposed into foreground and background, which are mixed with a certain probability, that is, I = (1-α)B + αF, where I represents the original image, B represents the background image, F represents the foreground image, and α represents the foreground mask probability map. Each pixel in the image is superimposed with the foreground and background with a probability in α. The cutout task focuses more on the boundary part of the object. For portrait segmentation, the boundary of the person may not be separated from the background for various reasons. Therefore, the deep cutout method can also be used to optimize the segmentation edge.
[0110] However, existing pixel-level segmentation methods that use supervised learning for semantic segmentation generally use a global loss function to evaluate the quality of model training during model training. Generally speaking, the global loss function lacks the ability to fine-tune local features and does not pay attention to the accuracy of edge pixel segmentation of objects, resulting in insufficient accuracy of edge pixel segmentation.
[0111] Based on this, an embodiment of the present invention provides an image semantic segmentation method to improve the accuracy of semantic segmentation.
[0112] See also Figure 1 , Figure 1 It is a flow chart of an image semantic segmentation method provided by an embodiment of the present invention; wherein, the image semantic segmentation method is applied to an electronic device, specifically, to at least one processor of the electronic device, that is, the executor of the image semantic segmentation method is at least one processor of the electronic device.
[0113] like Figure 1 As shown, the image semantic segmentation method includes:
[0114] Step S10: Acquire an image to be detected, and determine a semantic probability map and a true value map corresponding to the image to be detected;
[0115] Specifically, the image to be detected is an RGB image, and the determination of the semantic probability map corresponding to the image to be detected includes: when the image to be detected I is input, the image to be detected is input into the original image semantic segmentation model, and the semantic probability map predicted by the original image semantic segmentation model is determined, wherein the semantic probability map is the semantic probability map of the image to be detected predicted by the neural network, and the original image semantic segmentation model includes: Densenet model, resnet model, FCN model, DeepLabv1 model, DeepLabv2 model, DeepLabv3 model, DeepLabv3+ model, preferably, the present invention uses the DeepLabv3+ model with MobileNetv2 as the backbone network to predict the probability map of semantic segmentation of the image to be detected, so as to determine the semantic probability map corresponding to the image to be detected, and the semantic probability map includes the probability of the foreground object pixels in the image to be detected at the corresponding position of the image to be detected. It can be understood that the semantic probability map of the image to be detected is different from the real probability map of the image to be detected.
[0116] Specifically, determining the true value map corresponding to the image to be detected includes: performing binarization processing on the image to be detected to obtain the true value map of the image to be detected.
[0117] Step S20: determining an object edge detection map according to the semantic probability map;
[0118] Specifically, the semantic probability map is filtered by a filter, wherein the filter is a filter with an edge detection function, and the filter includes a sobel filter, a Roberts filter, a Prewitt filter, a Laplacian filter, a Marr filter, a Canny filter, a Kirsch filter, and a Nevitia filter. Preferably, the filter in the embodiment of the present invention is a sobel filter.
[0119] For details, please refer to Figure 2 , Figure 2 yes Figure 1 A detailed flowchart of step S20 in FIG.
[0120] like Figure 2 As shown, step S20: determining an object edge detection map according to the semantic probability map includes:
[0121] Step S21: performing horizontal filtering on the semantic probability map to determine a horizontal object edge detection map;
[0122] Specifically, the semantic probability map is filtered in the horizontal direction by a horizontal sobel filter, so as to determine a horizontal object edge detection map.
[0123] Step S22: vertically filtering the semantic probability map to determine a vertical object edge detection map;
[0124] Specifically, the semantic probability map is filtered in the vertical direction by a vertical sobel filter, so as to determine the vertical object edge detection map.
[0125] Step S23: merging the horizontal object edge detection map and the vertical object edge detection map along the channel direction to determine the object edge detection map.
[0126] Among them, the tensor of the image after being processed by the neural network of the original image semantic segmentation model is 4-dimensional, namely [N, H, W, C], wherein N, H, W, C represent: the number of samples, the tensor height, the tensor width and the number of tensor channels, that is to say, the horizontal object edge detection map and the vertical object edge detection map are both four-dimensional tensor images. Assuming that the horizontal object edge detection map and the vertical object edge detection map are both 1024*512 RGB images, the horizontal object edge detection map and the vertical object edge detection map are both images with a tensor of [1,1024,512,3]. The object edge detection map generated after the horizontal object edge detection map and the vertical object edge detection map are merged along the channel direction is an image with a tensor of [1,1024,512,6]. It can be understood that the object edge detection map is the total object edge detection map generated after the horizontal object edge detection map and the vertical object edge detection map are merged.
[0127] In the embodiment of the present invention, by using a horizontal Sobel filter and a vertical Sobel filter to filter the semantic probability map in the horizontal direction and the vertical direction respectively, more attention is paid to the edge during the network training process, so that the semantic segmentation processes the edge more accurately.
[0128] Step S30: determining a mask edge detection map according to the true value map;
[0129] Specifically, the truth image is filtered by a filter, wherein the filter is a filter with an edge detection function, and the filter includes a sobel filter, a Roberts filter, a Prewitt filter, a Laplacian filter, a Marr filter, a Canny filter, a Kirsch filter, and a Nevitia filter. Preferably, the filter in the embodiment of the present invention is a sobel filter.
[0130] Please refer to Figure 3 , Figure 3 yes Figure 1 A detailed flowchart of step S30 in FIG.
[0131] like Figure 3 As shown, step S30: determining a mask edge detection map according to the true value map includes:
[0132] Step S31: performing horizontal filtering on the true value image to determine a horizontal mask edge detection image;
[0133] Specifically, the truth image is filtered in the horizontal direction by a horizontal sobel filter, so as to determine the horizontal mask edge detection image.
[0134] Step S32: vertically filtering the semantic probability map to determine a vertical mask edge detection map;
[0135] Specifically, the semantic probability map is filtered in the vertical direction by a vertical sobel filter, so as to determine the vertical object edge detection map.
[0136] Step S33: merging the horizontal mask edge detection map and the vertical mask edge detection map along the channel direction to determine the mask edge detection map.
[0137] Among them, the tensor of the image after being processed by the neural network of the original image semantic segmentation model is 4-dimensional, namely [N, H, W, C], wherein N, H, W, C represent: the number of samples, the tensor height, the tensor width and the number of tensor channels, that is to say, the horizontal mask edge detection map and the vertical mask edge detection map are both four-dimensional tensor images. Assuming that the horizontal mask edge detection map and the vertical mask edge detection map are both 1024*512 RGB images, the horizontal mask edge detection map and the vertical mask edge detection map are both images with a tensor of [1,1024,512,3]. Then, the mask edge detection map generated after the horizontal mask edge detection map and the vertical mask edge detection map are merged along the channel direction is an image with a tensor of [1,1024,512,6]. It can be understood that the mask edge detection map is the total mask edge detection map generated after the horizontal mask edge detection map and the vertical mask edge detection map are merged.
[0138] In the embodiment of the present invention, by using a horizontal Sobel filter and a vertical Sobel filter to filter the truth image in the horizontal direction and the vertical direction respectively, more attention is paid to the edges during the network training process, so that the semantic segmentation processes the edges more accurately.
[0139] Step S40: determining an edge loss function according to the object edge detection image and the mask edge detection image;
[0140] Specifically, determining the edge loss function according to the object edge detection map and the mask edge detection map includes:
[0141] The object edge detection image and the mask edge detection image are subjected to norm processing to construct an edge loss function.
[0142] The norm processing includes calculating by using a p-order norm, where p is a positive integer, for example, p is 2, that is, the object edge detection map and the mask edge detection map are processed using the L2 norm to obtain the result of the edge loss function.
[0143] Specifically, the edge loss function is:
[0144]
[0145] Among them, L Edge is the marginal loss function, is the value of the i-th position in the true probability map of the image to be detected, For The value after filtering, is the value of the i-th position in the semantic probability map of the image to be detected predicted by the neural network, For The value after filtering, p is a positive integer and p≥2.
[0146] Specifically, the edge loss function L Edge , f() is a convolution operation used to extract edge information, and in the embodiment of the present invention, refers to the result after passing through the filter.
[0147] Among them, L Edge The principle of the function is that the Sobel filter is more sensitive to areas with sudden changes in values in the image, so it is used to detect the edges of objects in the image. and After being processed with a filter, the foreground background edge of the predicted image and the foreground background edge of the real image can be obtained, and then the p norm is used to measure whether the foreground background edge of the predicted image and the foreground background edge of the real image are consistent. Wherein, the value of p refers to the pth power, and the value of p in the embodiment of the present invention is set according to actual needs, for example: p takes a value of 2, which represents a common mean square error. It is understandable that the penalty degree of the difference under different numerical values corresponding to different values of p is also different.
[0148] Step S50: performing edge detection on the image to be detected, determining an original image edge detection map, and determining an edge consistency loss function according to the original image edge detection map and the mask edge detection map;
[0149] For details, please refer to Figure 4 , Figure 4 yes Figure 1 A detailed flowchart of step S50 in FIG.
[0150] like Figure 4 As shown, step S50: performing edge detection on the image to be detected to determine an edge detection map of the original image includes:
[0151] Step S51: Filter the image to be detected in the horizontal direction to determine a horizontal original image edge detection image;
[0152] Specifically, the image to be detected is an unprocessed original image, and the image to be detected is filtered in the horizontal direction by a horizontal Sobel filter, so as to determine a horizontal original image edge detection image.
[0153] Step S52: vertically filter the image to be detected to determine a vertical original image edge detection image;
[0154] Specifically, the image to be detected is filtered in the vertical direction by using a vertical sobel filter, so as to determine a vertical original image edge detection image.
[0155] Step S53: merging the horizontal original image edge detection map and the vertical original image edge detection map along the channel direction to determine the original image edge detection map.
[0156] Specifically, the horizontal original image edge detection image and the vertical original image edge detection image are both four-dimensional tensor images. Assuming that the horizontal original image edge detection image and the vertical original image edge detection image are both 1024*512 RGB images, the horizontal original image edge detection image and the vertical original image edge detection image are both images with a tensor of [1,1024,512,3]. The original image edge detection image generated after the horizontal original image edge detection image and the vertical original image edge detection image are merged along the channel direction is an image with a tensor of [1,1024,512,6]. It can be understood that the original image edge detection image is the total original image edge detection image generated after the horizontal original image edge detection image and the vertical original image edge detection image are merged.
[0157] Specifically, determining the edge consistency loss function according to the original image edge detection map and the mask edge detection map includes:
[0158] The edge consistency loss function is determined according to the horizontal original image edge detection map and the vertical original image edge detection map.
[0159] Wherein, the edge consistency loss function is:
[0160]
[0161] Among them, L Mag is the edge consistency loss function, I x is the horizontal original image edge detection image, I y is the vertical original image edge detection image, M x is the horizontal mask edge detection map, M y It is the vertical mask edge detection map.
[0162] Specifically, the edge consistency loss function L Mag The principle of processing data is: for an input image to be detected, which is an RGB original image, the position corresponding to the image to be detected and the probability map predicted by the network is strongly correlated. For example, if the edge of the character in the original image is inconsistent with the edge position of the character in the predicted probability map, it means that the error of the loss function will be large. Due to the annotation error in the true value image, the edge may not be consistent with the original image. Therefore, the edge consistency loss function is needed to constrain whether the edges of the annotated true value image and the original image are consistent. If they are consistent, the penalty for this item is 0. Otherwise, the greater the deviation, the greater the penalty.
[0163] Understandably, I x and I y It refers to the edge map of the image to be detected (RGB original image) after it passes through the filter (edge detection operator). The edge map has a larger value only in the area where the edge is detected, and the value in other areas is 0.
[0164] In an embodiment of the present invention, the input original image (image to be detected) is filtered using a horizontal Sobel filter and a vertical Sobel filter respectively, and the horizontal original image edge detection map and the vertical original image edge detection map of the image to be detected are extracted. The horizontal mask edge detection map and the vertical mask edge detection map are further used to construct an edge consistency loss function. The embodiment of the present invention can pay more attention to edge processing in the semantic segmentation process and improve the accuracy of edge processing.
[0165] Step S60: determining a composite loss function according to an original loss function, an edge loss function and an edge consistency loss function, wherein the original loss function is used to measure the gap between the semantic probability map corresponding to the image to be detected and the true probability map;
[0166] For details, please refer to Figure 5 , Figure 5 yes Figure 1 A detailed flowchart of step S60 in FIG.
[0167] like Figure 5As shown, step S60: determining a composite loss function according to the original loss function, the edge loss function and the edge consistency loss function, wherein the original loss function is used to measure the gap between the semantic probability map corresponding to the image to be detected and the true probability map, including:
[0168] Step S61: determining a first scaling factor and a second scaling factor corresponding to the edge loss function and the edge consistency loss function;
[0169] Among them, the edge loss function corresponds to a first proportional factor, the edge consistency loss function corresponds to a second proportional factor, and the first proportional factor and the second proportional factor are respectively used to adjust the composition ratio of the edge loss function and the edge consistency loss function.
[0170] For details, please refer to Figure 6 , Figure 6 yes Figure 5 A detailed flowchart of step S61 in FIG.
[0171] like Figure 6 As shown, the step S61: determining the first scaling factor and the second scaling factor corresponding to the edge loss function and the edge consistency loss function includes:
[0172] Step S611: presetting a first value range corresponding to the first scale factor and a second value range corresponding to the second scale factor;
[0173] The first value range and the second value range are both (0, 1), the first value range corresponds to a number of values, and the second value range corresponds to the same number of values as the first value range, for example: the first value range and the second value range both correspond to ten values;
[0174] Step S612: constructing a plurality of scaling factor combinations based on the value of the first scaling factor and the value of the second scaling factor;
[0175] Specifically, all values of the first scale factor and all values of the second scale factor are traversed to construct multiple scale factor combinations. For example, if the value of the first scale factor is 10 and the value of the second scale factor is also 10, then the number of scale factor combinations is 10*10=100.
[0176] Step S613: Calculate the evaluation index corresponding to each combination of proportional factors according to the plurality of combinations of proportional factors;
[0177] Specifically, the evaluation indicators include mean intersection over union (MIoU) and / or accuracy, wherein the mean intersection over union (MIoU) is a standard metric for semantic segmentation, which represents the average value of the ratio of the intersection and union of all categories, and the accuracy is the accuracy of network prediction.
[0178] Step S614: Determine the optimal combination of proportional factors according to the evaluation index.
[0179] Specifically, the proportional factor combination with the highest evaluation index is determined, and the proportional factor combination with the highest evaluation index is determined as the optimal proportional factor combination. It can be understood that the intersection-and-union ratio is positively correlated with the accuracy rate, that is, the higher the intersection-and-union ratio, the higher the accuracy rate.
[0180] Step S62: Determine the composite loss function according to the first scaling factor and the second scaling factor.
[0181] Specifically, the composite loss function is: L = L CE +λL Edge +βL Mag ,
[0182] Among them, L is the composite loss function, L CE is the original loss function, λ is the first scaling factor, L Edge is the edge loss function, β is the second scale factor, L Mag is the edge consistency loss function.
[0183] Specifically, the original loss function L CE is the cross entropy loss function, which is:
[0184]
[0185] Among them, L CE is the cross entropy loss function, is the value of the i-th position in the true probability map of the image to be detected, It is the value of the i-th position in the semantic probability map of the image to be detected predicted by the neural network.
[0186] Specifically, the original loss function L CE It is used to measure the gap between the predicted probability map and the true probability map. The larger the gap, the larger the loss function value. Here, the predicted probability map is the size and original Figure 1 The probability value ranges from 0 to 1. The closer it is to 1, the greater the probability that the pixel at that position is classified as foreground, and vice versa. It refers to the value of the i-th position in the true probability map. It can be understood that the value of the i-th position in the true probability map can only be 0 or 1, because each pixel in the original image can be determined to be the foreground or the background. It refers to the value of the i-th position in the semantic probability map predicted by the neural network. The value of the i-th position in the semantic probability map is a floating point number between 0 and 1. Since it is a predicted probability, it may be a floating point number.
[0187] Step S70: Performing training based on the composite loss function to obtain an image semantic segmentation model, and performing semantic segmentation on the image to be detected based on the image semantic segmentation model.
[0188] Specifically, the embodiment of the present invention uses the MobileNetV2 network as the backbone network as the encoder, uses transposed convolution and hole convolution networks to construct a decoder network, and trains through the composite loss function to obtain an image semantic segmentation model, and based on the image semantic segmentation model, performs semantic segmentation on the image to be detected to generate a semantic segmentation map after semantic segmentation.
[0189] Unlike other segmentation algorithms optimized for image edges that require the introduction of additional computing modules, the embodiments of the present invention do not require the introduction of additional modules during the use of the model forward transmission, thereby reducing the amount of calculation and can be applied to lightweight and mobile terminals for model deployment to enrich the application scenarios of the embodiments of the present invention.
[0190] In an embodiment of the present invention, a method for image semantic segmentation is provided, which is applied to an electronic device. The method includes: obtaining an image to be detected, determining a semantic probability map and a true value map corresponding to the image to be detected; determining an object edge detection map according to the semantic probability map; determining a mask edge detection map according to the true value map; determining an edge loss function according to the object edge detection map and the mask edge detection map; performing edge detection on the image to be detected, determining an original image edge detection map, and determining an edge consistency loss function according to the original image edge detection map and the mask edge detection map; determining a composite loss function according to the original loss function, the edge loss function and the edge consistency loss function, wherein the original loss function is used to measure the difference between the semantic probability map corresponding to the image to be detected and the true probability map; training based on the composite loss function to obtain an image semantic segmentation model, and semantically segmenting the image to be detected based on the image semantic segmentation model. On the one hand, by adding the edge loss function and the edge consistency loss function to generate a composite loss function, the semantic segmentation can process the edge more accurately. On the other hand, by training the composite loss function to obtain the image semantic segmentation model, the present invention can improve the accuracy of semantic segmentation.
[0191] See also Figure 7 , Figure 7 It is a structural schematic diagram of an image semantic segmentation device provided by an embodiment of the present invention; wherein the image semantic segmentation device 700 is applied to an electronic device, specifically, to one or more processors of the electronic device.
[0192] like Figure 7 As shown, the image semantic segmentation device 700 includes:
[0193] The semantic probability map unit 710 is used to obtain an image to be detected and determine a semantic probability map corresponding to the image to be detected;
[0194] A truth map unit 720, used to determine a truth map corresponding to the image to be detected;
[0195] An object edge detection map unit 730, configured to determine an object edge detection map according to the semantic probability map;
[0196] A mask edge detection map unit 740, configured to determine a mask edge detection map according to the true value map;
[0197] An edge loss function unit 750, configured to determine an edge loss function according to the object edge detection map and the mask edge detection map;
[0198] The original image edge detection map unit 760 is used to perform edge detection on the image to be detected and determine the original image edge detection map;
[0199] An edge consistency loss function unit 770, configured to determine an edge consistency loss function according to the original image edge detection map and the mask edge detection map;
[0200] A composite loss function unit 780 is used to determine a composite loss function according to an original loss function, an edge loss function and an edge consistency loss function, wherein the original loss function is used to measure the difference between the semantic probability map corresponding to the image to be detected and the true probability map;
[0201] The semantic segmentation unit 790 is used to perform training based on the composite loss function to obtain an image semantic segmentation model, and perform semantic segmentation on the image to be detected based on the image semantic segmentation model.
[0202] In the embodiment of the present invention, the object edge detection image unit is specifically used for:
[0203] Performing horizontal filtering on the semantic probability map to determine a horizontal object edge detection map;
[0204] Specifically, the semantic probability map is filtered in the horizontal direction by a horizontal sobel filter, so as to determine a horizontal object edge detection map.
[0205] Performing vertical filtering on the semantic probability map to determine a vertical object edge detection map;
[0206] Specifically, the semantic probability map is filtered in the vertical direction by a vertical sobel filter, so as to determine the vertical object edge detection map.
[0207] The horizontal object edge detection map and the vertical object edge detection map are combined along the channel direction to determine the object edge detection map.
[0208] Among them, the tensor of the image after being processed by the neural network of the original image semantic segmentation model is 4-dimensional, namely [N, H, W, C], wherein N, H, W, C represent: the number of samples, the tensor height, the tensor width and the number of tensor channels, that is to say, the horizontal object edge detection map and the vertical object edge detection map are both four-dimensional tensor images. Assuming that the horizontal object edge detection map and the vertical object edge detection map are both 1024*512 RGB images, the horizontal object edge detection map and the vertical object edge detection map are both images with a tensor of [1,1024,512,3]. The object edge detection map generated after the horizontal object edge detection map and the vertical object edge detection map are merged along the channel direction is an image with a tensor of [1,1024,512,6]. It can be understood that the object edge detection map is the total object edge detection map generated after the horizontal object edge detection map and the vertical object edge detection map are merged.
[0209] In the embodiment of the present invention, by using a horizontal Sobel filter and a vertical Sobel filter to filter the semantic probability map in the horizontal direction and the vertical direction respectively, more attention is paid to the edge during the network training process, so that the semantic segmentation processes the edge more accurately.
[0210] In the embodiment of the present invention, the mask edge detection image unit is specifically used for:
[0211] Performing horizontal filtering on the true value image to determine a horizontal mask edge detection image;
[0212] Specifically, the truth image is filtered in the horizontal direction by a horizontal sobel filter, so as to determine the horizontal mask edge detection image.
[0213] Performing vertical filtering on the semantic probability map to determine a vertical mask edge detection map;
[0214] Specifically, the semantic probability map is filtered in the vertical direction by a vertical sobel filter, so as to determine the vertical object edge detection map.
[0215] The horizontal mask edge detection map and the vertical mask edge detection map are combined along a channel direction to determine the mask edge detection map.
[0216] Among them, the tensor of the image after being processed by the neural network of the original image semantic segmentation model is 4-dimensional, namely [N, H, W, C], wherein N, H, W, C represent: the number of samples, the tensor height, the tensor width and the number of tensor channels, that is to say, the horizontal mask edge detection map and the vertical mask edge detection map are both four-dimensional tensor images. Assuming that the horizontal mask edge detection map and the vertical mask edge detection map are both 1024*512 RGB images, the horizontal mask edge detection map and the vertical mask edge detection map are both images with a tensor of [1,1024,512,3]. Then, the mask edge detection map generated after the horizontal mask edge detection map and the vertical mask edge detection map are merged along the channel direction is an image with a tensor of [1,1024,512,6]. It can be understood that the mask edge detection map is the total mask edge detection map generated after the horizontal mask edge detection map and the vertical mask edge detection map are merged.
[0217] In the embodiment of the present invention, by using a horizontal Sobel filter and a vertical Sobel filter to filter the truth image in the horizontal direction and the vertical direction respectively, more attention is paid to the edges during the network training process, so that the semantic segmentation processes the edges more accurately.
[0218] In the embodiment of the present invention, the edge loss function unit is specifically used for:
[0219] The object edge detection image and the mask edge detection image are subjected to norm processing to construct an edge loss function.
[0220] In this embodiment of the present invention, the edge loss function is:
[0221]
[0222] Among them, L Edge is the marginal loss function, is the value of the i-th position in the true probability map of the image to be detected, For The value after filtering, is the value of the i-th position in the semantic probability map of the image to be detected predicted by the neural network, For The value after filtering, p is a positive integer and p≥2.
[0223] Specifically, the edge loss function L Edge , f() is a convolution operation used to extract edge information, and in the embodiment of the present invention, refers to the result after passing through the filter.
[0224] Among them, L Edge The principle of the function is that the Sobel filter is more sensitive to areas with sudden changes in values in the image, so it is used to detect the edges of objects in the image. and After being processed with a filter, the foreground background edge of the predicted image and the foreground background edge of the real image can be obtained, and then the p norm is used to measure whether the foreground background edge of the predicted image and the foreground background edge of the real image are consistent. Wherein, the value of p refers to the pth power, and the value of p in the embodiment of the present invention is set according to actual needs, for example: p takes a value of 2, which represents a common mean square error. It is understandable that the penalty degree of the difference under different numerical values corresponding to different values of p is also different.
[0225] In the embodiment of the present invention, the original image edge detection unit is specifically used for:
[0226] Performing horizontal filtering on the image to be detected to determine a horizontal original image edge detection image;
[0227] Performing vertical filtering on the image to be detected to determine a vertical original image edge detection image;
[0228] The horizontal original image edge detection map and the vertical original image edge detection map are merged along the channel direction to determine the original image edge detection map.
[0229] In the embodiment of the present invention, the edge consistency loss function unit is specifically used for:
[0230] The edge consistency loss function is determined according to the horizontal original image edge detection map and the vertical original image edge detection map.
[0231] In this embodiment of the present invention, the edge consistency loss function is:
[0232]
[0233] Among them, L Mag is the edge consistency loss function, I x is the horizontal original image edge detection image, I y is the vertical original image edge detection image, M x is the horizontal mask edge detection map, M y It is the vertical mask edge detection map.
[0234] Specifically, the edge consistency loss function LMag The principle of processing data is: for an input image to be detected, which is an RGB original image, the position corresponding to the image to be detected and the probability map predicted by the network is strongly correlated. For example, if the edge of the character in the original image is inconsistent with the edge position of the character in the predicted probability map, it means that the error of the loss function will be large. Due to the annotation error in the true value image, the edge may not be consistent with the original image. Therefore, the edge consistency loss function is needed to constrain whether the edges of the annotated true value image and the original image are consistent. If they are consistent, the penalty for this item is 0. Otherwise, the greater the deviation, the greater the penalty.
[0235] Understandably, I x and I y It refers to the edge map of the image to be detected (RGB original image) after it passes through the filter (edge detection operator). The edge map has a larger value only in the area where the edge is detected, and the value in other areas is 0.
[0236] In an embodiment of the present invention, the input original image (image to be detected) is filtered using a horizontal Sobel filter and a vertical Sobel filter respectively, and the horizontal original image edge detection map and the vertical original image edge detection map of the image to be detected are extracted. The horizontal mask edge detection map and the vertical mask edge detection map are further used to construct an edge consistency loss function. The embodiment of the present invention can pay more attention to edge processing in the semantic segmentation process and improve the accuracy of edge processing.
[0237] In this embodiment of the present invention, the composite loss function unit is specifically used for:
[0238] Determine a first scaling factor and a second scaling factor corresponding to the edge loss function and the edge consistency loss function;
[0239] The composite loss function is determined according to the first scaling factor and the second scaling factor.
[0240] In the embodiment of the present invention, the composite loss function is: L = L CE +λL Edge +βL Mag ,
[0241] Among them, L is the composite loss function, λ is the first scale factor, and L Edge is the edge loss function, β is the second scale factor, L Mag is the edge consistency loss function.
[0242] In some embodiments, the original loss function is a cross entropy loss function.
[0243] In some embodiments, the cross entropy loss function is:
[0244]
[0245] Among them, L CE is the cross entropy loss function, is the value of the i-th position in the true probability map of the image to be detected, It is the value of the i-th position in the semantic probability map of the image to be detected predicted by the neural network.
[0246] It should be noted that the above device can execute the method provided in the embodiment of the present application, and has the functional modules and beneficial effects corresponding to the execution method. For technical details not fully described in the device embodiment, please refer to the method provided in the embodiment of the present application.
[0247] In an embodiment of the present invention, an image semantic segmentation device is provided, which is applied to an electronic device. The device includes: a semantic probability map unit, which is used to obtain an image to be detected and determine a semantic probability map corresponding to the image to be detected; a truth map unit, which is used to determine a truth map corresponding to the image to be detected; an object edge detection map unit, which is used to determine an object edge detection map according to the semantic probability map; a mask edge detection map unit, which is used to determine a mask edge detection map according to the truth map; an edge loss function unit, which is used to determine an edge loss function according to the object edge detection map and the mask edge detection map; an original image edge detection map unit, which is used to detect the image to be detected; The detected image is subjected to edge detection to determine the edge detection map of the original image; an edge consistency loss function unit is used to determine the edge consistency loss function according to the edge detection map of the original image and the edge detection map of the mask; a composite loss function unit is used to determine the composite loss function according to the original loss function, the edge loss function and the edge consistency loss function, wherein the original loss function is used to measure the difference between the semantic probability map corresponding to the image to be detected and the true probability map; a semantic segmentation unit is used to perform training based on the composite loss function to obtain an image semantic segmentation model, and perform semantic segmentation on the image to be detected based on the image semantic segmentation model. On the one hand, by adding the edge loss function and the edge consistency loss function to generate a composite loss function, the semantic segmentation can process the edge more accurately; on the other hand, by training the composite loss function to obtain the image semantic segmentation model, the present invention can improve the accuracy of semantic segmentation.
[0248] See also Figure 8 , Figure 8 A schematic diagram of the hardware structure of an electronic device according to various embodiments of the present invention;
[0249] like Figure 8As shown, the electronic device 80 includes but is not limited to: a radio frequency unit 81, a network module 82, an audio output unit 83, an input unit 84, a sensor 85, a display unit 86, a user input unit 87, an interface unit 88, a memory 89, a processor 810, and a power supply 811, and the electronic device 80 also includes a camera. Those skilled in the art will understand that Figure 8 The structure of the electronic device shown in the figure does not constitute a limitation on the electronic device, and the electronic device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently. In the embodiments of the present invention, the electronic device includes but is not limited to a television, a mobile phone, a tablet computer, a laptop computer, a PDA, a vehicle-mounted terminal, a wearable device, and a pedometer.
[0250] Processor 810 is used to obtain an image to be detected, determine a semantic probability map and a true value map corresponding to the image to be detected; determine an object edge detection map based on the semantic probability map; determine a mask edge detection map based on the true value map; determine an edge loss function based on the object edge detection map and the mask edge detection map; perform edge detection on the image to be detected, determine an original image edge detection map, and determine an edge consistency loss function based on the original image edge detection map and the mask edge detection map; determine a composite loss function based on the original loss function, the edge loss function and the edge consistency loss function, wherein the original loss function is used to measure the gap between the semantic probability map corresponding to the image to be detected and the true probability map; perform training based on the composite loss function to obtain an image semantic segmentation model, and perform semantic segmentation on the image to be detected based on the image semantic segmentation model.
[0251] In an embodiment of the present invention, on the one hand, by adding an edge loss function and an edge consistency loss function to generate a composite loss function, the semantic segmentation processing of edges is made more accurate. On the other hand, by training the composite loss function to obtain an image semantic segmentation model, the present invention can improve the accuracy of semantic segmentation.
[0252] It should be understood that in the embodiment of the present invention, the radio frequency unit 81 can be used for receiving and sending signals during information transmission or communication. Specifically, after receiving downlink data from the base station, it is sent to the processor 810 for processing; in addition, uplink data is sent to the base station. Generally, the radio frequency unit 81 includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier, a duplexer, etc. In addition, the radio frequency unit 81 can also communicate with the network and other devices through a wireless communication system.
[0253] The electronic device 80 provides the user with wireless broadband Internet access through the network module 82, such as helping the user to send and receive emails, browse web pages, and access streaming media.
[0254] The audio output unit 83 can convert the audio data received by the RF unit 81 or the network module 82 or stored in the memory 89 into an audio signal and output it as sound. Moreover, the audio output unit 83 can also provide audio output related to a specific function performed by the electronic device 80 (for example, a call signal reception sound, a message reception sound, etc.). The audio output unit 83 includes a speaker, a buzzer, a receiver, etc.
[0255] The input unit 84 is used to receive audio or video signals. The input unit 84 may include a graphics processor (GPU) 841 and a microphone 842, and the graphics processor 841 processes the target image of a static picture or video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The processed image frame can be displayed on the display unit 86. The image frame processed by the graphics processor 841 can be stored in the memory 89 (or other storage medium) or sent via the radio frequency unit 81 or the network module 82. The microphone 842 can receive sound and can process such sound into audio data. The processed audio data can be converted into a format output that can be sent to a mobile communication base station via the radio frequency unit 81 in the case of a telephone call mode.
[0256] The electronic device 80 also includes at least one sensor 85, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor includes an ambient light sensor and a proximity sensor, wherein the ambient light sensor can adjust the brightness of the display panel 861 according to the brightness of the ambient light, and the proximity sensor can turn off the display panel 861 and / or the backlight when the electronic device 80 is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary, which can be used to identify the posture of the electronic device (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; the sensor 85 can also include a fingerprint sensor, a pressure sensor, an iris sensor, a molecular sensor, a gyroscope, a barometer, a hygrometer, a thermometer, an infrared sensor, etc., which will not be repeated here.
[0257] The display unit 86 is used to display information input by the user or information provided to the user. The display unit 86 may include a display panel 861, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.
[0258] The user input unit 87 can be used to receive input digital or character information, and to generate key signal input related to user settings and function control of the electronic device. Specifically, the user input unit 87 includes a touch panel 871 and other input devices 872. The touch panel 871, also known as a touch screen, can collect the user's touch operation on or near it (such as the user's operation on the touch panel 871 or near the touch panel 871 using any suitable object or accessory such as a finger, stylus, etc.). The touch panel 871 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch orientation, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact coordinates, and then sends it to the processor 810, receives the command sent by the processor 810 and executes it. In addition, the touch panel 871 can be implemented in various types such as resistive, capacitive, infrared and surface acoustic waves. In addition to the touch panel 871, the user input unit 87 may also include other input devices 872. Specifically, other input devices 872 may include, but are not limited to, a physical keyboard, function keys (such as volume control keys, switch keys, etc.), a trackball, a mouse, and a joystick, which are not described in detail here.
[0259] Furthermore, the touch panel 871 may be overlaid on the display panel 861. When the touch panel 871 detects a touch operation on or near it, it transmits the information to the processor 810 to determine the type of touch event. Then, the processor 810 provides corresponding visual output on the display panel 861 according to the type of touch event. Figure 8 In the figure, the touch panel 871 and the display panel 861 are two independent components to realize the input and output functions of the electronic device. However, in some embodiments, the touch panel 871 and the display panel 861 can be integrated to realize the input and output functions of the electronic device, which is not limited here.
[0260] The interface unit 88 is an interface for connecting an external device to the electronic device 80. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, etc. The interface unit 88 may be used to receive input (e.g., data information, power, etc.) from an external device and transmit the received input to one or more elements within the electronic device 80 or may be used to transmit data between the electronic device 80 and an external device.
[0261] The memory 89 can be used to store software programs and various data. The memory 89 can mainly include a program storage area and a data storage area, wherein the program storage area can store at least one application 891 required for a function (such as a sound playback function, an image playback function, etc.) and an operating system 892, etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory 89 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0262] The processor 810 is the control center of the electronic device. It uses various interfaces and lines to connect various parts of the entire electronic device. It executes various functions of the electronic device and processes data by running or executing software programs and / or modules stored in the memory 89, and calling data stored in the memory 89, so as to monitor the electronic device as a whole. The processor 810 may include one or more processing units; preferably, the processor 810 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 810.
[0263] The electronic device 80 may also include a power source 811 (such as a battery) for supplying power to each component. Preferably, the power source 811 may be logically connected to the processor 810 through a power management system, thereby implementing functions such as charging, discharging, and power consumption management through the power management system.
[0264] In addition, the electronic device 80 includes some functional modules not shown, which will not be described in detail here.
[0265] Preferably, an embodiment of the present invention further provides an electronic device, including a processor 810, a memory 89, and a computer program stored in the memory 89 and executable on the processor 810. When the computer program is executed by the processor 810, each process of the above-mentioned image semantic segmentation method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it will not be described here.
[0266] The embodiment of the present invention further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by one or more processors, each process of the above-mentioned image semantic segmentation method embodiment is implemented, and the same technical effect can be achieved. To avoid repetition, it is not repeated here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0267] It should be noted that, in this article, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprises a ..." does not exclude the existence of other identical elements in the process, method, article or device including the element.
[0268] The above described device or equipment embodiments are merely illustrative, wherein the unit modules described as separate components may or may not be physically separated, and the components displayed as module units may or may not be physical units, that is, they may be located in one place, or may be distributed on multiple network module units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment.
[0269] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus a necessary general hardware platform, and of course by hardware, but in many cases the former is a better implementation method. Based on such an understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for a terminal (which can be a mobile terminal, a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment of the present invention or some parts of the embodiments.
[0270] Finally, it should be noted that the embodiments described above in conjunction with the accompanying drawings are only used to illustrate the technical solutions of the present invention. The present invention is not limited to the above-mentioned specific implementation methods, which are merely illustrative and not restrictive. Under the idea of the present invention, the technical features in the above-mentioned embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other changes in different aspects of the present invention as described above, which are not provided in detail for the sake of simplicity. Although the present invention has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that they can still modify the technical solutions recorded in the above-mentioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.
Claims
1. An image semantic segmentation method, applied to electronic equipment, characterized in that: The method comprises: Acquire an image to be detected, and determine a semantic probability map and a truth map corresponding to the image to be detected; Determining an object edge detection map according to the semantic probability map; Determining a mask edge detection map according to the true value map; Determining an edge loss function according to the object edge detection map and the mask edge detection map; Performing edge detection on the image to be detected to determine an original image edge detection map, and determining an edge consistency loss function according to the original image edge detection map and the mask edge detection map; Determine a first scaling factor and a second scaling factor corresponding to the edge loss function and the edge consistency loss function; Determine a composite loss function according to the first scale factor and the second scale factor, wherein the composite loss function is a sum of weighted sums of the edge loss function and the edge consistency loss function according to the first scale factor and the second scale factor, and is obtained by adding the composite loss function to the original loss function, the original loss function is used to measure the gap between the semantic probability map corresponding to the image to be detected and the true probability map, and the original loss function is a cross entropy loss function; Training is performed based on the composite loss function to obtain an image semantic segmentation model, and semantic segmentation is performed on the image to be detected based on the image semantic segmentation model.
2. The method according to claim 1, characterized in that Determining an object edge detection map according to the semantic probability map includes: Performing horizontal filtering on the semantic probability map to determine a horizontal object edge detection map; Performing vertical filtering on the semantic probability map to determine a vertical object edge detection map; The horizontal object edge detection map and the vertical object edge detection map are combined along the channel direction to determine the object edge detection map.
3. The method according to claim 1, characterized in that: Determining a mask edge detection map according to the true value map includes: Performing horizontal filtering on the true value image to determine a horizontal mask edge detection image; Performing vertical filtering on the true value image to determine a vertical mask edge detection image; The horizontal mask edge detection map and the vertical mask edge detection map are combined along a channel direction to determine the mask edge detection map.
4. The method according to claim 3, characterized in that The step of determining an edge loss function according to the object edge detection map and the mask edge detection map comprises: The object edge detection image and the mask edge detection image are subjected to norm processing to construct an edge loss function.
5. The method according to claim 4, characterized in that The edge loss function is: Among them, L Edge is the marginal loss function, is the value of the i-th position in the true probability map of the image to be detected, For The value after filtering, is the value of the i-th position in the semantic probability map of the image to be detected predicted by the neural network, For The value after filtering, p is a positive integer and p≥2.
6. The method according to claim 1, characterized in that The performing edge detection on the image to be detected to determine the edge detection map of the original image includes: Performing horizontal filtering on the image to be detected to determine a horizontal original image edge detection image; Performing vertical filtering on the image to be detected to determine a vertical original image edge detection image; The horizontal original image edge detection map and the vertical original image edge detection map are merged along the channel direction to determine the original image edge detection map.
7. The method according to claim 6, characterized in that The step of determining an edge consistency loss function according to the original image edge detection map and the mask edge detection map comprises: The edge consistency loss function is determined according to the horizontal original image edge detection map and the vertical original image edge detection map.
8. The method according to claim 6 or 7, characterized in that: The edge consistency loss function is: Among them, L Mag is the edge consistency loss function, I x is the horizontal original image edge detection image, I y is the vertical original image edge detection image, M x is the horizontal mask edge detection map, M y It is the vertical mask edge detection map.
9. The method according to claim 1, characterized in that: The composite loss function is: L = L CE +λL Edge +βL Mag , Among them, L is the composite loss function, L CE is the original loss function, λ is the first scaling factor, L Edge is the edge loss function, β is the second scale factor, L Mag is the edge consistency loss function.
10. The method according to claim 1, characterized in that The cross entropy loss function is: Among them, L CE is the cross entropy loss function, is the value of the i-th position in the true probability map of the image to be detected, It is the value of the i-th position in the semantic probability map of the image to be detected predicted by the neural network.
11. An image semantic segmentation device, applied to electronic equipment, characterized in that: The device comprises: A semantic probability map unit, used to obtain an image to be detected and determine a semantic probability map corresponding to the image to be detected; A truth map unit, used to determine a truth map corresponding to the image to be detected; An object edge detection map unit, used to determine an object edge detection map according to the semantic probability map; A mask edge detection map unit, used to determine a mask edge detection map according to the true value map; An edge loss function unit, used to determine an edge loss function according to the object edge detection map and the mask edge detection map; An original image edge detection map unit is used to perform edge detection on the image to be detected and determine an original image edge detection map; An edge consistency loss function unit, used to determine an edge consistency loss function according to the original image edge detection map and the mask edge detection map, wherein the original loss function is used to measure the gap between the semantic probability map corresponding to the image to be detected and the true probability map; A composite loss function unit, used to determine a first scale factor and a second scale factor corresponding to the edge loss function and the edge consistency loss function; according to the first scale factor and the second scale factor, determine a composite loss function, the composite loss function is a sum of weighted sums of the edge loss function and the edge consistency loss function according to the first scale factor and the second scale factor, and is added to the original loss function, the original loss function is used to measure the gap between the semantic probability map corresponding to the image to be detected and the true probability map, and the original loss function is a cross entropy loss function; The semantic segmentation unit is used to perform training based on the composite loss function to obtain an image semantic segmentation model, and perform semantic segmentation on the image to be detected based on the image semantic segmentation model.
12. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image semantic segmentation method as described in any one of claims 1-10.
13. A non-volatile computer-readable storage medium, characterized in that: The non-volatile computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to enable an electronic device to execute the image semantic segmentation method as described in any one of claims 1-10.
Citation Information
Patent Citations
Character detection method and device
CN105574513A
Image segmentation model training method and device
CN111260665A